Skip to content

Research

Adaptive robotics education and digital twin research

HCR is an academic research instrument for Blockly programming assessment, adaptive learning, and digital-twin robotics education. Participants choose whether to join before the client creates a player or session.

Participation

Choose whether to participate before the study starts

The client creates no simulator session or player identifier until a person actively chooses to participate in this academic study.

Required study response

Compiled Blockly Program IR, challenge version, score, and technical session or submission identifiers required for replay and analysis.

Optional context

Primary language as a coarse code such as zh or en, and current UTC offset in minutes. The primary Participate action enables both; More settings can disable either one.

Not intentionally collected

No browser fingerprint, precise IANA time-zone name, inferred location, advertising identifier, or cross-site tracking profile.

How we describe the data

We do not call the data fully anonymous. Study records may retain session identifiers, and the server keeps routine security logs, so we describe the data as de-identified.

Framing

Two problems, one system

Pedagogical

A class moves at one speed

Hardware costs and supervision limit robotics teaching, so a class often receives one fixed sequence. Students do not learn at one fixed pace, and the mismatch affects both ends of the class.

Measurement

Programming ability is hard to score

Unlike a multiple-choice answer, a program has an open-ended response space and a continuous outcome. Standard adaptive-testing methods assume neither, so they must be adapted before they can measure programming performance.

Contributions

What is new here

  1. 01

    Adaptive selection over a generated item bank

    Challenges do not come from a fixed list. Item families generate and calibrate candidates, then Fisher information selects one near the learner’s current ability estimate. The bank can grow without someone authoring every level by hand.

  2. 02

    Guaranteed-solvable generated items

    Procedural generation and reachability are in tension: a plausible-looking target may be unreachable. Solving each candidate before it is served converts that from a hope into a precondition.

  3. 03

    Continuous scores in a dichotomous estimator

    Programming performance is continuous; the estimator available is not. An order-preserving remap around a per-item mastery threshold preserves the ordering while keeping the raw score intact for analysis.

  4. 04

    Determinism as a fairness property

    The server replays each program and estimates time from joint travel. Client hardware cannot change the result, which is necessary if a score is used competitively or diagnostically.

Method

How the adaptive layer is built

After every attempt the platform re-estimates how the learner is doing and picks the next challenge to match. Finish comfortably and the next one is harder; struggle and it steps back. One level per student, not one level per class.

Model

Two-parameter logistic

Ability θ on a logit scale, item difficulty b, discrimination a. The guessing parameter is fixed at zero: the response space is a program, not a set of choices, so there is nothing to guess into. Estimating a third parameter here would fit noise.

Response

The program is the response

A learner’s Program IR is replayed server-side and scored. The normalized score is remapped around the item’s mastery threshold τ so that “above τ” and “mastered” coincide, then passed to the estimator. The raw score is persisted separately.

Selection

Information, then exposure control

Candidates are ranked by information at the current θ and then exposure-capped, so learners at the same level do not all receive the same items and the bank is not burned through.

Generation

Difficulty as a target, not an outcome

Item families use features that predict difficulty, including clearance, reachability strain, budget pressure, and loop structure. Each candidate is solved for a target b and rejected during generation if the reference solver finds no solution.

Calibration

Provisional until it has evidence

A new item enters provisional: exposure-capped and excluded from ability updates that count. Once it has responses, difficulty is refit by Newton iteration on the marginal likelihood with θ held at posterior means.

This version treats ability as one composite θ and records dimension tags on every response for reporting. A genuinely multidimensional model is a future step, not a claim made here.

Validity

An item nobody can finish measures nothing

Drawing a target haircut first and hoping the arm can reach it does not work. One early challenge asked for 91 pieces of hair when the arm could reach only 20. It was impossible, yet the level gave the learner no warning.

Now the solver runs first. A candidate target is swept against the collision-free joint space, then handed to a reference solver; anything the solver cannot finish is rejected at generation and never served. The solution becomes the level’s reference cost and reference time.

The reachable set is the union of every voxel the tool can contact while sweeping the collision-free joint space. Because that sweep is expensive, it is computed once and cached as a fixture. Tests then verify that the reference solution can be replayed and that no target lies in a dead zone.

The failure that motivated it

91target voxels requested by a hand-drawn challenge
20reachable by the arm under head-clearance constraints

A learner could mistake that structural impossibility for personal failure. The ability estimate would then treat the failure as evidence.

Reproducibility

What a result depends on

Deterministic generation

Hairstyle generation is a pure function of the challenge configuration. The same configuration always yields the same ordered voxel set, with no randomness in the pipeline.

Server-side replay

Competitive scores come from replaying the submitted program on the service, never from a number the browser reports.

Estimated, not measured, time

Execution time is derived from joint travel and configured speeds, so it is a property of the program rather than of the machine that ran it.

Frozen wire contract

The app and service share the same scoring and program types. Protocol changes are additive, so older clients keep working.

Elsewhere

Papers and people