# Research Basis and Logic Model

Version: 2026-07-22.3  
Evidence claim: design rationale only; no demonstrated product effect

## Problem

Open-ended sandbox play can be engaging without producing clear STEM goals, reliable feedback, reflection, transfer, or valid evidence of learning.

## Logic model

Inputs include the construction/physics environment, a validated local concept catalog, non-identifying play signals, worked physical examples, specific feedback, learner support and pace controls, and human review.

Activities are to observe relevant play, select a bounded concept, label AI content, model an example, invite criterion-linked practice, provide one actionable next step, fade or add support, invite a changed-context task, and allow an override or concern report.

Outputs include validated teaching plans, criterion-level attempts, performance-only evidence, source provenance, band transitions, overrides, concern reports, and candidate curriculum crosswalks.

Short-term outcomes proposed for study are clearer task goals, more actionable error feedback, more revision of physical models, and greater learner agency. Long-term outcomes proposed for study are improved transfer of selected STEM concepts and stronger planning, testing, revision, and reflection. These are hypotheses, not established results.

## Research-linked design decisions

1. Model, then practice: mentors construct a worked physical example before guided or independent practice. The programming study below supports worked examples and metacognitive scaffolding, but transfer to this sandbox must be tested.
2. Criterion-specific feedback: feedback names the failed criterion and gives one next action. Feedback and metacognitive-prompt research informs this design; product effectiveness is unknown.
3. Retrieval and retention: the evidence ledger schedules later retrieval and never awards mastery after one success. Retrieval research supports this rationale, but game performance is not yet a validated measure.
4. Adaptive scaffolding: the mentor adjusts modeling, granularity, support, and challenge from completed assessment evidence. Intelligent-tutoring research informs the design; fairness and classroom validity require study.
5. Structured play and choice: players retain material, shape, construction, support, and pace choices inside a visible goal. The secondary play-based research base is limited, so this is a cautious hypothesis.
6. Human agency over AI: generated teaching is labeled; users can skip a recommendation without penalty and report a concern. A middle-school teachable-agent study informs the design, but educator usability must be evaluated.

## Empirical sources

- Karpicke, J. D., and Roediger, H. L. (2008). *The Critical Importance of Retrieval for Learning*. https://doi.org/10.1126/science.1152408
- Shin and colleagues (2023). *The Effects of Worked-Out Example and Metacognitive Scaffolding on Problem-Solving Programming*. https://doi.org/10.1177/07356331231174454
- Hattie, J., and Timperley, H. (2007). *The Power of Feedback*. https://doi.org/10.3102/003465430298487
- Guo, L. (2022). *Using Metacognitive Prompts to Enhance Self-Regulated Learning and Learning Outcomes*. https://doi.org/10.1111/jcal.12650
- Dean, C. G., and Wenner, J. A. (2025). *Patterns and Representation in Play-Based Learning*. https://doi.org/10.3389/feduc.2025.1557001
- VanLehn, K. (2011). *The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems*. https://doi.org/10.1080/00461520.2011.611369
- Xing and colleagues (2025). *Development of a Generative AI-Powered Teachable Agent for Middle School Mathematics Learning*. https://doi.org/10.1111/bjet.13586

## Applicability limits

The cited populations, subjects, interventions, and settings differ from this game. The sources justify testable design choices, not claims that the product improves learning. Separate middle-school, high-school, and college studies are required. College evaluation must use suitable higher-education outcomes and cannot be inferred from the middle-school teachable-agent evidence.

## Evidence plan

Start with externally reviewed qualitative studies for each band. Revise the product from recorded findings. Then preregister quantitative studies using validated, concept-relevant measures, appropriate comparison conditions, sufficient samples, transparent exclusions, attrition reporting, and independent analysis. Publish favorable, null, and unfavorable findings.
