
Last week we covered what counts as proof that a learning game works and how to build reporting teachers will open. Both depend on a question that must be answered by the design team: how do you gather evidence of learning without stopping the game to administer a test? Our design philosophy avoids the “chocolate-covered broccoli” approach, where players are presented with traditional curriculum content and then rewarded by straightforward arcade gameplay after completing the traditional content. Learning games can be so much more than that! Stealth assessment is the established answer to this problem, and it carries a longer research history than designers might expect.
Valerie Shute introduced stealth assessment roughly two decades ago to describe measurement embedded inside a digital environment, so that inferences about a learner's skills are drawn continuously from ordinary interaction data instead of a separate test. A 2026 special issue of the Journal of Research on Technology in Education surveys where the field stands now, covering current trends, new directions, and the ethical questions the approach raises. The consistent finding across twenty years of work is that embedded measurement holds up psychometrically when the measurement is specified in advance of production.
Stealth assessment rests on evidence-centered design, formalized by Mislevy, Steinberg, and Almond in 2003. It separates an assessment into three linked models. The competency model describes the skills you want to infer. The evidence model specifies which player behaviors bear on which competencies, expressed as conditional probabilities. The task model describes the situations that elicit those behaviors. Concepting and designing for all three before production is what elevates stealth assessment from mere data collection around player actions and after-the-fact analysis.
Of the three models, the task model will feel most familiar to an educational game designer, since it describes situations that reliably produce the behavior you need to observe. This is level design with one added constraint, because each situation has to present a meaningful choice whose options carry different evidentiary weight. A puzzle with a single correct path produces very little evidence about the player's reasoning. A system that permits several routes to success tells you which route this player reached for, which is the observable your evidence model needs.
Beats Empire, built with Teachers College at Columbia University alongside SRI, UW-Madison, Georgia Tech, and Digital Promise under NSF funding, was designed from the outset as a formative assessment game. Players manage a music studio, analyze listener data across the boroughs of a fictional city, and use what they find to sign artists, record songs, and target releases. Data interpretation is the business decision in question, which means every action a player takes to advance their studio doubles as an observable insight about how they read and apply data. Those actions are logged and surfaced live on a teacher dashboard with suggestions for how to support students who are struggling.
Building an assessment game presents different problems than building a learning game, and the sharpest of them is keeping the experience open-ended while it produces usable measurement. A game that funnels every player down one legible path will yield clean data about an experience nobody wanted to play. Beats Empire addresses this by letting players define their own version of a successful studio, so the evidence comes from how a player pursues the goals they set rather than from compliance with goals the game assigned. The project's design process is documented at length in Playful Testing: Designing a Formative Assessment Game for Data Science.
Shute's work is direct about the need to check in-game measures against external ones, and this step is instrumental in making data-based claims that are both viable and plausible. Run a validated instrument alongside the game with a subset of players, then compare what your evidence model inferred against what the instrument reported. Disagreement between the two is useful information at this stage, since it usually points at a task that failed to elicit the behavior you designed it for rather than at a flaw in the underlying competency model.
Stealth in this context refers to the absence of interruption in the player's experience. Learners, teachers, and families should still understand that the game measures performance, what it records, and where that information goes. Building this disclosure into onboarding and into your data documentation costs very little and aligns the product with the COPPA and FERPA requirements governing student data. The 2026 special issue treats these questions as central to the method, which aligns to what districts now ask during procurement.
~
The appeal of stealth assessment for a design team is that it moves measurement out of the pause menu and into the mechanics, where the player was already making the decisions you wanted to observe. Getting there requires committing to the competency, the evidence, and the tasks early, while the core loop can still be shaped around them. Interested in building assessment into your next game? Let's talk!
Best practices for preventing motion sickness while maximizing learning outcomes.
Best practices for preventing motion sickness while maximizing learning outcomes.