Candy Crush is often described through its number of levels, but scale is the consequence of a repeatable design system. King's public GDC material emphasizes theory, thought, tools, testing, and data-informed tuning—useful ideas for any puzzle team.
The aim is not to reproduce Candy Crush boards. Use the principles to design original goals and blockers, then test them inside a match-3 prototype or compare other puzzle games.
Quick read
Key takeaways
- Give each level one dominant lesson or tension before combining mechanics.
- Tune goal, moves, spawn rules, blockers, and board geometry as one system.
- Use simulation and analytics to find outliers, then keep human playtests for fairness and fun.
- Design difficulty as a rhythm of pressure and relief rather than a straight climb.
Start With a Level Contract
Write the intended goal, primary decision, new or reinforced mechanic, target tension, expected failure mode, and role in the surrounding sequence. A level can be difficult for the wrong reason, so name what the player is meant to learn or master.
Keep early drafts simple enough to diagnose. When board geometry, blockers, spawn rules, special pieces, and multiple goals all change together, playtest feedback cannot reveal which variable created the problem.
Map the Difficulty Drivers
Moves limit opportunity; board shape limits space; blockers consume actions; spawn rules change probability; goals decide which matches matter. Track these levers explicitly and change one or two at a time.
King's published GDC material describes tools and testing as part of level design. The transferable lesson is that scalable content needs a data representation, fast editor, repeatable simulation, and visible history—not thousands of one-off scenes.

Put Luck Where It Creates Variety, Not Confusion
Randomness can create replayable boards, surprising chains, and different tactical paths. It becomes frustrating when outcomes feel disconnected from the player's decisions or when an unwinnable opening is not detected.
Record reshuffles, dead boards, pass rate, remaining moves, booster dependence, and failure reason. These signals locate candidates for review; they do not decide whether the level is enjoyable.
Design a Difficulty Rhythm
A sequence can alternate teaching, practice, challenge, recovery, and recombination. Microsoft's 2026 discussion with King and Mojang describes continuous rebalancing in long-running games, underscoring that difficulty is an ongoing relationship with the audience.
Review the beat chart, not only individual pass rates. Several fair levels can still create a tiring sequence if they demand the same kind of attention without relief.
Combine Simulation, Analytics, and Human Play
Simulation can estimate solvability and reveal extreme distributions. Analytics can show where real players stop or repeat. Human playtests can explain why the board feels clever, dull, unfair, or unreadable.
Pair the level-design review with the match-3 art guide so the board communicates goals and state as clearly as the underlying rules deserve.
Worked Example: Design a Five-Level Teaching Sequence
Introduce one original blocker in level one with generous moves and a board that makes the intended interaction likely. Level two asks the player to repeat it under a mild spatial constraint. Level three combines it with a familiar goal. Level four creates a focused challenge, and level five provides relief while asking for one strategic variation.
Track the sequence as a beat chart: concept introduced, concept practiced, difficulty driver, expected failure, target emotional role, and observed player behavior. A single pass-rate number cannot show whether the level teaches, surprises, frustrates, or simply takes too long.
| Level role | Main variable | Evidence to collect |
|---|---|---|
| Introduce | Clear board and generous moves | Can players state the blocker rule? |
| Practice | Repeated interaction | Do they plan around it? |
| Combine | Add one familiar mechanic | Which rule is forgotten? |
| Challenge | Tighter move economy or geometry | Failure reasons and remaining moves |
| Relief | Lower pressure with a twist | Does confidence recover? |
Balance With a Difficulty-Driver Model
Represent board geometry, move limit, goal quantity, blocker layers, spawn distribution, special-piece opportunity, reshuffle rules, and tutorial support as explicit variables. This makes tuning reproducible and allows simulations to compare versions.
Separate difficulty from friction. A level can be statistically hard because it creates interesting planning, or because the goal is hidden, the opening is unstable, or the player must wait through repeated cascades. Analytics locates the problem; replay review and interviews explain it.
- Goal pressure: how much useful work each move must accomplish
- Space pressure: how board shape and blockers restrict options
- Probability pressure: how spawn and refill rules affect opportunity
- Knowledge pressure: what the player must recognize or remember
- Execution friction: clarity, animation delay, mis-taps, and repeated setup
Investigate Difficulty Outliers Without Flattening the Game
When pass rate moves outside the intended range, segment by player experience, booster use, attempts, version, and entry state. Inspect first failures and late failures separately. New players may misunderstand the goal while experienced players may encounter a probability bottleneck near completion.
King's GDC talks describe tools, testing, and data as parts of scalable level design. Use automated analysis to prioritize human review, not to replace intent. The goal is a designed rhythm, not identical pass rates across every level.
| Symptom | Likely cause | Next check |
|---|---|---|
| Many failures in first moves | Goal or mechanic unclear | Simplify opening and improve teaching |
| Failures cluster near completion | Move economy or final spawn bottleneck | Review target and refill rules |
| High reshuffle rate | Opening distribution or geometry | Add opening validation |
| Booster dependence | Core board lacks viable paths | Test without power assistance |
| Players quit after winning | Sequence fatigue or weak next goal | Review beat chart and reward pacing |
Match-3 Level Review Template
Review the level as data, as a simulated distribution, and as a human experience. Each view catches a different failure. Store the designer's intent beside the level configuration so future tuning does not erase the reason it exists.
Test neighboring levels together. A fair challenge can still be poorly placed after several high-pressure boards or before a new mechanic that also demands attention.
- The level has one stated learning, mastery, or pacing role.
- Goal, moves, board shape, blockers, spawn rules, and special opportunities are versioned.
- Opening boards are checked for dead or misleading states.
- Simulation results include distributions, not only average completion.
- Fresh-player and experienced-player replays explain major failure clusters.
- The surrounding beat chart contains pressure, variation, and recovery.
What the Primary Sources Establish About Candy Crush Level Design
Our evidence baseline starts with the GDC: Level Design Saga, accessed August 20, 2026. We use it to establish documented behavior, terminology, or constraints—not to claim that the source endorses Elseland's workflow or conclusions. The practical artifact under review is a level hypothesis, simulator distribution, observed playtest sequence, and documented tuning history.
That distinction is central to E-E-A-T. A first-party page can establish what a format, tool, platform, model, or game team publicly documents. It cannot prove that a particular asset is fast, accessible, legally cleared, fun, or production-ready. Those conclusions require separate observation, measurement, specialist review, or player evidence tied to the actual project.
For this topic, the decision is whether the level creates the intended lesson and pressure without relying on opaque frustration. The following observations turn the official reference into a reviewable production record rather than a decorative citation:
| Evidence layer | What it can support | What it cannot support alone |
|---|---|---|
| Official source | Documented feature, rule, format, or published design context | Project-specific quality or universal performance |
| Project measurement | Observed behavior in a named build, scene, device, or sample | Unmeasured platforms or future versions |
| Human review | Usability, visual, editorial, and production judgment | Legal certainty or population-level player behavior |
| Release record | Who approved what, when, with which evidence | Permanent compliance after inputs or rules change |
- 1. separate goal, board geometry, blockers, move economy, refill, and special-piece effects. Store the result with the asset or build identifier so another reviewer can reproduce the conclusion.
- 2. model distributions rather than trusting one successful run. Store the result with the asset or build identifier so another reviewer can reproduce the conclusion.
- 3. observe where players look and what they believe caused failure. Store the result with the asset or build identifier so another reviewer can reproduce the conclusion.
- 4. review a level in the sequence that teaches and tests its mechanics. Store the result with the asset or build identifier so another reviewer can reproduce the conclusion.

A Field Review Protocol for Candy Crush Level Design
Use this protocol after the first plausible output exists and before scaling the workflow. Keep one untouched baseline, one candidate revision, and one deliberately stressed case. The stressed case should expose the topic's likely failure mode—crowded scenes, extreme poses, small-screen play, unusual inputs, or a release-rule change—rather than merely repeat the easiest success case.
Run the review in the real delivery context whenever possible. Capture the tool or model version, source files, settings, target device or engine, date, and reviewer. If the work depends on a changing external service, record the response or exported artifact instead of assuming the same output can be recreated later.
A useful review ends with a decision and a next action. “Looks good” is not a gate. State whether the candidate passes, passes with a bounded exception, needs revision, or should be rejected; identify the evidence behind that status and the owner of the next check.
| Review status | Meaning | Required next action |
|---|---|---|
| Pass | All defined visual, technical, and release gates are supported by evidence | Freeze the reviewed artifact and link it to the build |
| Conditional pass | A known limitation is bounded and does not invalidate the intended use | Document the exception, owner, and trigger for re-review |
| Revise | The direction is viable but one or more gates remain unsupported | Change one controlled variable and repeat the affected checks |
| Reject | The candidate conflicts with the intended use, evidence, rights, safety, or budget | Preserve the record and choose a different approach |
- write the intended player lesson and emotional curve. Record the expected result before the check, then attach the observed result and any exception after it.
- identify controllable and stochastic difficulty drivers. Record the expected result before the check, then attach the observed result and any exception after it.
- simulate enough runs to inspect tails as well as averages. Record the expected result before the check, then attach the observed result and any exception after it.
- playtest with people who did not author the level. Record the expected result before the check, then attach the observed result and any exception after it.
- classify failures as comprehension, planning, execution, or variance. Record the expected result before the check, then attach the observed result and any exception after it.
- record each tuning change and its expected effect before retesting. Record the expected result before the check, then attach the observed result and any exception after it.
Expert Interpretation and Limits of This Game Guides Guide
The strongest conclusion this guide can support is a conditional production recommendation: use the workflow when its documented assumptions match the project, and keep the evidence needed to revisit the decision. We do not infer universal model quality, player preference, legal clearance, or performance from an official screenshot, a provider example, or a single successful asset.
Experience matters here because candy crush level design crosses creative judgment and implementation detail. The practical review should include the people who will edit the source, integrate the result, test it in play, maintain it after release, and answer rights or policy questions. A narrow expert handoff often misses problems that appear only when those responsibilities meet.
Before publishing or shipping, repeat time-sensitive checks against the current source and exact build. Preserve dated evidence, disclose the evaluation method, and distinguish measured results from editorial inference. That record is more valuable than a confident conclusion that future reviewers cannot reproduce.
| Claim type | Editorial treatment |
|---|---|
| Documented fact | Link to GDC: Level Design Saga and include the access date |
| Observed project result | Name the build, environment, sample, and method |
| Expert judgment | State the criteria, reviewer role, and tradeoff |
| Inference or forecast | Label it explicitly and describe what evidence could change it |
- Public talks explain principles but do not expose current proprietary level data.
- A Candy Crush case is an analogy, not a transferable formula for every audience.
- Simulation quality depends on how accurately the agent represents player choices.
- Completion rate alone cannot explain fairness, clarity, or long-term satisfaction.
Frequently asked questions
Are Candy Crush levels procedurally generated?
King's public talks describe a level-design workflow supported by tools, testing, simulation, and data. Treat levels as authored data that can be analyzed and tuned, not as a claim that every board is generated automatically.
What makes a match-3 level difficult?
Difficulty emerges from goal, move count, board shape, blockers, spawn rules, available combinations, randomness, and the player's prior knowledge.
Should puzzle difficulty always increase?
No. A varied rhythm can introduce mechanics, provide practice, create peaks, and offer recovery. The sequence matters as much as each level.
Can AI design match-3 levels?
AI and simulation can draft or analyze boards, but designers should still define intent, inspect fairness, review unusual strategies, and playtest the experience.
What metrics matter for match-3 level design?
Pass rate, attempts, remaining moves, reshuffles, booster use, churn, completion time, and failure location can all help. Segment the data and combine it with replay or playtest evidence because the same number can reflect very different player experiences.
How can simulations help puzzle designers?
Simulations can estimate solvability, opportunity distribution, dead boards, and sensitivity to rule changes. They are strongest as an outlier detector and comparison tool; human review still judges clarity, strategy, pacing, and fun.
What is a beat chart for puzzle levels?
It is a sequence view that records when mechanics appear, how they combine, and what emotional or difficulty role each level plays. It prevents the team from balancing every board in isolation.
How should randomness be tested?
Run many seeds, inspect the distribution and extremes, validate openings, and replay cases near the failure boundary. Randomness should create tactical variety while preserving understandable agency and recovery.
Sources and further reading
- GDC: Level Design Saga
King level-design presentation covering theory, thought, tools, testing, and data-informed tuning.
- GDC: Candy Crush Postmortem
A first-party postmortem on design decisions and the role of luck.
- Microsoft: evergreen games
A 2026 panel recap with leaders from King and Mojang on maintaining and rebalancing long-running games.
Next step



