VR Playtesting from Alpha to Early Access and Beyond

Building a Licensed IP in VR: How ARVORE Playtested The Boys: Trigger Warning

17 months of playtesting. A power system rebuilt from scratch. A narrative tested as its own discipline — separately from everything else.

17+ moTesting on the platform
New + returning participantsFor longitudinal comparison
NarrativeTested as its own discipline

ARVORE

ARVORE is an award-winning XR studio based in São Paulo, Brazil — recipients of Emmy, Cannes XR, and many more awards. Known for narrative-driven immersive experiences including the Pixel Ripped series.

ARVORE partnered with Sony Pictures Virtual Reality to create the first official VR game set in the world of The Boys — the globally acclaimed Prime Video series — built alongside the show's creators and featuring voice performances by cast members from the series.

The Boys: Trigger Warning

A stealth-action VR game set inside the brutal, darkly comedic universe of The Boys. Lucas Costa's life detonates when the Armstrongs — a washed-up but lethal family of Vought Superheroes — turn a family outing into carnage. The Boys drag Lucas back from the brink, juice him with powers, and throw him headfirst into their underground war against Vought and its Supes.

Available on Meta Quest.

When ARVORE started testing The Boys on the platform, the game was a vertical slice — an early build representing only a small portion of the intended game, but real enough to test the central question of whether players could feel like a supe inside the world of The Boys.

Working on a globally recognized IP with a major studio publisher meant every creative and design decision carried real weight. ARVORE needed to validate mechanics, how players feel, and narrative comprehension — across a long development period, with full confidentiality required at every step, from the first participant through launch.

A globally recognized IP demands a specific kind of infrastructure

ARVORE's production required testing infrastructure that could match the complexity and confidentiality demands of the project at every stage — and grow more demanding as the game matured.

1
Highly specific participant profiles — and they got stricter over time

Every test round defined a precise combination of criteria: headset and hardware requirements, age range, location, language, specific VR experience level, genre preferences, games and apps engagement history, and more. As the game matured, requirements became more detailed and strict with each test round.

2
Full confidentiality, enforced at every touchpoint

Every participant signed an NDA before receiving any access to the game or related information. Access was gated behind that signature. With a major licensed IP, cast involvement, and Sony Pictures VR as publisher, this was a firm requirement.

3
New and returning participants — each answering a different question

Some test rounds required participants who had never seen the game to measure first impressions and first-time comprehension. Others required participants who had experienced specific earlier builds to directly evaluate what had changed. Some test rounds included both simultaneously. Some focused exclusively on returning participants to get a direct longitudinal comparison against what had been tested before.

4
Synced VR gameplay and third-person video — consistently, across all test rounds

Every test round required two synchronized video recordings captured simultaneously: VR gameplay footage and a third-person video of the participant. The third-person view reveals behavioral cues invisible in gameplay alone — how participants physically interact with the game, whether they play seated or standing, and how they respond in real space. Each submission included both synced recordings, complete and formatted consistently across all participants within the same test round. The total playtime across all sessions had to meet the required or estimated duration for that test round, since for longer sessions, participants split their play across multiple sessions.

5
A different test design at every stage

From short, focused sessions on specific levels to multi-hour full-game playthroughs — and from large groups to single in-depth sessions — each test round had its own session length, group size, task structure, and deliverables.

Seventeen months. Every stage of development tested.

ARVORE ran test rounds from an early vertical slice all the way through a dedicated narrative test round before launch. The phases below represent the types of test rounds they ran — each designed around what the team needed to learn at that specific moment in development.

1
Vertical Slice
Concept and core mechanics validation

Testing began with a large group of participants matched to a precise profile. Goals were foundational: could players understand and execute the core mechanics? Did the controls make sense? Were the interactions comprehensible from the start? All participants brought substantial VR experience, and profile requirements included genre background, games and apps engagement history, and more.

2
Mid-Development · Mechanics Iteration
Did the changes work? New and returning participants answer differently.

As the game's mechanics evolved significantly, separate test rounds ran for new participants and returning participants. New participants measured whether updated systems were immediately comprehensible. Returning participants — those who had experienced specific earlier builds — directly evaluated whether changes had landed and what had improved. The perspective that only someone who played the earlier version can give.

3
Mid-Development · Level-Specific Test Rounds
Focused feedback on specific parts of the game

When ARVORE needed feedback on specific sections of the game in isolation, test rounds were structured accordingly — participants were directed to specific levels and given defined tasks to complete. These test rounds ran with automatic participant acceptance once the required profile and availability were matched, without the need for manual review.

4
Full Beta
Full-game sessions with mixed participant groups

With a substantially more complete game, ARVORE ran extended sessions across a mix of new and returning participants. Participants completed the full tutorial, watched all cutscenes, and progressed as far as possible. Survey coverage expanded significantly — covering onboarding pacing, boss comprehension, combat feel, the injectable power system, navigation, and presence alongside all core mechanics. All deliverables arrived in the dashboard when each session completed.

5
Narrative Test Round
Story comprehension and emotional clarity — tested on its own

A dedicated narrative test round ran with new participants — focused entirely on whether the story came through clearly: who is the protagonist, what is the world, what are the stakes, and how does the opening sequence make players feel. No combat metrics. All qualitative. This test round gave the team a targeted signal on narrative comprehension, independent of mechanics performance.

Fully configured by the studio. Fully autonomous in execution.

Every test round was configured by the ARVORE team directly on the platform, building on previous configurations rather than starting from scratch. Surveys changed substantially as the game evolved: early test rounds asked whether mechanics were comprehensible at all; later test rounds asked detailed questions about specific boss fights, difficulty curves, and emotional responses to narrative moments. Every test round included synchronized video recordings alongside survey responses.

All participants signed an NDA before receiving any access. The studio could monitor progress and completion status in real time — who had signed, who hadn't, who had completed their session, and whose post-play deliverables — survey responses and video recordings — were ready to review. ARVORE configured each test round to accept matched participants automatically or hold applications for manual review, depending on what that test round required. Whenever needed, the platform automatically matched a replacement participant and kept the test round on track.

After access was granted, participants completed sessions independently and remotely and results were delivered to the dashboard without the team needing to be present.

What the platform handled

  • Participant matching across a detailed and evolving profile — headset specs, VR experience level, genre preferences, games and apps engagement history, age, region, language, and more — with criteria tightening as each test round grew more specific.
  • NDA management and signature tracking before any access was granted, with real-time visibility for the studio into who had signed, who hadn't, and who had completed each stage.
  • Participant pools for new and returning participants — with ARVORE able to selectively invite returning participants for test rounds where longitudinal comparison was the goal, with participant privacy and compliance maintained throughout.
  • VR gameplay footage synced with a third-person video of the participant, per session per participant — with each submission consistent across all participants within the same test round, and total playtime verified against the required or estimated session duration.
  • Custom survey builder allowing the ARVORE team to create surveys from scratch for each test round, directly on the platform, fully customizable.
  • Automatic participant replacement when needed, so test rounds stayed on track without manual intervention from the team.
  • Multi-member team access — to configure test rounds, monitor progress, and review results, so testing was a shared practice across the studio.
  • All results — survey responses, video recordings, participant data — consolidated in one dashboard with automatically generated reports.

Before launch, ARVORE ran a test round with no mechanics questions at all.

All questions focused on narrative: the protagonist's identity and motivations, the relationships between characters, the emotional response to key dramatic scenes, and whether the world and its stakes were clear to someone who had only experienced the opening

New participants only — people playing the game for the first time, including people who hadn't watched the show. The goal was a clear read on whether the story worked, separate from how the mechanics were feeling at that stage.

Across participants, the protagonist's identity, motivation, and emotional journey were described clearly and consistently. The intended emotional response to the first major dramatic scene — shock, dark humor, fear — was what participants reported. Participants who hadn't seen the show understood the world and the stakes.

One finding emerged: some participants described not feeling the urgency to escape — the story made clear that danger was imminent, but what players could physically do in that moment didn't yet match the tension the narrative was building. Exactly the kind of finding that only surfaces when story is tested on its own.

The game that shipped looked almost nothing like what participants encountered at the start.

The Boys: Trigger Warning launched on Meta Quest.

The clearest measure of what the testing period produced is the distance between the first test round and the last. The mechanics that participants encountered at the start of testing looked almost nothing like what shipped — the movement system had significantly evolved, the power set had grown, and participants were describing the experience in terms that matched The Boys universe.

That shift was tracked test round by test round, confirmed by returning participants who had played earlier builds and could speak directly to what had changed.

Does your project have this level of complexity and stakes?

If any of the following describes your situation, the platform is built for it.

1
You're building a VR game with high creative and commercial stakes — licensed IP, narrative-driven, or both

Where every build is sensitive, all participants need to be vetted, and the game has to deliver on expectations players already bring to the IP.

2
Your target audience is highly specific — and who plays your game during development directly determines the quality of your data

Headset and hardware specs, VR experience level, genre preferences, games and apps engagement history, and more. Requirements that can be configured precisely for every test round and tightened as the production demands it.

3
You need to know whether your changes actually worked — not just whether the new build feels different

Returning participants who played earlier builds give you a direct answer. Having infrastructure to manage them separately from new participants is what makes that comparison meaningful.

4
You want to stop spending time on the overhead of running tests

Screening participants, sending NDAs, chasing signatures, coordinating and checking video recordings, participant replacement — all managed through the platform, so your team focuses on what the data reveals.

How Mystic Interactive ran 26 test rounds — from alpha to live ops — tested with new and returning players, shipped Crystal Conquest on schedule, and user testing became part of their development process.

About Mystic Interactive and Crystal Conquest

Mystic Interactive is a small independent studio building Crystal Conquest — a free-to-play VR strategy game where players harness the elements in PvP multiplayer combat. The game supports single-player and multiplayer. 

The game runs on both Meta and Steam, supporting standalone Meta Quest as well as PCVR headsets, and features multiple game modes (1v1 Duel mode, 3v3 MOBA mode, and a PvE Gauntlet mode), a soft-currency economy, and a new character pipeline that was actively expanding throughout testing.

Where they started

When Mystic started testing on the platform, Crystal Conquest was in Alpha — not yet public, with no player community to draw from. This is the hardest moment to get fresh-eye feedback, and also the most critical.

They needed to test everything continuously — experience across devices, build distribution flow, new characters, economy systems, multiplayer networking and user experience, and onboarding throughout their entire development cycle as a strategy to de-risk their Early Access release.

The Challenge

Coordinating external VR playtesting is harder than it looks.

Mystic needed participants with particular VR headsets and hardware specs, matched to a precise profile, organized into synchronized multiplayer groups — often scheduled and run without the team needing to be present at all. Sessions ran within a defined context, with participants completing tasks, so results were actionable at every stage. That's not something you can coordinate through a Discord community, even if you had one.

1 Multiplayer coordination at scale

Testing multiplayer, especially 3v3 combat, requires the right number of participants online simultaneously — each matching a specific profile, headset, and hardware requirement. The team needed this to happen reliably, without managing it themselves.

2 Precise participant profiles that varied test by test

Requirements varied — sometimes any Meta Quest headset or PCVR, sometimes all Meta models as standalone, sometimes PC specs strict enough to guarantee the game would run. One multiplayer playability test required participants to be physically located on the US East Coast. The platform matched for all of it.

3 Cross-platform complexity

Crystal Conquest runs on Meta Quest headsets and PCVR headsets — distributed via Meta, SideQuest, and Steam — each with a different install and distribution flow. Testing across all of them, sometimes simultaneously, required infrastructure — not just coordination.

4 No public players yet — by design

Studios at this stage can't test with their own community because they don't have one yet. But this is precisely when first-time player perspective matters most: before habits form, before players learn to work around rough edges.

5 New vs. returning players see the game differently

Testing both groups separately — often on the same day with separately matched cohorts — surfaces onboarding gaps that experienced players miss entirely.


“Opening testing to remote players from anywhere in the world allowed us to get a wide range of perspectives on the project, with feedback coming from people with varying levels of vision.”
- Jazmin Cano, Accessibility Product Manager

Observations and feedback received from the playtesters with blindness and low vision helped Owlchemy to prioritize iterations, validate hypotheses, and uncover new VR accessibility features to integrate into Cosmonious

Image credits: Mystic Interactive and Crystal Conquest