How We Test AI Girlfriend Apps

We test every app for 8+ hours minimum, score it on five criteria, and publish everything we found. Here's exactly how we do it.

Our process

  1. 1

    Sign up & test for 8h+

  2. 2

    Run our scoring rubric

  3. 3

    Cross-check with 2nd tester

  4. 4

    Publish & update over time

The 5 scoring criteria

Every app gets a score on each criterion from 0 (poor) to 5 (excellent). The overall score is a weighted average.

Character Diversity

Weight 20%

How many distinct personalities are available out of the box? Can users build custom characters with meaningful options (appearance, backstory, kink-level)? Are the differences cosmetic or do characters actually behave differently in chat?

10 / 10
Hundreds of pre-built characters AND a deep custom builder. Different characters genuinely behave differently.
6 / 10
A reasonable starter library, basic customization. Characters feel similar at some point.
2 / 10
Very few characters, or all characters feel like the same persona with a different skin.

Chat Experience

Weight 25%

Quality of the conversation. Does the AI stay in character? Does it remember earlier messages? How does it handle awkward turns, sensitive topics, repeated questions? We test 200+ messages minimum per app.

10 / 10
Natural flow, coherent memory across days, handles edge cases gracefully.
6 / 10
Mostly natural in short sessions, repetitive in long ones, occasional memory drops.
2 / 10
Generic responses, forgets context within 10 messages, breaks character often.

Image Generation

Weight 20%

Quality, consistency and variety of generated images. We test specific prompts (3 scenarios per character, 5 generations each) and grade against a consistency rubric: same face, same body, plausible anatomy, lighting variety.

10 / 10
Consistent identity across generations, varied poses and lighting, photorealistic or stylistically consistent.
6 / 10
Identity drifts on edge prompts, occasional anatomy issues, good in 70% of generations.
2 / 10
Identity changes every generation, frequent anatomy errors, censored heavily.

Video Generation

Weight 15%

When available: quality of short video clips. We test 3 prompts per character (idle, action, dialogue). We assess motion smoothness, face consistency, audio sync.

10 / 10
Smooth motion, identity preserved, lipsync works, multiple durations available.
6 / 10
Acceptable motion, occasional artifacts, identity mostly preserved.
2 / 10
Heavy artifacts, identity drifts mid-clip, or feature simply absent on paid plans.

Pricing Fairness

Weight 20%

What do you actually get for the money? We compare entry-tier pricing against feature gates, generation credits, and what's hidden behind upsells. We also flag opaque billing (credit systems, surprise charges).

10 / 10
Transparent pricing, generous free or low tier, no surprise charges.
6 / 10
Reasonable pricing but some features gated unfairly, opaque credit pack pricing.
2 / 10
Aggressive upsells, hidden credit consumption, non-discrete billing.

What we don't do

Just as important as our methodology — the things we refuse to do.

  • Accept payment for higher scores. Ever.
  • Cover apps we haven't personally tested for at least 8 hours.
  • Hide our affiliate relationships — see our full disclosure.
  • Republish content from press releases or third-party reviewers.
  • Generate fake user accounts to inflate testing volume.
  • Use AI to write our reviews. Every word is written by a human tester.

See the methodology in action

Read any of our 10+ reviews and you'll find the exact scores per criterion, the rationale, and the data we gathered.