Ember’s rule is simple: a review should show how it reached a conclusion. A polished screenshot, an operator promise or one charming reply may be interesting evidence, but none deserves to become a universal verdict by itself.

EmberGF reviews AI companion products as software with emotional design, not as contestants in a fantasy pageant. We examine what a service documents, what a user can reproduce, which questions remain unanswered and whether commercial relationships are visible. A five-star score is not a lab coat.

Name the evidence level first

Every material statement should fit one of four labels:

  • Documented: stated in a current primary source such as terms, privacy notices, help pages or a live checkout screen.
  • Observed: visible in a dated account or interface check, with the device, locale and account state recorded.
  • Reproduced: tested more than once with a written procedure and comparable conditions.
  • Unknown: not established by the available evidence.

An operator claim can be reported without being endorsed. A desk review must say that no account test occurred. A hands-on review must describe what was actually done; “we tested it” is not a method.

A six-part companion review

  1. Scope and identity. Record the operator, official domain, intended audience, source date and exact product being reviewed. Similar names are not proof that two services share an owner or policy.
  2. Product boundaries. Separate documented chat, memory, media and voice features from adjectives such as “realistic,” “private” or “unlimited.” Those words need definitions and test conditions.
  3. Privacy path. Inspect data categories, model-development language, human review, retention, export and deletion. Our privacy checklist covers the questions to ask before using an intimate prompt.
  4. Billing path. Capture the charge, currency, billing period, renewal, included usage and paid extras shown at the final decision point. Cancellation, refund, chat deletion and account deletion are four different checks.
  5. Behavior check. When testing is authorised, use fictional low-risk prompts, a defined sequence and more than one session. Record instruction following, contradiction, unwanted escalation and whether account controls behave as described.
  6. Exit and update. Locate support, cancellation and deletion before recommending a paid commitment. Preserve the source date so future policy changes can be reported rather than silently erased.

Measure behavior without grading affection

We do not score whether a bot feels like the ideal partner. That judgment is personal and easily shaped by novelty. We can examine whether a stated boundary persists, whether a memory claim survives a repeatable fictional test, and whether an interface distinguishes generated output from a human message.

The NIST AI Risk Management Framework organises voluntary risk-management work around Govern, Map, Measure and Manage. Its Generative AI Profile adds suggested actions for risks specific to generative systems. Neither document certifies consumer companion apps, but both support a useful principle: evaluation needs context, documented limits and continuing monitoring.

The research proposal known as Model Cards was designed for reporting model uses, evaluation conditions and limitations. EmberGF applies the same documentation instinct to reviews: state what evidence covers and where it stops.

Commercial independence has a visible location

A referral relationship must never create a test result. When a relevant approved commercial link exists, the disclosure belongs beside the recommendation—not hidden in a distant policy page. The FTC Endorsement Guides FAQ explains why material connections should be clear and conspicuous.

If no approved destination exists, the review remains non-commercial. We do not substitute an unrelated offer simply because a button-shaped space looks lonely.

How readers can audit us

Check whether the article names its sources, date, test limits and commercial relationship. Compare the live checkout with any billing summary. Treat an unsupported score, precise latency number or broad privacy guarantee as a question, not a fact.

Use Synthetic Chemistry for matching claims, Filters and Consent for safety design, and The Real Cost of an AI Companion for the spending checklist.

A trustworthy review does not eliminate uncertainty. It labels uncertainty clearly enough that a reader can make a smaller, safer and reversible decision.

What we know / what remains uncertain

Known: This article distinguishes observed product behavior, published policy, and expert analysis.

Uncertain: AI products change quickly. Material updates are reflected in the modified date above.

Sources & method

EmberGF links primary sources where available, separates testing from opinion, and labels commercial relationships beside the relevant recommendation.