Track prompt versions, test conversations with simulated voice, text, or human testers, and score each run against your own criteria. See how your prompt changes affect the results.
Yes. Refunds post within 30 days.
Keep a history of your prompt versions and test how each one behaves in conversation. Compare results to see where a change helps and where it introduces new problems.
Set evaluation criteria for your use case, from following instructions to completing a task or answering accurately. Score each conversation against those criteria to see where your prompt needs work.
Run quick text checks, simulate spoken conversations, and test with people. Use each mode to investigate different issues and refine your prompt before putting it in front of users.
Bring prompt versions, test conversations, and custom scores together so your team can find issues and decide what to improve.
Select a supported voice model, add your system prompt, and adjust the voice settings for the behavior you want to test.
Describe the conversation you want to test and set criteria for a successful run, such as instruction following, task completion, or answer accuracy.
Run a text conversation, simulate a voice session, or test with a human participant. Choose the mode that helps you investigate the issue.
Review scores and flagged issues, save a new prompt version, and test again. Compare results across versions using the same scenarios and criteria.
See how people across accents, ages, and languages interact with your agent. Collect ratings and feedback on what felt clear, confusing, or frustrating.
Run quick text checks or test spoken conversations with simulated callers. Create personas with their own voice, pace, background noise, and temperament, then run them in parallel.
Save each prompt and voice configuration with a note. Review earlier versions, compare their scores, and restore a previous configuration.
Define criteria for task completion, instruction following, and answer accuracy. Add policies or reference documents when checking factual responses.
Dual-channel audio with synced transcripts. Share a clip of the exact failing turn.
Turn-by-turn response time, barge-in recovery and dead air, measured on the audio itself.
Run a suite on every prompt or model change. Block the merge when pass rate drops.
Review run scores, inspect flagged issues, and decide what to test next. Example data shown for illustration.
Choose from supported voice models and speech providers, or connect an existing agent. Compare how your prompt behaves across different model and voice configurations.
Prices shown billed annually. Monthly billing is 20% more.
For one agent getting ready to launch
For teams shipping agent changes weekly
For regulated and high-volume contact centers
A vetted, paid pool of testers screened for clear audio and attention to detail. You choose languages, accents and demographics per run, and every tester follows your scenario brief.
Yes. If your agent answers a phone number, SIP trunk or WebRTC room, Cullman can call it. Platform integrations add richer logs but are optional.
Each run is evaluated against the criteria you define, such as instruction following, task completion, and answer accuracy. For accuracy checks, add policies, FAQs, or a knowledge base to review responses against your source material.
Most runs of up to 50 calls finish within a few hours during business hours. Simulated runs finish in minutes.
Yes. Trigger a suite from GitHub Actions, GitLab or any CI with one API call, and fail the build on a pass-rate threshold.
Recordings are encrypted at rest and kept for 90 days by default. Enterprise plans can set retention and data region.
Bring a prompt to a live demo. We'll walk through a test conversation, score it against your criteria, and show you how to compare results across prompt versions.
Request a demo