Onboarding checklist: After your first login, you’ll see an Onboarding Checklist in the left sidebar with five steps: Connect your agent, Generate scenarios, Run a simulation, Run an evaluation, and Set up metrics.Each step is checked off automatically as you complete it. Once you’ve finished all five, the checklist disappears. Use it to track your progress as you work through this guide.
Prerequisites
- An Arklex account. Sign up here with your work email, or ask an admin for an invite link if your team already has an organization.
- An agent with an HTTP endpoint that follows the Chat Completions (
/chat/completions) schema or the A2A protocol. - Any auth headers your endpoint requires (e.g. an API key or bearer token).
Step 1: Connect an agent
Navigate to Agents in the left sidebar and click Connect Agent. Fill in the form:1
Name the agent
Enter a descriptive Agent name — e.g.
Banking Assistant.2
Choose the API type
Select Chat Completions for most agents, or A2A for the Agent-to-Agent protocol.
3
Add the endpoint URL
Paste the full URL Arklex will POST conversations to.
4
Add auth headers
Add any headers your endpoint requires (e.g.
Authorization: Bearer <token>). Values are encrypted at rest.5
Set the body
Add the default parameters or system messages your endpoint expects.
Step 2: Add scenarios
A scenario represents one type of simulated user — who they are, what they want, and what they know. Before generating scenarios, each one needs at least one knowledge base to draw on, so the simulated user has realistic information to work with during the conversation. Navigate to Scenarios and click Generate Scenarios (or Import if you already have a scenario file).1
Add a knowledge base
Add at least one knowledge source for the simulated user to draw on. Arklex supports two types:
Both types are chunked and indexed so the simulated user can reference their contents when forming conversation turns. For our example, upload the bank’s policy and FAQ documents so simulated customers ask grounded, realistic questions.
2
Describe your agent
Describe your agent and the kinds of users it serves — for our example, “a banking assistant for a retail bank, used by customers checking balances, making transfers, and resolving card issues.”
3
Review generated scenarios
Arklex generates a set of scenario personas, each with a user goal and profile — for instance, a customer disputing an unrecognized charge, or someone trying to send a transfer above their daily limit.
4
Save to a group
Edit any you want to adjust, then save them to a scenario group. Groups organize related scenarios together — e.g. all scenarios for transfers, card support, or fraud reports.
You’ll know it worked when your new scenario group appears with its personas listed underneath.
Step 3: Run a simulation
A simulation runs your scenarios against the connected agent and produces full conversation transcripts. Navigate to Simulations and click New Simulation.1
Select agent
Select the agent you connected in Step 1.
2
Name the simulation
Give it a descriptive name (e.g. “Banking Assistant — transfers & card support”).
3
Select scenarios
Choose the scenarios to include.
4
Configure conversation settings
Set Conversations per scenario (default 1, capped so the total stays at or below 50) and Max turns per conversation (default 5, max 10).
5
Run
Click Run Simulation.
Step 4: Run an evaluation and review results
An evaluation scores your transcripts with an LLM judge. Once the simulation shows Completed, navigate to Evaluations and click New Evaluation.1
Select simulation
Select the simulation you just ran.
2
Choose a judge
Pick an LLM Judge provider and model (e.g. GPT-4o, Claude 3.5 Sonnet, Gemini 2.0 Flash).
3
Select metrics
Choose the metrics to score. Goal Completion is always included. Add Helpfulness, Coherence, or any custom metrics you care about.
4
Run the evaluation
Click Run Evaluation. Arklex scores each conversation turn and updates the status to Completed when done.
5
Review the results
Click the completed evaluation to open its detail page, which breaks results into four sections:
- Quantitative Metrics — numeric scores per metric, grouped into turn-level and conversation-level, with band labels (Excellent, Good, Needs Improvement, Poor).
- Qualitative Metrics — label distributions for categorical metrics.
- Unique Errors — behavioral failures detected by the judge, grouped by severity with suggested fixes.
- Conversations — the full list of scored conversations with per-conversation scores and status.
Next steps
Your first evaluation is the loop you’ll repeat as your agent evolves. From here:- Invite your team so reviewers can add annotations and calibrate judge scores against human judgment.
- Add more agents to compare versions or A/B test prompts and models against the same scenarios.
- Customize metrics to score behaviors specific to your domain.
- Upload knowledge documents to give simulated users access to your product docs or FAQs.