Features

Everything you need to test voice systems—and understand the answer.

Explore what SIPforge does, when to use each test engine, and how to read the result without digging through raw tooling.

How the workflow fits together

From setup to evidence in six clear steps.

01

Choose the right engine

Use SIP for infrastructure, Twilio Voice for carrier and Programmable Voice calls, or AI Voice Agent for real conversation testing.

02

Prepare what the test needs

Add servers, accounts, numbers, media, voice-agent prompts, and AI provider keys once, then reuse them safely.

03

Start with a clean baseline

Begin with low concurrency, short durations, and a clear cost cap. Increase load only after the first result looks correct.

04

Monitor the run live

Watch active calls, ASR, failures, spend, voice quality, AI latency, and infrastructure resources while the test is moving.

05

Inspect the evidence

Open call detail for SIP errors, provider status, DTMF, transcripts, callbacks, timing, and per-call cost.

06

Export or repeat

Share a PDF, export raw CSV or JSON, compare a change, or save the configuration as a reusable template.

Platform capabilities

Every feature supports a decision.

SIPforge keeps traffic generation, live analysis, troubleshooting, and reporting connected.

One testing workspace

Launch and review every voice test from the same history.

  • SIP / SIPp load tests
  • Twilio Voice call flows
  • AI voice-agent conversations

Live test visibility

See whether the system is holding up before the run finishes.

  • Call rate and concurrency
  • ASR and terminal outcomes
  • MOS, jitter, loss, and latency

Failure investigation

Move from a red number to the evidence needed to troubleshoot.

  • SIP response details
  • Provider and callback errors
  • Call logs and transcripts

Infrastructure context

Connect call behavior with what happened on the target server.

  • CPU and RAM
  • Open file descriptors
  • Asterisk channels and drops

Repeatable operations

Turn an ad-hoc test into a dependable quality process.

  • Reusable templates
  • Scheduled regression runs
  • Metric threshold alerts

Team-ready evidence

Give every stakeholder the right level of visibility.

  • Admin, runner, and viewer roles
  • Run-to-run comparison
  • PDF, CSV, and JSON exports

Choose the right engine

Practical guidance before you run.

Open a guide to see what to prepare, how to run the test, and what its results mean.

SIP / SIPp testingGenerate signaling and media traffic against Asterisk, FreeSWITCH, Kamailio, SBCs, proxies, and gateways.

Before you start

  • Create a server with host, port, transport, and credentials.
  • Confirm firewall, NAT, RTP range, and SIP transport reachability.

How to use it

  • Begin with OPTIONS or INVITE/BYE at a low call rate.
  • Set total calls, concurrency, hold time, transport, and codec.
  • Increase CPS only after a clean baseline.

How to read results

  • 200 responses show successful signaling.
  • SIP errors identify rejected or timed-out call legs.
  • MOS, jitter, loss, and RTT explain media quality.
Twilio Voice testingPlace real Programmable Voice calls and capture status, DTMF, cost, callbacks, and call-flow evidence.

Before you start

  • Connect a Twilio account and sync caller numbers.
  • Upload media when the flow plays audio and set a first-run cost cap.

How to use it

  • Choose a single number or a sequential or random number pool.
  • Add destinations and select playback, gather, or DTMF flow.
  • Watch callbacks to confirm Twilio reaches SIPforge.

How to read results

  • Completed means the call connected and finished.
  • Busy and no-answer affect ASR but are not platform failures.
  • Failed means Twilio reported a real call failure.
AI voice-agent testingRun real phone conversations, capture transcripts, measure responsiveness, and score whether the agent achieved its goal.

Before you start

  • Create a test persona, opening line, goal, and evaluation criteria.
  • Connect telephony plus the required LLM, speech-to-text, and text-to-speech providers.

How to use it

  • Start with one call to a known test number.
  • Listen for pacing, interruptions, and first-response delay.
  • Tune the prompt and providers before scaling.

How to read results

  • Transcript shows what the system heard and said.
  • Latency isolates delays in STT, LLM, TTS, or telephony.
  • End reason explains hangups, timeouts, stream loss, or provider errors.

Read the result

The metrics, in plain language.

ASR

Answer-Seizure Ratio

The percentage of call attempts that connected successfully.

MOS

Mean Opinion Score

A 1–5 estimate of call audio quality. Four or higher is excellent.

Jitter

Packet timing variation

Lower is better. High jitter often sounds choppy or robotic.

Packet loss

Audio packets not received

Under 1% is usually healthy; rising loss damages intelligibility.

Latency

Time before a response

Shows where a voice or AI experience begins to feel slow.

Failure

A real engine or provider error

Different from busy or no-answer outcomes, which mainly affect ASR.

Want to see these metrics in a finished test?

Explore the annotated example before creating your first workspace.

View sample report

From a readiness run

The conclusion, ready to repeat.

Anonymized examples of what a representative SIP and RTP report makes easy to say.

Sustained the planned call rate with voice quality still in the excellent range.
Contact-center operations
The failure spike sat next to CPU, so the capacity question had an answer.
Voice infrastructure
Sign-off used the PDF conclusion. Engineering kept the SIP detail underneath.
Release lead

Ready when you are

Know what to test—and what the answer means.

Create your workspace, choose the right engine, and begin with a clear baseline.