Aviz Logo
Contact
Modern network stack built for the AI era.
Build AI Factories
Start here: Neocloud, sovereign AI, and private AI cloud users
Network Automation with ONESToken Economics with AI ParserAgentic Operation with AI-NOCPOC Validation with FTAS
Modernize your Network
Start here: Cisco, Arista, Juniper network users; Gigamon and Netscout observability users
Autonomous Agentic AI PlatformNetwork CopilotAviz Packet BrokerAviz Service NodeVirtual Aviz Service Node / vTAP
SONiC & AI Experience Hub
Validate open networking and AI-driven infrastructure with real-world, multi-vendor testing environments.
Partner with Aviz Networks
Join our ecosystem of channel and technology partners. Together we deliver open networking solutions that drive innovation and growth.
Make Networks for AI. Introduce AI in your Networks.
End-to-end solutions, any NOS, any switch, any ASIC, any LLM, any application , backed by partner best practices, proven tech, and SLAs.
Explore why Aviz is the best partner to modernize your network with.
Explore Case Studies, TCO and ROI Calculators, Certifications, Community and News room
Aviz Training and Certification
Learn and Certify in SONiC and AI
24/7 World-Class SONiC Support & Proven Services.
Our dedicated team delivers round-the-clock, world-class SONiC support with unmatched quality, scalability, and efficiency, keeping your network optimized, secure, and always running at its best.
Hamburger
Aviz Logo

Make every AI request measurable.

AI-Parser — AI Traffic Observability
AI-Parser turns AI inference traffic into per-request, per-tenant intelligence: tokens, latency, throughput, identity, prompt signals, and completion status — without instrumenting the application.
Skip the browsing. Ask AI.
Hello! How may I help you today?
The Problem

GPUs can be healthy while AI applications perform poorly.

You're billed for tokens and blamed for latency, but the traffic between the user and the model is still a black box. Traditional observability tells you traffic is flowing — AI-Parser tells you what that traffic means.

01

Know who is consuming AI

Separate traffic by user IP, tenant, application, or stable API-key identity, not only by model.

02

Measure real token usage

Capture input, output, and total token counts carried in the AI response, then roll them up by tenant.

03

See the experience users receive

Track time to first token, end-to-end response time, inter-token latency, and tokens per second.

04

Find waste and contention

Identify aborted streams, truncated answers, over-allocation, noisy neighbors, and struggling workloads.

AI-Parser observes the request and response as delivered on the wire, creating an independent transaction record for AI usage and experience.

How It Works

See the AI call, not just the GPU.

AI-Parser reconstructs OpenAI-compatible HTTP and JSON traffic, including streamed responses, to expose the information needed for cost, performance, governance, and operations.

End users and AI agents talk to an AI workload node that runs the inference engine (vLLM, Dynamo, TGI, Triton, TensorRT-LLM, SGLang) alongside Aviz AI-Parser, which provides per-app token visibility, latency and TTFT, prompt and usage signals, and chargeback-ready billing. The node also runs the training model on GPUs, and AI-Parser feeds an observability stack of Elastic with Aviz Elastic Node, Splunk, New Relic, Prometheus, Datadog, Grafana, and Aviz AI-NOC.

AI-Parser runs as a lightweight service alongside the inference engine, on the same node — no mirrored tap, no separate DPU or x86 appliance required.

Turn shared GPU infrastructure into accountable services.

When the orchestration platform maps each tenant or application to an IP and GPU allocation, AI-Parser records align naturally with that model.

Token usage and cost by tenant for chargeback or showback
TTFT, end-to-end latency, inter-token latency, and throughput by tenant
Workload profiles by model, prompt size, request rate, and call type
Noisy-neighbor, fairness, and under-utilization insight
Aborted streams and wasted token-budget signals by tenant

From raw AI transactions to live operational views

Per-request dashboard view showing token counts and latency

Per-request dashboard

Review user identity, model, prompt, token counts, and delivered latency.

Usage trends view showing prompt activity over time and a frequent-prompts word cloud

Usage trends

Track prompt activity, frequent requests, and top users over time.

Transaction metadata view showing a structured field-value record

Transaction metadata

Identity, tokens, prompt signals, performance, and completion status.

Features

What AI-Parser adds beyond model-server metrics

Engine metrics stay the best source for GPU internals. AI-Parser adds identity, content, and delivered experience.

AreaModel-server viewAI-Parser viewWhy it matters
IdentityAggregate, labelled by modelClient IP, tenant, stable API-key identityChargeback, SLA, and accountability per tenant
Prompt contentNever exposedParsed from the body — fingerprint or textGovernance, DLP, prompt-cache analysis
TokensSelf-reported by the engineCounted independently on the wireAn audit trail against mis-billing and drift
End-to-end latencyRequest entry to completionFirst request byte to last response byteThe latency the user actually experienced
Time to first tokenAdmit to first tokenRequest in to first byte outCatches host stack, API server, and queue delay
Aborts and wastePartial failure countersTCP FIN/RST with no completion markerSees the client hang up, and attributes it

Vendor-neutral: vLLM, TGI, Triton, TensorRT-LLM, Dynamo, and SGLang all speak the same OpenAI-compatible schema, so AI-Parser reads them the same way.

Use Cases

A different use case for every team, from one read.

The same passive read of AI traffic answers a different question depending on who's asking.

Performance

Performance as delivered

Time to first token, wire end-to-end time, inter-token latency, throughput, and stream completion.

Governance

Identity and privacy controls

Per-IP and per-tenant records, stable identity fingerprints, prompt hashing by default, and optional raw text.

Agents

Agentic-flow visibility

Observe the multiple backend model calls created by one user question and regroup related calls through identity.

Accuracy

Honest, structured records

Read or compute fields from observed traffic. When a value is not present, leave it blank instead of guessing.

Visibility

End-to-end visibility

Understand AI traffic from the application and agent through the network to the LLM and GPU.

Consumption

AI consumption

Measure prompt, completion, and total token usage by application, workload, or model.

Efficiency

GPU efficiency

Spot expensive GPUs waiting on traffic, requests, or application dependencies.

Telemetry

Open telemetry

Export enriched inference KPIs via Kafka into your existing stack.

How to Buy

Talk to us about licensing.

Pricing details coming soon

AI-Parser licensing terms are being finalized. Talk to an Aviz architect for current pricing and deployment options for your environment.

Technical questions.

Does AI-Parser require instrumenting my application?

No. It observes traffic as delivered on the wire — there's nothing to instrument in the application or inference engine.

Where does it deploy?

As a lightweight service on the GPU node, alongside the inference engine — no mirrored tap, no separate DPU or x86 appliance, and nothing inline in the request path.

Which inference engines does it support?

vLLM, TGI, Triton, TensorRT-LLM, Dynamo, and SGLang — all read the same way, since they share an OpenAI-compatible schema.

Does it store my users' prompts?

Prompts are converted to one-way fingerprints by default. Raw prompt text is retained only when explicitly enabled.

What happens to fields AI-Parser can't observe?

They're left blank — never guessed or fabricated.

How does this relate to my model server's own metrics?

They're complementary, not competing. Model-server metrics remain the best source for GPU internals; AI-Parser adds the identity, content, and delivered-experience context that's only visible on the wire.

Bring token-level visibility to your AI infrastructure.

Connect AI traffic, tenant allocation, network context, and GPU service performance in one operational view.

Explore Capabilities