[·] indexedbyagents.com
Live
← the full index

Agent Observability & Evals

52 records · last verification sweep 2026-09-08

This app lets you browse interactive leaderboards that compare performance across different categories and methods. Simply pick a category, methodology, and metric to see updated tables and charts....

Verified2026-09-04
Statuslive
RecordIA-0191
IA-0193

Arga Labs

101

Build, seed, and call twins of the services your software acts on via API, CLI, or MCP.

Verified2026-09-04
Statuslive
RecordIA-0193
IA-0194

Arize AI

101

AI observability platform with open-source Phoenix.

Domainarize.com
Verified2026-09-04
Statuslive
RecordIA-0194
101

CLIRank scores 420 APIs on agent-friendliness: auth method, CLI tools, pricing transparency, headless operation. Search by capability, compare alterna

Verified2026-09-04
Statuslive
RecordIA-0196
IA-0197

Comet Opik

101

Comet's open-source LLM evaluation and tracing platform.

Domaincomet.com
Verified2026-09-04
Statuslive
RecordIA-0197
IA-0200

Datanika

101

Build, orchestrate, and monitor data pipelines with dlt and dbt. Multi-tenant, multi-language, open-source.

Verified2026-09-04
Statuslive
RecordIA-0200
IA-0207

Instrukt

101

Integrated AI environment in the terminal. Build, test and instruct agents. - blob42/Instrukt

Verified2026-09-04
Statuslive
RecordIA-0207
IA-0209

Kusho

101

KushoAI is AI-native infrastructure for software maintenance. Autonomous agents that live in your CI/CD and handle testing, healing, and monitoring co

Domainkusho.ai
Verified2026-09-04
Statuslive
RecordIA-0209
IA-0211

LangSmith

101

LangChain's platform for debugging, testing, and monitoring agents.

Verified2026-09-04
Statuslive
RecordIA-0211
101

This web app lets you browse the Agent Memory Leaderboard, where textual and coding memory systems are ranked side‑by‑side. Choose academic or commercial entries, view overall scores, detailed capa...

Verified2026-09-04
Statuslive
RecordIA-0213
101

This app shows interactive leaderboards that rank AI models on the Gaia2 benchmark, displaying scores like pass@1, search, execution, adaptability, and more. Users simply open the page and can refr...

Verified2026-09-04
Statuslive
RecordIA-0214
IA-0215

Lemma

101

Surface silent failures, pull context across traces, improve your agent before users churn.

Verified2026-09-04
Statuslive
RecordIA-0215
IA-0217

Lunary

101

Open-source observability and evaluation for LLM apps.

Domainlunary.ai
Verified2026-09-04
Statuslive
RecordIA-0217
IA-0219

MOVEdot

101

MOVEdot puts AI agents on the manual load of vehicle development, so testing, simulation, and validation keep pace. Runs in your own cloud.

Verified2026-09-04
Statuslive
RecordIA-0219
IA-0221

NOFire AI

101

NOFire connects your stack, builds a live production graph, and gates every agent action at runtime. Ship fast without leaving production exposed.

Domainnofire.ai
Verified2026-09-04
Statuslive
RecordIA-0221

Browse and filter through detailed leaderboard results of Open Agent performance across various math and multi-modal benchmarks. Select evaluation dimensions, algorithms, datasets, and models to cu...

Verified2026-09-04
Statuslive
RecordIA-0222
101

Orbit Sentinel monitors FCC, ITU, UNOOSA & FAA filings, updated daily. Track spectrum applications, launch permits & regulatory changes across agencie

Verified2026-09-04
Statuslive
RecordIA-0223
IA-0225

Pibit.ai

101

The fastest path from submission to better decisions.

Domainpibit.ai
Verified2026-09-04
Statuslive
RecordIA-0225
101

Matched from 47,000+ marketing agencies in 30 seconds. Ranked on verified review data - never on payment. Free, no commission, no paid placements.

Verified2026-09-04
Statuslive
RecordIA-0226
IA-0228

Portkey

101

AI gateway and observability platform for agents.

Verified2026-09-04
Statuslive
RecordIA-0228
IA-0230

Ragas

101

Open-source evaluation for RAG pipelines and agents.

Domainragas.io
Verified2026-09-04
Statuslive
RecordIA-0230
IA-0231

Raindrop

101

Monitor your AI Agent the right way. Get alerted when your agent fails in production, trace exactly what went wrong, and prove your fix worked.

Verified2026-09-04
Statuslive
RecordIA-0231
IA-0232

Rectify

101

All-in-one operations platform for SaaS. Session replay, monitoring, support, code scanning, roadmap, and changelogs in one visual platform.

Verified2026-09-04
Statuslive
RecordIA-0232

Find your research gap, match your journal, and get an 8-agent AI review in 15 minutes. 25 free credits to start — pay-as-you-go, no subscription.

Verified2026-09-04
Statuslive
RecordIA-0233
IA-0234

Sentry

101

Error monitoring with AI agent observability.

Domainsentry.io
Verified2026-09-04
Statuslive
RecordIA-0234

This web app shows the current ranking of students in the Agents Course Unit 4 Challenge, displaying each user’s name, score, and submission time. The list refreshes automatically every minute, and...

Verified2026-09-04
Statuslive
RecordIA-0235
101

The @testdriverai in any GitHub repo and TestDriver writes UI tests and catches regressions before they merge. AI-powered end-to-end testing.

Verified2026-09-04
Statuslive
RecordIA-0236
IA-0239

Vellum

101

Platform for building, evaluating, and deploying AI agents.

Domainvellum.ai
Verified2026-09-04
Statuslive
RecordIA-0239
IA-0240

wcagc

101

Scan SaaS and e-commerce pages for automated WCAG 2.1 AA issues, mapped to ADA, Section 508, or EN 301 549. Then plan manual review and fixes.

Domainwcagc.com
Verified2026-09-04
Statuslive
RecordIA-0240
IA-0241

Website

101

Cognee is the open-source agent memory platform for LLM agents. Build persistent memory across sessions with graph, vector, and relational retrieval that runs self-hosted, in Docker, on-prem, or on Co

Domaincognee.ai
Verified2026-09-04
Statuslive
RecordIA-0241