Skip to content
CAI
Produce a survey ↗Verify a survey

stanfordnlp/dspy

This system is a Python-based framework for building and managing Large Language Model (LLM) applications, specifically focusing on programmatic control of model behavior. It provides a unified API for defining, executing, and evaluating LLM-based programs, including features for structured output formatting, tool calling, and multi-step reasoning. The codebase has been significantly refactored to remove legacy components, introducing a modern, flat public API and a robust testing infrastructure to ensure reliability across different model providers.

63.2

Adequate · 2 August 2026

27k

lines of production code

Python

primary language

9

bus factor · 473 authors in all

2

measurements over time

CAI band scale
CAI trend line

How it got here

2023 · API unification and legacy removal

This period focused on modernizing the project's development workflow and significantly refactoring the public API. Legacy modules, caching mechanisms, and internal primitives were removed in favor of a flattened, unified interface with improved adapter support.

10 changes

2024 · Comprehensive test coverage expansion

This period focused on establishing a robust testing infrastructure for the DSPy codebase, introducing extensive unit and integration tests across all major modules. The work covered core primitives, prediction modules, evaluation metrics, and client integrations, while also adding reliability testing frameworks and test utilities for external dependencies like LiteLLM.

12 changes

2025–2026 · Comprehensive test coverage expansion

This period focused on significantly expanding the test suite across the codebase, with a particular emphasis on new types, adapters, and core components. The work included adding comprehensive tests for language model types, retrievers, and proposers to ensure robustness and correct behavior.

5 changes

CAI lens gauges

Survey your own repository

stanfordnlp/dspy was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point the surveyor at a repository you know and see whether you agree with it.

Survey a repository

About this page

  • The description of this project is derived from its own commit history, not from its README.
  • The score is its highest published measurement, taken on 2 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit c69136b29a — the exact code this score is about.
  • Scored under rubric rubric-2026.08.18. Score the same commit under that rubric and you get the same number.
CAI link cards