Skip to content
CAI
Produce a survey ↗Verify a survey

apache/flink

This system is a distributed stream and batch processing engine that provides a comprehensive data processing framework. It features a modern DataStream V2 API and a robust Table/SQL API for defining complex data pipelines. The platform supports a wide array of connectors for file systems, databases, and external services, alongside state management, metrics, and testing utilities.

49.2

Weak · 6 August 2026

1.4M

lines of production code

Java

with Scala

28

bus factor · 2,036 authors in all

2

measurements over time

CAI band scale
CAI trend line

How it got here

2010–2018 · Hadoop compatibility and metrics expansion

This period focused on deepening Hadoop ecosystem integration by implementing a comprehensive compatibility layer for MapRed and MapReduce APIs, alongside adding support for Hadoop Writable types. Concurrently, the project significantly expanded its observability by introducing multiple new metrics reporters for systems like Prometheus, Datadog, and InfluxDB, while also refactoring the client submission architecture and modernizing the codebase with JUnit 5 and improved build tooling.

80 changes

2019–2020 · Table API and File Connector modernization

This period focused on modernizing the Table API with a new type inference framework and SQL parser overhaul, while simultaneously refactoring the File Source and Sink connectors to support new APIs and features like file compaction. The work also included significant improvements to the Kubernetes integration, Web Dashboard, and comprehensive test coverage for these new components.

63 changes

2021–2023 · RPC migration and architectural hardening

This period focused on migrating the RPC layer from Akka to Pekko and introducing a new ChangelogStateBackend for state management. The team also established stricter architectural constraints using ArchUnit and enhanced the SQL Gateway API, while simultaneously expanding test coverage across connectors and file systems.

61 changes

2024–2026 · DataStream V2 API and native integrations

This period focused on the introduction of the DataStream V2 API, providing a new execution environment and context model for stream processing. It also expanded Flink's ecosystem with a new ForSt state backend, native S3 filesystem support, and OpenTelemetry metrics integration, alongside significant test coverage improvements for the new APIs.

15 changes

CAI lens gauges

Survey your own repository

apache/flink was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point the surveyor at a repository you know and see whether you agree with it.

Survey a repository

About this page

  • The description of this project is derived from its own commit history, not from its README.
  • The score is its highest published measurement, taken on 6 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 7232ad9c2c — the exact code this score is about.
  • Scored under rubric rubric-2026.08.19. Score the same commit under that rubric and you get the same number.
CAI link cards