FlowGraph

An open-source, zero-infrastructure developer observability suite for local development. Combines a lightweight OpenTelemetry Node.js SDK, an embedded CLI relay, and an interactive React DAG visualizer that dynamically maps architecture topology and streams live distributed execution traces in real time.

TypeScriptNode.jsOpenTelemetryReact 19React FlowWebSocketsDagreZustandVitetsup

Architecture

Building FlowGraph: Zero-Infrastructure OpenTelemetry & Live Architecture Visualization for Local Dev

When building backend services and distributed microservices, understanding runtime execution flow is surprisingly painful during local development.

Developers are usually caught between two extremes:

  1. Littering their codebase with messy console.log statements that pollute terminal stdout and lack causal context or timing info.
  2. Spinning up heavy enterprise observability stacks—Docker Compose clusters with Jaeger, Zipkin, Tempo, Prometheus, and Grafana—just to inspect a few local HTTP requests and database queries.

Because setting up local APM tools is heavy and tedious, most engineers delay instrumenting OpenTelemetry until right before shipping to production. That is when they discover broken context propagation, missing child spans, or unexpected latency bottlenecks.

I built FlowGraph to solve this. FlowGraph is an open-source, zero-infrastructure developer observability suite designed specifically for local development. It consists of a lightweight Node.js SDK (@pundir07/flowgraph-sdk), a self-contained CLI relay (@pundir07/flowgraph), and an interactive React DAG execution dashboard that visualizes system architecture topology and streams live distributed traces in real time.


The Vision: Zero Setup, 100% Production Portability

The goal for FlowGraph was simple:

  • Zero Configuration: A developer should be able to run npx @pundir07/flowgraph and start seeing live traces immediately on a clean web dashboard with zero Docker containers or external daemon setups.
  • Drop-in Tracing: Instrumenting a Node.js backend should take fewer than 3 lines of code.
  • 100% OpenTelemetry Compliant: Developers shouldn't learn a proprietary tracing API. FlowGraph is a transparent proxy layer over standard @opentelemetry/api. When ready to ship to production, the exact same code routes to Datadog, Honeycomb, Tempo, or AWS X-Ray without changing a single line of instrumentation.
  • Dynamic Topology Learning: Instead of static architecture diagrams that go stale, FlowGraph dynamically infers service relationships, database queries, and downstream API calls from incoming execution spans.

Architecture Overview

FlowGraph is organized as an npm monorepo with clean separation between trace collection, relay transport, and browser visualization:

1. The SDK Layer (@pundir07/flowgraph-sdk)

  • Wraps @opentelemetry/sdk-trace-node and @opentelemetry/auto-instrumentations-node to auto-capture incoming HTTP requests, outbound network calls, database queries (PostgreSQL, MongoDB, Redis), and framework routing.
  • Exposes a 100% compliant OpenTelemetry proxy interface (startActiveSpan, setAttribute, addEvent, recordException, setStatus).
  • Implements a custom OpenTelemetry SpanProcessor that serializes finished spans and dispatches them asynchronously over a WebSocket connection to the local FlowGraph CLI relay.
  • Distributed as dual ESM/CJS bundles using tsup for full compatibility with modern Node.js and legacy CommonJS codebases.

2. The CLI Relay & Server (@pundir07/flowgraph)

  • Built with commander as a standalone CLI executable.
  • Runs a lightweight Node.js HTTP server that simultaneously serves the pre-bundled React visualizer SPA and handles duplex WebSocket upgrades on a single unified port.
  • Acts as a real-time multiplexer: accepts incoming span streams from multiple backend producer services and broadcasts parsed trace events and topology diffs to connected browser clients with sub-millisecond latency.
  • Features fail-fast port collision detection (EADDRINUSE) with helpful troubleshooting output and clean signal handlers (SIGINT, SIGTERM).

3. The DAG Execution Visualizer (Web Dashboard)

  • Built with React 19, @xyflow/react (React Flow), Dagre (hierarchical graph layout engine), Zustand, and Tailwind CSS.
  • Automatically calculates rank-based DAG layouts from parent-child span trees and service dependencies without visual jitter as new spans arrive.
  • Provides interactive span inspection, execution waterfalls, error stack highlights, duration metrics, and metadata payloads.

Key Technical Challenges & Deep Dives

1. OpenTelemetry Auto-Instrumentation & Module Loading Order

One of the trickiest parts of Node.js observability is auto-instrumentation. Libraries like @opentelemetry/auto-instrumentations-node work by monkey-patching core modules (http, https) and third-party drivers (express, pg, ioredis) when they are required.

If application code imports or requires express before the OpenTelemetry SDK initializes its patches, auto-instrumentation silently fails to intercept incoming requests.

To make FlowGraph bulletproof, I architected the SDK initialization to register instrumentations synchronously before any other module imports execute, while isolating the WebSocket transport to run asynchronously. This ensures that:

  • Core hooks are registered at the earliest phase of the event loop.
  • The developer only needs to call initFlowGraph({ serviceName: "api-gateway" }) at their application's entry point.

```typescript import { initFlowGraph, tracer } from "@pundir07/flowgraph-sdk";

// Initialize FlowGraph before importing route handlers initFlowGraph({ serviceName: "auth-service", relayUrl: "ws://localhost:4100", });

// Use standard OpenTelemetry API anywhere in your code export async function handleLogin(req, res) { return tracer.startActiveSpan("auth.verify_credentials", async (span) => { try { span.setAttribute("user.id", req.body.userId); const result = await verifyUser(req.body); span.setStatus({ code: 1 }); // OK return res.json(result); } catch (err) { span.recordException(err); span.setStatus({ code: 2, message: err.message }); // ERROR throw err; } finally { span.end(); } }); } ```

2. SDK Non-Blocking Resilience & Bounded Ring Buffering

A fundamental rule of observability tooling is: the telemetry tool must never degrade or crash the host application.

If the local FlowGraph CLI dashboard is closed or restarting while the backend is serving traffic, the SDK must not crash, throw unhandled promise rejections, or leak memory buffering unsent spans.

I implemented several safety layers in @pundir07/flowgraph-sdk:

  • Bounded In-Memory Ring Buffer: Spans are queued in a circular ring buffer with a strict size ceiling. If the relay is offline and the buffer fills up, oldest spans are discarded gracefully rather than causing an OOM (out-of-memory) condition.
  • Exponential Reconnection Backoff with Jitter: When the WebSocket connection drops, the SDK attempts automatic reconnects with jittered backoff to prevent socket thundering herds.
  • Fail-Safe Try-Catch Isolation: All serialization and dispatch routines are wrapped in defensive boundaries so telemetry bugs are isolated from the application request-response lifecycle.

3. Single-Port Multiplexing for CLI and Browser

Traditional local APM stacks often require configuring separate ports for the ingestion collector (e.g. OTLP gRPC 4317, HTTP 4318) and the web UI (e.g. Jaeger UI 16686).

To deliver a true zero-config developer experience, FlowGraph runs everything on a single port (default 4100):

  • Standard HTTP GET requests serve the pre-compiled static React visualizer assets.
  • HTTP Upgrade headers (Upgrade: websocket) are intercepted by the server to handle WebSocket handshakes.
  • The WebSocket server protocol distinguishes between producers (backend Node services publishing trace spans) and consumers (browser visualizer tabs subscribing to live events), routing payloads efficiently across channel pools.

4. Real-Time DAG Layout with Dagre and React Flow

Visualizing distributed traces is straightforward as a timeline waterfall, but understanding the macroscopic architecture topology requires rendering a Directed Acyclic Graph (DAG) of services, endpoints, and data stores.

When multiple services are streaming hundreds of spans per second, computing graph layouts can quickly cause severe UI lag and distracting element jumping:

  • Node & Edge Extraction: Incoming spans are parsed to extract unique service nodes (api-gateway, users-db, payment-service) and directed invocation edges with latency weights.
  • Incremental Dagre Auto-Layout: Using Dagre's rank-based hierarchical algorithm, the visualizer calculates optimal coordinate positions (x, y) for nodes and connection curves.
  • Zustand State Isolation: Graph topology state and active trace selection are decoupled in Zustand stores to prevent full-tree React re-renders when high-frequency span events arrive.

Production Portability: Zero Vendor Lock-In

Because FlowGraph is built directly on standard OpenTelemetry APIs and data models:

  1. Developers use the exact same @opentelemetry/api conventions they would in enterprise production environments.
  2. In local development, @pundir07/flowgraph-sdk points to the local CLI relay for instant visual feedback.
  3. In staging and production environments, teams can swap or augment the exporter to standard OTLP exporters (@opentelemetry/exporter-trace-otlp-http or @opentelemetry/exporter-trace-otlp-grpc) to stream directly to Datadog, Grafana Tempo, Honeycomb, or New Relic without refactoring business logic.

Summary & What's Next

FlowGraph bridges the gap between raw console.log debugging and heavyweight production APMs, giving developers an instant, zero-setup visual dashboard for OpenTelemetry traces right on their local machines.

The project is completely open source under MIT and available on npm:

  • CLI package: @pundir07/flowgraph
  • Node.js SDK: @pundir07/flowgraph-sdk