LionPact Hub
In development

Catch late and broken streaming data first.

LionPact Watch checks your streaming pipelines while data is moving. When a stage falls behind, a schema changes or volume drops, your team knows within minutes, with the stage and the cause in the alert.

Status
In active development
Works with
Apache Kafka, Apache Flink, Spark Structured Streaming, Apache Iceberg
Runs in
Your own cloud account or cluster
Access
Read-only. No code changes to your jobs.
Cost
Free and open source, always

The problem

In a streaming system, data is used seconds after it is produced. Dashboards refresh, fraud models score events and other teams read the same topics.

Most data quality checks still run on tables after data has landed, often once an hour or once a day. By the time a check fails, the bad or missing data has already been shown to users and fed into decisions.

The failures are usually quiet. Nothing crashes. A consumer slowly falls behind, a producer renames a field, or one partition stops receiving events. The pipeline looks green while the numbers drift.

  • Silent lagA job keeps running but falls minutes or hours behind
  • Schema driftA producer adds, renames or retypes a field
  • Dead partitionsSome partitions or devices stop sending events
  • Volume dropsEvent counts fall far below the normal pattern
  • DuplicatesRetries write the same event more than once
  • Bad valuesNulls, negatives or out-of-range values spike

What it helps your organisation achieve

LionPact Watch is built for data platform, analytics and ML teams that depend on real-time data.

  • Find problems in minutes

    Know about a late or broken stream before a business user or customer reports it.

  • Trust real-time numbers

    Show how fresh each dataset is, so people know when a dashboard can be relied on.

  • Protect downstream teams

    Schema contracts catch breaking changes at the source, before every consumer fails.

  • Know the impact straight away

    Lineage shows which tables, reports and models a failure affects, so the right people are told.

  • Spend less time debugging

    Each alert names the stage, the check that failed and when it started, so on-call engineers start in the right place.

  • Keep data inside your walls

    It runs in your environment with read-only access. Your data is never sent anywhere else.

How it works

Watch sits beside your pipelines. It reads signals your systems already produce and never sits in the data path, so it cannot slow down or break a job.

LionPact Watch architecture Your pipelines send read-only signals to Watch collectors. A checks engine evaluates freshness, contracts and anomalies, and sends results to your team as a dashboard, alerts, an impact view and a metrics export. Your pipelinesKafka topicsFlink jobsSpark Streaming jobsIceberg tablesSource databases (CDC) CollectorsConsumer offsets and lagSampled messagesJob and checkpoint metricsTable snapshots Checks engineFreshness and latencySchema contractsLearned baselinesDuplicate detectionLineage graph Your teamHealth dashboardSlack and email alertsImpact viewMetrics export LionPact Watch architecture Your pipelines send read-only signals to Watch collectors. A checks engine evaluates freshness, contracts and anomalies, and sends results to your team as a dashboard, alerts, an impact view and a metrics export. Your pipelinesKafka topicsFlink jobsSpark Streaming jobsIceberg tablesSource databases (CDC) CollectorsConsumer offsets and lagSampled messagesJob and checkpoint metricsTable snapshots Checks engineFreshness and latencySchema contractsLearned baselinesDuplicate detectionLineage graph Your teamHealth dashboardSlack and email alertsImpact viewMetrics export
All connections are read-only. Watch never writes to your systems.

How to use it

Watch is in development, so the steps and configuration below show the planned setup. Details may change before the first release.

  1. Run it next to your pipelines

    Deploy Watch as a container in your own cloud account or Kubernetes cluster. Nothing is installed inside your jobs, and no pipeline code changes.

  2. Connect your sources, read-only

    List the topics, jobs and tables you want watched. Watch reads offsets, metrics and a small sample of messages. It never writes to your systems.

    # watch.yaml (planned format)
    sources:
      - name: orders
        type: kafka
        bootstrap_servers: kafka-1:9092
        topic: orders
      - name: orders_clean
        type: iceberg
        table: analytics.orders_clean
  3. Say what healthy looks like

    For each stream, set how fresh it must be and which fields and types consumers rely on. Volume and value ranges are learned automatically from normal traffic.

    checks:
      - stream: orders
        freshness:
          warn_after: 60s
          fail_after: 5m
        contract:
          required: [order_id, customer_id, amount, created_at]
          types: { amount: decimal, created_at: timestamp }
        volume:
          baseline: learned   # learns the normal hourly pattern
  4. Send alerts to the right people

    Route warnings and failures to Slack channels or email by stream or team. Each alert says which stage, which check and since when.

    alerts:
      - channel: slack
        webhook_env: SLACK_WEBHOOK_URL
        streams: [orders, payments]
        on: [fail]
  5. Review and tune

    Use the dashboard to see the health of every stream. After the first few days, tighten thresholds where alerts are too loose and relax them where they are noisy.

Where organisations can use it

Any team that makes decisions on data that is seconds or minutes old.

Use caseTypical streamWhat Watch checks
E-commerceOrder and payment eventsOrders feed freshness, payment amount ranges, duplicate order IDs
Fraud and riskReal-time feature pipelinesLate features and null spikes before models score transactions
Supply chainSupplier, parts and shipment feedsMissing partner feeds, schema changes from partners, volume drops
IoT and telemetryDevice and sensor streamsDevices that stop reporting, out-of-range readings
Lakehouse CDCDatabase changes into IcebergReplication lag, schema drift from source tables
Real-time analyticsDashboards on streaming tablesFreshness shown next to each metric, so users know what to trust

Free and open source. Always.

LionPact Watch is a personal R&D project, not a product. There is nothing to buy, now or later.

  • No paid tier. Every feature is free for individuals and organisations.
  • No sign-up, no licence key. Use it without asking anyone.
  • Open code. Read it, change it and run it inside your company.
  • Your data stays with you. It runs in your environment. Nothing is sent to LionPact Hub.

Questions

Is it really free?

Yes. LionPact Watch is free and open source for individuals and organisations, and it will stay that way. There is no paid tier and no licence key.

When can my team start using it?

Watch is in active development. Connect with Rajesh Kaushik on LinkedIn to hear about the first release and early access.

Does it need access to our data?

It needs read-only access to offsets, job metrics and a small sample of messages to check contracts and values. It runs inside your environment, so your data never leaves it.

Will it slow down our pipelines?

No. Watch is not in the data path. It reads signals beside your jobs and makes no changes to them.

We already run data quality checks on our warehouse. Do we still need this?

Warehouse checks look at data after it lands. Watch looks at data while it is moving, so it catches problems earlier. The two work well together.

How can I suggest a feature or contribute?

Send a message on LinkedIn. Feedback from teams running streaming pipelines shapes what gets built first.

Questions about LionPact Watch?

For early access, feature ideas or a walkthrough of how Watch could fit your pipelines, connect with Rajesh Kaushik on LinkedIn.

Connect on LinkedIn