The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

Monitoring 101

Docker monitoring with Netdata

Docker Monitoring

What Is Docker?

Docker is an open platform that packages applications and their dependencies into lightweight, portable units called containers. A container bundles everything the software needs to run — code, runtime, system tools, libraries, and settings — so it behaves the same way on a developer’s laptop, in staging, and in production. Because containers share the host operating system’s kernel instead of shipping a full guest OS, they start in milliseconds and use far fewer resources than virtual machines.

Under the hood, containers are built from two Linux kernel features: namespaces, which isolate what a process can see (its own filesystem, network, and process tree), and control groups (cgroups), which limit and account for the resources a group of processes can use. Understanding cgroups matters for monitoring, because that is exactly where the most useful container metrics come from.

What Is the Docker Engine?

The Docker Engine is the core of the platform — the daemon (dockerd) that builds images, creates and runs containers, and manages their networks and volumes. It is the industry’s de facto container runtime and runs on the major Linux distributions (CentOS, Debian, Fedora, Oracle Linux, RHEL, SUSE, Ubuntu) as well as Windows Server. When people say “monitor Docker,” they are usually interested in two related layers: the containers themselves (their resource usage and health) and the Docker Engine daemon that orchestrates them.

Why Is Docker Monitoring Important?

Containers are lightweight and efficient, but without visibility they become a black box. Monitoring Docker lets you:

  • Detect performance bottlenecks early — catch CPU spikes, memory leaks, and disk I/O saturation before they reach users.
  • Understand behavior at scale — see how services behave under load or during autoscaling when tracking metrics per container.
  • Meet SLAs and uptime targets — get alerted on resource saturation, restart storms, and unexpected exits.
  • Right-size and control cost — use real usage data to avoid overprovisioning and prevent noisy-neighbor problems.
  • Troubleshoot faster — drill into real-time and historical metrics to isolate an issue in a single container without recreating the bug.

What Are the Benefits of Docker Monitoring Tools?

A good Docker monitoring tool gives you:

  • Real-time visibility into container state, health, and resource usage.
  • Insight into which container — and which process inside it — is consuming resources.
  • The ability to evaluate images and containers and track their current state.
  • Alerting that catches health-check failures, out-of-memory kills, and restart loops.

How Netdata Monitors Docker

Netdata collects Docker telemetry from three complementary sources, all with zero configuration and at per-second granularity:

  1. Per-container resource metrics via cgroups. Netdata reads the virtual files Linux exposes (usually under /sys/fs/cgroup/) to report each container’s CPU, memory, disk I/O, and network usage. It scans for new and removed cgroups every few seconds (configurable via check for new cgroups every in netdata.conf), so containers are picked up automatically as they start and cleaned up when they stop. This is the same cgroups mechanism Docker itself uses for resource isolation.

  2. Container state and health via the Docker collector. Netdata connects to the Docker instance over a TCP or UNIX socket and calls the Docker API — System info, List images, and List containers — to report how many containers are running, paused, or stopped, their health-check status, and image counts and sizes.

  3. Daemon-level metrics via the Docker Engine collector. For the dockerd daemon itself, Netdata uses the built-in Prometheus exporter to collect engine-wide metrics: container actions, build failures, health-check failures, and Docker Swarm manager state.

Because collection is per-second and auto-discovered, you see the exact moment a problem begins — not a 30- or 60-second average that hides it.

Key Docker Metrics to Monitor — and Why

Container state and health (Docker collector)

  • docker.containers_state — total number of containers by state (running, paused, stopped). Sudden shifts here often signal crash loops or a failed deploy.
  • docker.containers_health_status — health status across containers (healthy, unhealthy, starting). A rising unhealthy count is an early outage warning.
  • docker.images — count of active and dangling images. Dangling images are a common cause of creeping disk usage.
  • docker.images_size — total size of all images, useful for capacity planning.
  • docker.container_state — per-container state (running, paused, exited, and more).
  • docker.container_health_status — per-container health-check result.
  • docker.container_writeable_layer_size — the writable-layer size of each container, a key capacity-planning metric.
MetricDescription
docker.containers_stateNumber of containers in each state
docker.containers_health_statusHealth status across all containers
docker.imagesActive and dangling image counts
docker.images_sizeTotal size of Docker images
docker.container_stateState of an individual container
docker.container_health_statusHealth status of an individual container
docker.container_writeable_layer_sizeWritable-layer size per container

Per-container resource usage (cgroups)

For every container, Netdata tracks CPU usage and throttling, memory (used, cache, swap, and pressure), disk reads and writes, and network bandwidth, packets, and errors per interface — all per second, on both cgroups v1 and v2 with automatic detection. These are the metrics you reach for when an application slows down and you need to know which container is starving for CPU or leaking memory.

Docker Engine daemon metrics (Docker Engine collector)

  • Engine daemon container actions — the rate of container actions (create, delete, start, commit, change) the daemon performs. Useful for spotting churn and unexpected spikes in container activity.
  • Engine daemon container states — how many containers the daemon is managing in each state (running, paused, stopped).
  • Builder builds failed total — the rate of failed image builds, broken down by reason. A spike points to a broken base image, registry issue, or bad Dockerfile.
  • Engine daemon health checks failed total — the rate of failed container health checks the daemon runs.
  • Swarm manager leader — whether this node is the current Swarm manager leader. Flapping leadership indicates cluster instability.
  • Swarm manager object store — number of objects (nodes, services, tasks, networks, secrets, configs) held by the Swarm manager.
  • Swarm manager nodes per state — Swarm nodes by state (ready, down, unknown, disconnected).
  • Swarm manager tasks per state — Swarm tasks across their lifecycle states (running, failed, ready, rejected, shutdown, complete, and more).

Zero-Configuration Auto-Discovery

The best part of monitoring Docker with Netdata is that it is zero-configuration. If containers are already running when you install Netdata, it auto-detects them and starts collecting metrics immediately. Spin up new containers afterward — docker compose up, a docker run, a Swarm task, or a Kubernetes pod — and they appear in the dashboard within seconds, with the right charts and alerts already attached. Stop a container and it is cleaned up automatically. Netdata handles ephemeral, short-lived containers without complaint, which is exactly what dynamic environments need.

Monitor Many Containers — and the Apps Inside Them

Netdata auto-populates its menus with your containers by ID or name and scales to any number of them, whether you run 1, 100, or 1,000. It organizes CPU and memory charts into families so you can quickly see which containers use the most CPU, memory, disk I/O, or network, and correlate that with the rest of the system.

Monitoring doesn’t stop at the container boundary. Netdata can also monitor the applications running inside containers — a MySQL database, an Nginx web server, a PostgreSQL instance — using the same auto-discovery, so you get application-specific metrics and pre-configured alerts alongside the container’s resource usage.

Pre-Configured Alarms

Netdata ships with alarms for every running container out of the box. As soon as a container is detected, Netdata attaches CPU-utilization, RAM-usage, and RAM+swap-usage alarms for its cgroup, calculated against the limits you set — so they adapt automatically to any container size. You can edit health.d/cgroups.conf to change thresholds or add your own. On top of static thresholds, Netdata’s edge-based machine learning trains models per metric to flag genuine anomalies — restart storms, throttling, OOM kills, and health-check failures — without threshold guesswork.

Is Docker Monitoring Secure With Netdata?

Security matters when you instrument production containers. Netdata is designed with this in mind:

  • Read-only by default — Netdata collects container metrics without elevated privileges or modifying containers.
  • Minimal footprint — the agent is lightweight and exposes little attack surface.
  • Container isolation — metrics are collected externally via cgroups and system files, so monitoring never interferes with container behavior.
  • Data stays local — when connected to Netdata Cloud, metric data is encrypted in transit and stays on your infrastructure; only metadata leaves the node.

For stricter environments, Netdata can run with reduced permissions and inside its own container for additional isolation.

Advanced Docker Monitoring and Root-Cause Analysis

Beyond charts and alarms, effective Docker monitoring means being able to answer “why” quickly. Netdata’s Anomaly Advisor correlates thousands of metrics to surface the handful most likely responsible for an incident, and its browser-based troubleshooting tools show the processes and network connections inside a container — with history — so you can diagnose problems without docker exec or SSH. For hands-on, symptom-specific procedures, see the Docker operations guides: step-by-step runbooks for disk exhaustion, OOM cascades, daemon hangs, log explosions, container restart loops, and network isolation.

When you are ready to see it running against your own containers, the Live Demo needs no login, and you can learn how the full platform approaches this on the Docker monitoring solution page.

Troubleshooting Docker Monitoring in Netdata

If container metrics or alarms are not appearing as expected, work through these checks:

  • Verify cgroup access — make sure the Netdata agent can read from /sys/fs/cgroup/.
  • Check container timing — containers started before Netdata are detected on the next scan; if in doubt, restart Netdata to trigger immediate detection.
  • Confirm cgroup version — Netdata supports both cgroups v1 and v2, but a misconfigured kernel can interfere. Confirm your system’s cgroup version.
  • Review container limits — on hosts running many containers, consider adjusting Netdata’s resource limits or update frequency.
  • Inspect alarm configuration — check health.d/ to confirm alarms are enabled and thresholds are calibrated for your environment.

FAQs

What is Docker monitoring?

Docker monitoring is the practice of tracking the performance, resource usage, and health of Docker containers and the Docker Engine daemon, using tools that collect metrics such as CPU, memory, disk I/O, network, container state, and health-check status.

Why is Docker monitoring important?

It prevents downtime, keeps performance optimized, controls cost through right-sizing, and enables real-time troubleshooting when something goes wrong in a containerized application.

What does a Docker monitor do?

A Docker monitor tracks container states, health status, per-container resource utilization, and daemon-level metrics, and alerts you when key indicators cross a threshold or behave anomalously.

What is the difference between Docker and the Docker Engine?

Docker is the overall platform for building and running containers. The Docker Engine is the daemon (dockerd) at its core that actually builds images and runs containers. Netdata monitors both the containers and the engine.

Do I need to run Netdata inside every container?

No. One Netdata agent per host monitors all containers on that host. It reads cgroup metrics directly from the kernel — no sidecars and no per-container agents.

Why don’t I see my containers in Netdata?

Confirm the agent can read from /sys/fs/cgroup/ and that containers were running when Netdata started. If they were started afterward, wait for the next auto-discovery scan or restart Netdata to detect them immediately.

Does Netdata support Kubernetes and Docker Swarm?

Yes. Netdata monitors containers regardless of orchestrator. It reads per-container metrics via cgroups under Docker, Docker Compose, Docker Swarm, and Kubernetes, and adds Swarm-manager metrics via the Docker Engine collector.

How can I monitor Docker in real time?

Install Netdata on any host running the Docker daemon. It auto-discovers your containers and gives you per-second dashboards, alerts, and ML-based anomaly detection out of the box.