← Back to writing

Observe Before You Enforce: Why Segmentation Starts with a Flow Map

·–

Part 1 of Microsegmentation from First Principles, a series that works through the design of Innerwall, an open-source microsegmentation platform, one problem at a time.

Microsegmentation projects rarely fail because the firewall can't express the policy. They fail because the policy was wrong on the day it was switched on. A connection nobody knew about gets denied, something important stops working, and the rollback that follows usually takes the project's credibility with it.

This post argues that the root cause is ordering. Enforcement is attempted before observation. It sets out what an allowlist actually claims, why observation is the only practical way to support that claim, and what a collector has to get right for the evidence to be trustworthy. The rest of the series builds on it.

A segmentation policy is a claim about completeness

Microsegmentation means default deny between workloads, with an explicit allowlist of the connections that are permitted. Write the allowlist as a set of flows A, and the set of flows the application actually requires as R. Two things can go wrong:

  • Outage: some required flow is not allowed. Formally, R ∖ A ≠ ∅.
  • Exposure: some allowed flow is not required. Formally, A ∖ R ≠ ∅.

The goal is R ⊆ A with A ∖ R as small as practical. The two failure modes are not symmetric, and the asymmetry drives everything that follows. An outage is immediate and visible, and it gets blamed on the security change that caused it. Exposure is silent: an over-broad rule costs nothing until the day someone abuses it. So an allowlist has to be complete before it is enforced, while it only has to be tight eventually.

That puts the whole difficulty on knowing R. The usual sources are architecture documents, firewall request tickets, a configuration database, and interviews with application owners. All of them describe the flows someone intended, written down by people, at some point in the past. Call that estimate R_doc. The error that matters, R ∖ R_doc, is exactly the set of flows that will cause an outage, and by construction nobody knows what is in it. If anyone knew, it would be in the documents.

Observation is evidence, not authorization

The alternative is to measure. Let O be the set of flows observed on the network during some window. O is not R, and the way it differs is the most important picture in this series.

Two overlapping sets inside the space of all possible flows: R, the flows the application requires, and O, the flows observed during the window.All possible flowsRrequiredOobservedR ∖ OR ∩ OO ∖ R
R ∖ O · Needed, not seen
Paths rarer than the observation window: month-end jobs, failover, certificate renewal, break-glass access. A policy built only from the map denies them.
R ∩ O · Needed and seen
The evidence a policy can be built from. Candidates for rules.
O ∖ R · Seen, not needed
Stale dependencies, misconfiguration, scanners, an attacker. Observed traffic is not authorized traffic; copying the map into policy makes it permanent.

Observation does two useful things. It replaces an unknown error (R ∖ R_doc) with a smaller and better-characterized one (R ∖ O): the paths that are needed but too rare to have run during the window. And it surfaces O ∖ R for review, instead of leaving it hidden.

It does not do two other things. First, it does not make observed traffic legitimate. A dependency on a decommissioned service, a misconfigured client, a scanner, or an attacker moving laterally all show up in O. Turning a flow map directly into policy, rule for rule, encodes today's exposure as tomorrow's intent. The map is the input to a policy decision, not the decision.

Second, it cannot see R ∖ O. No amount of care in the collector reveals a path that did not run. Part of that gap can be closed by observing for longer, and part by enumerating rare paths explicitly. The remainder is what the next stage, simulation, is for. Part 2 covers it.

What a flow is, at this layer

Before evidence can be trusted, it has to be defined. On Linux, the kernel's connection tracker already holds state for every connection once tracking is active on the host, which any stateful firewall rule requires: the original tuple (as sent by the initiator), the reply tuple (as expected back), protocol state, and optionally packet and byte counters. A collector does not need to capture packets to learn about connections. It can subscribe to the tracker's events over netlink. Here is what a short PostgreSQL connection looks like from conntrack -E, with accounting enabled:

    [NEW] tcp      6 120 SYN_SENT src=10.0.1.23 dst=10.0.2.9 sport=48712 dport=5432 [UNREPLIED] src=10.0.2.9 dst=10.0.1.23 sport=5432 dport=48712
[DESTROY] tcp      6 src=10.0.1.23 dst=10.0.2.9 sport=48712 dport=5432 packets=14 bytes=2196 src=10.0.2.9 dst=10.0.1.23 sport=5432 dport=48712 packets=11 bytes=3842 [ASSURED]

Three properties of this source matter.

Direction comes from the original tuple, not from port heuristics. The first src= is the initiator. "The lower port is the server" is a guess that fails on high-port services and on clients that bind low ports. The tracker records who sent the first packet, so there is nothing to guess.

Every connection is seen, not a sample. Sampled flow export (sFlow, or IPFIX with packet sampling) is cheap because it inspects one packet in N. The probability that a connection of k packets is missed entirely is (1 − 1/N)^k. At a 1-in-1,000 sampling rate, a 14-packet connection like the one above is missed with probability 0.999^14 ≈ 0.986. Short, infrequent connections are almost invisible to sampling, and they are precisely the population that lands in R ∖ O. Sampling inflates the most dangerous region of the diagram. For segmentation, unsampled connection-level observation is the requirement, not a luxury.

Events report transitions, not state. A NEW event counts a connection; a DESTROY event carries its final byte counters (and only when net.netfilter.nf_conntrack_acct=1; otherwise the counters are zero). A connection established before the collector subscribed has no NEW event, so it surfaces only when it closes. A connection pool that was opened before the agent started and never recycles is invisible to an event stream alone. A collector should dump the existing table once at startup to close that gap. Separately, anything exempted from tracking in the raw table with notrack never reaches the tracker at all.

UDP and ICMP have no handshake, so the tracker synthesizes pseudo-connections bounded by timeouts (tens of seconds to a few minutes for UDP, per the nf_conntrack_udp_timeout* sysctls). That is the right granularity for a flow map: a client querying a DNS resolver from fresh source ports produces many tracker entries but one dependency.

Who reports a connection

If every host reported every connection it saw, a connection between two managed workloads would appear twice, once from each end. Deduplicating the two reports is solvable, but there is a simpler rule with better properties. Each host reports only connections whose original destination is one of its own addresses. Everything else the tracker holds (outbound, forwarded, loopback) is ignored at the source.

Two managed workloads

web-1

10.0.1.23

original dst not local · ignores

db-1

10.0.2.9

original dst local · reports

src 10.0.1.23 → db-1 :5432/tcp

Reported once, by the host that would enforce inbound policy on it.

Through a source-NAT proxy

client

10.8.0.5

proxy

10.0.3.4

rewrites source

db-1

10.0.2.9

reports

src 10.0.3.4 → db-1 :5432/tcp

db-1 sees the proxy, so the map shows the proxy as the peer. That is also the address db-1's firewall will match.

The rule has three consequences.

  1. Each connection between managed workloads is reported exactly once, by its destination. There is nothing to deduplicate.
  2. The reporter is the enforcement point. With inbound policy, the host that would allow or deny a connection is its destination. So the source address in the record is exactly the address that host's firewall will match. That matters most when the network rewrites addresses. Behind a source-NAT load balancer or proxy, the destination sees the proxy's address, and the map shows the proxy as the peer. This is not an inaccuracy to correct for. A rule written at that host can only ever match the proxy's address, so the map is showing the operator the truth they have to write policy against. The real clients behind the proxy are segmented at the proxy's own inbound boundary.
  3. Outbound dependencies to unmanaged destinations are out of the inbound map. A first release that enforces inbound only can live with that. Part 6 covers why inbound comes first.

The set of "this host's addresses" is itself dynamic (interfaces come and go, addresses are added), so the collector re-reads it periodically rather than once at startup.

The ephemeral port is noise

Of the four values in a TCP or UDP tuple, the source port carries no policy meaning. It is chosen by the client's kernel per connection (on Linux, from net.ipv4.ip_local_port_range, which defaults to 32768–60999), and no rule will ever be written against it. Keeping it makes storage scale with the number of connections. Dropping it makes storage scale with the number of distinct relationships.

So the collector aggregates on (source address, destination address, destination port, protocol) over a fixed window, keeping a connection count, a byte count, and the first and last instants seen. Consider 200 application instances each opening 50 connections a minute to one database port. That is 10,000 connection records a minute, or 200 aggregated records per one-minute window. The second number is bounded by the shape of the dependency graph, not by traffic volume, which is what makes it safe to ship continuously. One instance's record for one window, rendered as JSON (on the wire it is protobuf):

{
  "src_address": "10.0.1.23",
  "dst_address": "10.0.2.9",
  "dst_port": 5432,
  "protocol": "PROTOCOL_TCP",
  "direction": "DIRECTION_INBOUND",
  "decision": "POLICY_DECISION_OBSERVED",
  "connection_count": 50,
  "byte_count": 301900,
  "first_seen": "2026-09-14T15:02:00.412Z",
  "last_seen": "2026-09-14T15:02:59.871Z"
}

In Innerwall the window defaults to sixty seconds and is set by the control plane. Per-connection records never leave the host.

Resolve addresses to identities once, at ingest

The tracker reports addresses. Policy is written in terms of workload identity and labels, which is the subject of Part 3. Somewhere, 10.0.1.23 has to become "web-1, app=checkout, env=prod". The question is when.

The tempting answer is at query time: store addresses, then join them against the inventory when someone opens the map. That is wrong, because addresses get reassigned (DHCP leases, autoscaling, instance replacement) and labels change. Suppose 10.0.4.17 belonged to a billing workload until Tuesday and was then reassigned to a scratch batch host. A query-time join attributes all of billing's historical connections to the batch host, and shows them under the batch host's current labels. The map would be confidently wrong about the past, and a policy written from it would be wrong in ways nobody would think to check.

So resolution happens once, at ingest, against the inventory as it stands at that instant, and the result is stored with the record:

  1. If the address is a managed workload's current address, the peer is that workload, and a snapshot of its labels is stored alongside.
  2. Otherwise, if it falls inside a defined address group, the peer is the most specific (longest-prefix) group containing it.
  3. Otherwise, the peer is the address itself: an unmanaged peer.

Unmanaged peers are first-class, not errors. They are often the clients just outside the segmentation boundary (a monitoring system, a jump host, a partner network), and they are exactly the sources the policy will need explicit rules for.

The stored record is now immutable history: it names the peer and the labels that were true when the flow happened. Two different questions should not be conflated. "What scope was this flow in when it happened?" is answered by the snapshot. "What have the workloads currently in this scope seen?" is answered by resolving the scope now and reading their records. Both are legitimate, and a system that answers one when asked the other produces policy reviews that quietly disagree with reality. The identity of the reporting workload comes from its authenticated connection, never from the payload, so a compromised agent cannot report flows on another workload's behalf.

The whole pipeline, and where it can lose data

On each host

  1. 1

    Connection tracker

    The kernel already keeps per-connection state. The collector subscribes to NEW and DESTROY events over netlink.

    lossEvent socket overflow (ENOBUFS); flows marked NOTRACK; connections that predate the subscription

  2. 2

    Inbound filter

    Keep a connection only if its original destination is one of this host's addresses. Drop loopback, forwarded, and outbound.

  3. 3

    Window aggregation

    Fold events into one record per (src, dst, dst port, protocol) per window: connection count, bytes, first and last seen. The source port is discarded.

  4. 4

    Bounded buffer

    Closed windows wait here until the control plane is reachable.

    lossOverflow drops the oldest windows; every dropped record is counted and reported

one window per stream, over its own mTLS connection

Control plane

  1. 5

    Ingest-time resolution

    Resolve each source address once, against the registry at that instant: workload, else most specific address group, else the bare address. Snapshot the workload's labels.

  2. 6

    Windows and totals

    Append the window's rows for time-range questions (pruned after a horizon). Upsert one row per key for “since first seen” (kept).

An incomplete map must say so

The diagram marks the places where observations can be lost. They need attention because of how a flow map is used: absence on the map is read as absence of a dependency, and absence of a dependency becomes a deny. A map that silently drops data is worse than no map, because it produces confident policy with outages built in.

Loss can happen in three places:

  • Between the kernel and the collector. The netlink event socket has a finite receive buffer. If events arrive faster than the collector consumes them, the kernel drops them and the next receive returns ENOBUFS. Those events cannot be recovered. The collector can only detect the overflow and treat the affected interval as incomplete.
  • Between the collector and the control plane. Agents must keep working when the control plane is unreachable, so closed windows wait in a buffer. Any buffer that is bounded will eventually overflow. When it does, the oldest windows should go first (recent data matters more for an imminent policy decision), and every dropped record must be counted. In Innerwall, that count rides on the agent's heartbeat, so the operator sees that a workload's map is incomplete before trusting it.
  • Structural blind spots. Untracked flows, connections that predate the collector and never close, and zero byte counts where accounting is disabled. These are properties of the host, so the agent should report them once, plainly, rather than let them pass as "no traffic".

The rule that follows: a workload whose evidence is known to be incomplete should not be the basis for enforcement until a clean observation period has passed. The console should make that state impossible to miss.

How long is long enough?

Request-driven dependencies show up within hours: anything exercised on every business day saturates almost immediately. What remains are periodic processes, and each one appears only when its period comes round.

Distinct dependencies discovered over time
Illustrative model, not measured data. Each step is a periodic process running for the first time inside the window. A quarterly process would not appear at all.

There is a clean way to reason about this. Take a process that runs once per period T, at a phase you don't know, and an observation window of length W. It is observed with probability min(W / T, 1). A two-week window catches a monthly process about 46% of the time (14 / 30.4) and a quarterly one about 15% of the time (14 / 91.3). Observing for longer helps, but only up to the longest period you are prepared to wait for. Anything rarer than that is in R ∖ O by construction.

So "observe for N days" is the wrong stopping rule. A better one has three parts:

  1. The rate of new dependencies has decayed. The curve is flat apart from steps you can attribute to known periodic processes.
  2. The window spans the longest business cycle you can afford to wait for. For most applications that means at least one month-end.
  3. Rarer paths are enumerated, not awaited. Disaster-recovery failover, quarterly reporting, annual certificate renewal, and break-glass access are either exercised deliberately during observation or written into policy from their specifications and marked as unobserved, so the review knows those rules rest on documentation rather than evidence.

Storage has to support this. Innerwall keeps two shapes of every flow. Per-window rows answer time-range questions ("what did this workload see last Tuesday?") and are pruned after a horizon, thirty days by default. A single cumulative row per (workload, peer, port, protocol) answers "has this ever been seen, and when first and last?" and is never pruned. A quarterly job observed once, two months ago, has long since aged out of the windows, but its cumulative row still says it exists. Without that second shape, the retention horizon would quietly become the observation window, and the stopping rule above would be impossible to apply.

From map to policy

The flow map is where policy authoring starts, not where it ends. Two transformations stand between them.

Subtract. Review O ∖ R. Every observed dependency is either confirmed as intended or treated as a finding. The map's real value as a security tool is often here, before a single rule is enforced: it shows connections that should not exist.

Generalize. The map records specific pairs: this address talked to that workload on that port. Policy should be written against identity: workloads labeled app=checkout may reach workloads labeled app=payments-db on 5432/tcp. That way it holds when instances are replaced or scaled. Generalizing from pairs to labels is a judgment call about intent, and Part 3 covers how to make it safely.

What comes out is a candidate policy, and it is still a hypothesis about R. The next step tests it against live traffic without enforcing it: compile the policy to the same host firewall rules it will eventually run as, end each chain in accept instead of drop, and log every connection that would have been denied. That turns R ∖ O, the gap observation cannot close, into a log line rather than an outage. That is Part 2.

Summary

  • A segmentation policy is an allowlist under default deny, so it is a claim that every required flow is listed. Outages come from the gap between that claim and reality, and documentation cannot measure the gap.
  • Observation replaces an unknown gap with a characterized one (R ∖ O) and exposes traffic that should not exist (O ∖ R). Observed is not authorized.
  • Observe connections, not packet samples. Sampling preferentially misses the rare, short connections that cause outages.
  • Have each host report only the connections it receives. That reports each connection once, from the enforcement point, with the source address that point will actually match.
  • Aggregate away the ephemeral source port, so data volume tracks the dependency graph rather than the traffic.
  • Resolve addresses to workloads and labels at ingest and store the snapshot, because addresses and labels change and history must not.
  • Count every lost observation and surface it. On a flow map, absence becomes a deny.
  • Stop observing on evidence (a decayed discovery rate, a full business cycle, enumerated rare paths), not on a fixed number of days, and keep a never-pruned "since first seen" record so retention does not cap the evidence.

Next in the series: Simulation as a Deployment Stage: Running Policy Without Dropping a Packet.