9.0 KiB
Architecture Overview
LogWisp moves log entries through independent pipelines. Everything else — plugins, sessions, TLS, statistics — hangs off that spine.
Component Hierarchy
main
└── Service
├── Pipeline "app"
│ ├── Registry instance tracking, single-instance enforcement
│ ├── Session Manager per-pipeline connection/session bookkeeping
│ ├── Sources[] plugin instances, keyed by id
│ ├── Flow
│ │ ├── Rate Limiter optional, token bucket
│ │ ├── Filter Chain optional, ordered
│ │ ├── Formatter raw | txt | json, with sanitizer
│ │ └── Heartbeat optional generator
│ └── Sinks[] plugin instances, keyed by id
├── Pipeline "audit"
│ └── ...
└── Status Reporter optional, 30s interval
Package map:
| Package | Responsibility |
|---|---|
cmd/logwisp |
Entry point, help, logger bootstrap, signal loop, status reporter |
internal/config |
Typed config schema, loading, top-level validation |
internal/service |
Owns the pipeline set; start, stop, shutdown, global stats |
internal/pipeline |
Pipeline runtime and per-pipeline plugin registry |
internal/flow |
Rate limiter, filter chain invocation, formatting, heartbeat |
internal/filter |
Regex filter and filter chain |
internal/format |
Adapter over lixenwraith/log formatter + sanitizer |
internal/source/* |
Source plugins |
internal/sink/* |
Sink plugins |
internal/plugin |
Global factory registry populated by plugin init() |
internal/chain |
Chain wire protocol: hello preamble, entry codec, backoff |
internal/tlsx |
The single seam between TLSOptions and crypto/tls |
internal/session |
Session manager and per-instance proxy |
internal/core |
Shared types (LogEntry, TransportEvent), capabilities, constants |
internal/tokenbucket |
Rate limiter primitive |
internal/sanitize |
Standalone hex-escaping helpers |
Plugin Registration
Every plugin registers itself in an init() function, and
cmd/logwisp/bootstrap.go blank-imports each package to trigger those
init()s. Adding a plugin therefore means writing the package, calling
plugin.RegisterSource / plugin.RegisterSink, and adding one blank import.
Registration may attach metadata. The console source declares
MaxInstances: 1, because a process has only one stdin; the per-pipeline
registry rejects a second instance of any such type.
Data Flow
Entry lifecycle
- Source produces a
core.LogEntryand publishes it to every subscriber channel it has handed out. Publication is non-blocking: a full subscriber channel increments the source'sdropped_entriescounter. - Flow applies, in order: rate limit → filter chain → formatter. A drop at
any stage ends the entry's life and increments
flow.total_dropped. - The formatter output becomes a
core.TransportEvent, which carries both the formattedPayloadand the original structuredEntry. - Dispatch sends the event to every sink's input channel with a non-blocking send.
LogEntry fields:
| Field | Purpose |
|---|---|
Time |
Entry timestamp |
Node |
Origin node label for chained topologies; stamped at the first hop, preserved by relays |
Source |
Origin identifier within the node (filename, plugin id, …) |
Level |
DEBUG/INFO/WARN/ERROR/TRACE, when detected |
Message |
Log content |
Fields |
Optional structured metadata as raw JSON |
RawSize |
Original byte size, used by the entry-size cap |
Carrying Entry alongside Payload is what makes chain sinks
format-independent: a tcp_chain or http_chain sink re-serializes the
structured entry rather than shipping whatever text the local formatter chose.
Back-pressure and drops
There is exactly one drop policy and it is not configurable: never block.
| Stage | Full-buffer behaviour | Counter |
|---|---|---|
| Source → subscriber | Drop the entry | source dropped_entries |
| Flow | Drop on rate limit, filter, or format error | flow.total_dropped |
| Pipeline → sink | Drop for that sink only | pipeline total_dropped_by_sink |
| TCP/HTTP sink → client queue | Drop for that client only | sink dropped_writes |
The tcp_chain sink is the one deliberate exception. It holds a line across
reconnects until it is written or the process shuts down, so a downstream
outage propagates backwards as a full input buffer and surfaces as
total_dropped_by_sink on the pipeline rather than as silent data loss inside
the sink. The http_chain sink retries a batch with backoff, and drops it only
on a non-retryable response or on shutdown (dropped_batches).
Concurrency Model
- One goroutine per source drains that source's subscription and feeds the flow.
- The flow's formatter holds a mutex; the underlying formatter reuses an internal buffer and is not goroutine-safe.
- Each network sink runs one broadcast/broker goroutine plus, per connection, a writer goroutine (and for TCP, a reader goroutine that exists only to detect disconnects and refresh session activity).
- Chain sinks run a single run-loop goroutine that exclusively owns the connection or the pending batch, so no locking is needed around either.
- Statistics are atomics; configuration and registries use RW mutexes; shutdown is context cancellation plus wait groups.
Shutdown ordering
Pipeline.Stop is deliberately ordered so in-flight data drains:
- Stop all sources concurrently; each closes its subscriber channels.
- Wait for the run loop, which ends when every subscription channel closes.
- Stop all sinks concurrently.
Network Architecture
All listeners bind tcp4 and all dialers dial tcp4. IPv6 clients cannot
connect; this is deliberate, not an oversight.
| Plugin | Role | Protocol |
|---|---|---|
tcp sink |
Listener | Raw broadcast of formatted payloads |
http sink |
Listener | HTTP/1.1 SSE; HTTP/2 negotiated via ALPN when TLS is on |
tcp_chain source |
Listener | Chain protocol, persistent NDJSON stream |
http_chain source |
Listener | Chain protocol, NDJSON batches over POST |
tcp_chain sink |
Dialer | Chain protocol, persistent stream, auto-reconnect |
http_chain sink |
Dialer | Chain protocol, batched POST with retry |
TLS is built in exactly one place, internal/tlsx, which exposes
Server(opts) for listeners and Client(opts, host) for dialers. See
Security.
Sessions
Each pipeline owns a session.Manager. Plugins receive a session.Proxy
scoped to their instance id, so one plugin cannot see or remove another's
sessions. A session records the remote address, creation and last-activity
timestamps, and metadata — including tls and tls_peer_cn for TLS peers, and
auth_method / auth_identity for authorized ones.
Idle sessions are reaped every 5 minutes against a 30-minute idle limit. The HTTP sink's broker treats a vanished session as an eviction signal and closes the corresponding SSE client.
Authorization decisions do not read session metadata — they are made from the
handshake by internal/authz, at the point of connection or request, and their
outcome is recorded in the session. That ordering matters: a session exists
only for a peer that was already admitted. See the
mTLS authentication design.
Configuration Reload
Reload (signal or file watch) rebuilds the entire service:
- Re-read the config through the config manager.
- Build a new service from it. If construction fails, the old service keeps running untouched.
- Shut the old service down, start the new one, and restart the status reporter if it is enabled.
Because this is a full rebuild, listening sockets close and reopen and all clients are disconnected. Application logging is configured once at startup and is not re-applied on reload.
Resource Management
- Every buffer is bounded; the drop-not-block policy keeps memory flat under load.
- Network sinks and chain sources accept a
max_connectionscap. Admission is a load-then-check, so a burst can over-admit by roughly one connection. - Chain listeners bound a single line at
core.MaxLogEntryBytes(1 MiB); an oversized line is a protocol violation and terminates the connection. - The
http_chainsource caps each request body atmax_body_bytes. - File sinks rotate on size, cap total rotated size, and honour a retention window.
Performance Notes
- In-memory entry processing is sub-millisecond; the formatter mutex is the only shared serialization point in the hot path.
- File tailing detects new content within roughly 100 ms (fixed poll), while
check_interval_msgoverns how quickly a newly created file is noticed. http_chaintrades latency for efficiency: entries wait up toflush_interval_ms(default 1 s) before a batch is sent.- Scale out with more pipelines per process, more sinks per pipeline, or more nodes chained together.