Field notes

nginx Done Right—Well, How I Like It

Performance brought me to nginx in 2005. Two decades later, my definition of right includes where it writes, whom it trusts, what it serves, and what happens when an operation only partly succeeds.

By Mikinho, Chief Mad Hatter 34 minute read

It started on a Linode

My first Linode VPS, in 2003, ran Apache httpd. At the time, Apache and IIS were the realistic choices in front of me. I liked Apache considerably more than IIS 6.0, and I had no desire to build my public hosting around Windows Server. Even so, Apache felt clunky to me, and much of the web-server world I encountered seemed to revolve around PHP.

I discovered nginx in early 2005. Its event-driven, asynchronous approach hooked me almost immediately. A worker could handle many connections without dedicating a thread or process to each one. That architecture made sense to me, and the combination of performance and low resource use was compelling. The nginx team's later explanation of its process model describes the design that first caught my attention.

I still have an appreciation for Linode. That first VPS was where this part of my history began.

More than twenty years later, I still like nginx. My expectations of a good nginx installation have grown considerably.

Performance brought me here. Today, I want a tested agreement among the binary, its configuration, the operating system, the deployment process, and the application behind it.

The second half of this article's title matters. This is how I like to run nginx, based on the systems I maintain and the choices I am willing to own. My current implementation reflects a preference for CentOS Stream, RHEL, and Rocky Linux. The principles reach further than those distributions; the implementation has a particular home.

The configuration excerpts below draw on my baseline at this revision. They show the parts that explain a decision; their linked source files supply the surrounding configuration. Names such as sample_node are the repository's reusable examples.

A media box, and a familiar strength

Media was familiar territory from my time as a Microsoft MVP for Windows Media Center. Years later, work on a media box brought nginx back into the picture and reinforced my appreciation for its efficiency.

That experience belongs in this history because it broadened where I thought nginx could be useful. The original attraction remained the same: I liked how much work it could do with modest demands on the system around it.

I had opinions in 2016, too

In November 2016, I committed a systemd unit with the message “Better nginx systemd...welll how I like it.” Apparently, I've been qualifying this opinion for a while.

Those early files already show some of the concerns I still have. The unit created nginx's directories, used a private temporary directory, and restarted the service after failure. A default server made room for Let's Encrypt validation, and I soon moved the status endpoint into its own configuration file.

They also contain things I would handle differently today. The old default redirect says return 301 https://$hostcs;. Yes, $hostcs. The service unit has its nginx -t preflight commented out.

Those are useful details to leave visible. They show the distance between writing down what I wanted and consistently checking that the files expressed it. The current unit validates the configuration before both startup and reload.

nginx.service
Type=exec
ExecStartPre=/usr/sbin/%p -t -q -c /etc/%p/%p.conf -g "daemon off;"
ExecStart=/usr/sbin/%p -c /etc/%p/%p.conf -g "daemon off;"
ExecStartPost=!/usr/local/libexec/nginx-runtime-verify --master-pid $MAINPID
ExecReload=/usr/sbin/%p -t -q -c /etc/%p/%p.conf -g "daemon off;"
ExecReload=/bin/kill -s HUP $MAINPID

In this unit, %p expands to nginx. The two reload commands run in order: validation has to succeed before systemd sends the master process HUP. This also gives me the same preflight when a human reloads the service and when another service requests it. The nginx master performs its own checks on reload; the explicit preflight makes a rejected candidate visible in the service operation before the signal is sent.

The master stays in the foreground so systemd supervises the process it started. Type=exec confirms execution, not application readiness; the post-start checker adds bounded checks of the live master and workers. The service supplies daemon off; itself, so the nginx configuration must not duplicate that directive.

The repository itself has a small 2016 snapshot, followed by a long gap before a concentrated effort in August 2026 to formalize the accumulated practice. It is an incomplete record of the intervening years. The recent work brought the configuration, operating-system integration, deployment choices, and validation together in one place.

The definition got larger

Around 2020, my attention shifted. As cyberattacks became a more prominent concern in my work, I spent more time thinking about security, privacy, and the boundaries between systems. Performance and resource use still mattered. I wanted more explicit answers about everything surrounding them.

Which machine is allowed to tell the application who the client is? What happens when a request names a host I have never configured? What information am I putting into an access log? Where can the service write? What will an unattended maintenance job do when some of its work succeeds and some fails?

I don't have a major nginx production failure to offer as the dramatic turning point. Most of this is the gradual accumulation of expectations. One particularly useful lesson came from my own certificate-renewal automation.

A renewed certificate still has to reach the client

I once arranged for nginx to reload only if the entire certificate-renewal run succeeded. It worked beautifully for as long as nothing failed, which is how I know it was wrong.

The first time one certificate failed, Certbot renewed the others and my automation skipped the reload. nginx kept serving the old certificates while their replacements sat on disk—renewed, valid and unused. I found out the way nobody wants to: a client called. A manual reload finished the job. The automation had done exactly what I told it to, which was the problem.

A renewal run succeeds for two sites and fails for a third. The old all-success condition skips reload. The current service attempts validation and HUP while keeping the renewal failure visible. A fresh TLS connection must then verify the certificate the client receives.
Partial renewal needs both failure reporting and a reload attempt. The final check is the certificate served to a new client connection.

The mistake was in the condition I had put around the reload. I had treated renewal as one indivisible success or failure, even though it could produce several useful results and a failure in the same run. Certbot's renewal documentation makes that distinction important: the command operates across certificates, and failed renewals can give it a nonzero exit status.

The relevant part of my current renewal service is small:

certbot.service
ExecStart=/bin/certbot renew --quiet --agree-tos
ExecStopPost=/bin/systemctl reload nginx.service

The complete service attempts the reload after Certbot exits, including after a partial failure. Certbot's original failure remains visible. The nginx service runs its configuration check before sending the reload signal; a failed preflight or signal command is reported by systemd.

There is an important limit to that result. In this unit, a successful systemctl reload confirms the preflight and signal command succeeded, not that nginx finished adopting the new configuration. The master handles HUP asynchronously and can reject the change and continue with its old configuration. The nginx error log and the certificate presented on a fresh TLS connection complete the check.

For a direct TLS endpoint, this read-only example displays the leaf certificate a new connection receives. Replace example.com with the hostname being checked; this is an inspection command, not a complete certificate monitor:

set -o pipefail
timeout 10s openssl s_client \
    -connect example.com:443 \
    -servername example.com \
    -verify_hostname example.com \
    -verify_return_error </dev/null |
    openssl x509 -noout -serial -dates -fingerprint -sha256

The OpenSSL client verifies the hostname and chain instead of silently accepting errors, and the certificate output gives the serial, dates, and SHA-256 fingerprint to compare against the certificate that was meant to be there. A CDN or load balancer in front presents its own certificate; check every terminator, not one.

Successfully renewed certificates therefore get a chance to become active while unsuccessful renewals still require attention. Reloading cannot fix a certificate that failed to renew, and a successful reload does not erase that failure.

There is a tradeoff: this unit attempts a reload after every renewal run, including a run with nothing to renew. I accept that simple, consistent behavior. Certbot deploy hooks offer another design, but I keep this baseline's reload in one place so it has one service-level result to inspect. The unit is written for the webroot renewal flow; a standalone authenticator that needs to bind a port would require a different service policy.

A separate disk check has a narrower job: find incomplete certificate lineages and certificates approaching expiration. One healthy certificate must not hide another with a missing fullchain.pem, and a dangling symlink is a failure, not a quirk. What a client actually receives is the endpoint check's job, above.

First issuance needs the same explicitness. The issuance wrapper selects the Let's Encrypt staging or production ACME directory with --server, and staging adds --dry-run so no test certificate is persisted. These are the two endpoint selections, not complete issuance commands:

# Staging: test without saving a certificate.
--server https://acme-staging-v02.api.letsencrypt.org/directory --dry-run

# Production: selected only after the same names passed staging.
--server https://acme-v02.api.letsencrypt.org/directory

Pinning the endpoint is not enough on its own, because Certbot's server-selection rules let an inherited cli.ini redirect an explicit production selection to staging. So before invoking Certbot, the wrapper rejects staging, test-cert, and dry-run settings in inherited system and user cli.ini files, false-valued entries included. Testing mode belongs to the operation I selected, not to an ambient default.

For me, this is a particularly clear example of the agreement between components. Certbot obtains the certificate, systemd controls the operation and records its result, and nginx has to load the renewed material. The useful outcome is the certificate a client receives.

That same preference for explicit ownership extends to something much smaller: a directory under /run.

Yes, I want a dedicated runtime directory

The lack of a dedicated nginx directory under /run bothers me. It may sound like an unusually small thing to care about. I care about it anyway.

I want the service definition to describe where its runtime files belong, who owns those directories, and how their lifetime relates to the service. My current systemd unit declares runtime, state, and log directories explicitly. The PID file lives under /run/nginx.

nginx.service
# The system manager supplies the default root user.
Group=nginx
SupplementaryGroups=nginx
UMask=0027
RuntimeDirectory=%p lock/%p
RuntimeDirectoryMode=0755
StateDirectory=%p
StateDirectoryMode=0755
LogsDirectory=%p
LogsDirectoryMode=0750

systemd creates the runtime directories when the service starts; state and log directories have persistent lifetimes. It tracks the foreground master directly. nginx still writes its own PID file, which the startup checker compares with the supervised process. The corresponding nginx configuration agrees with those paths:

pid /run/nginx/nginx.pid;
lock_file /run/lock/nginx/nginx.lock;
error_log /var/log/nginx/error.log warn;

That agreement avoids a surprisingly ordinary class of problems: one file expects a PID in one place, another watches a different place, or a directory exists only because somebody created it during installation. I want a cold start after reboot to establish its own runtime layout.

The modes and ownership are deliberate. The master runs as root:nginx, using the system manager's default root user and an explicit group list without inherited root-group access. nginx changes its workers to the nginx account. Runtime directories use 0755; the log directory is root:nginx 0750, allowing workers to traverse it without granting directory write access. Keeping nginx in the foreground also preserves UMask=0027: its daemonization code otherwise resets the mask to zero. A creation mask does not repair existing files, and directory permissions remain a separate decision from permissions on files or sockets inside them.

That distinction connects directly to the filesystem boundary in Why I Use Unix-Domain Sockets Behind nginx. Its application-owned runtime directory uses 0710, while the Unix socket grants the proxy the access it needs. The owners and consumers differ, so the modes differ. A dedicated directory gives both services a place to express that relationship.

The next part of the nginx unit limits the environment around those paths:

nginx.service
NoNewPrivileges=yes
PrivateDevices=yes
PrivateTmp=yes
ProtectSystem=strict
ProtectHome=yes
ProtectKernelModules=yes
ProtectKernelTunables=yes

ProtectSystem=strict makes the service's filesystem view read-only apart from permitted writable areas, including the managed directories. Private temporary storage reduces sharing with other services. The kernel protections remove activities a web server should not need, and NoNewPrivileges prevents gaining privileges through an executable transition. These controls are described in systemd.exec. They constrain what a compromised process could attempt as well as catching accidental writes.

Declaring those settings does not apply them to a running master: systemctl daemon-reload rereads the unit, and nginx's HUP reload keeps the old process. Host setup therefore does a planned restart and checks the resulting process; routine configuration and certificate reloads keep the validated reload path.

The startup checker reads /proc for the master and active workers—identities, Umask, NoNewPrivs, capability sets, seccomp—and checks that the PID file agrees with the supervised process. Ordinary startup needs at least one active worker; host setup supplies the reviewed worker count and requires an exact match. Only the checker uses systemd's ! command prefix, which keeps the manager's root credentials so it can see workers across the /proc visibility boundary; nginx itself stays root:nginx, and the rest of the sandbox still applies.

Resource limits need arithmetic as well as intent. The shared unit keeps TasksMax=512, nginx runs worker_processes auto with aio threads, and the default pool has 32 threads per worker—so sixteen workers need 1 + 16 × (1 + 32) = 529 tasks before a graceful reload doubles them. Host setup checks that budget against the reviewed worker and thread counts and refuses to apply an insufficient one. A documented limit is only useful when the service can actually operate within it.

I dislike configurations that leave that operating-system side unexplored, including packaged defaults I have encountered. Installing a web server creates a service with responsibilities and permissions. I want to review both.

The unit gives the master an explicit capability allowlist for its privileged work and denies mount-related syscalls. Workers must have no effective, permitted, or ambient capabilities. Selecting QUIC BPF adds a separate capability extension, including a kernel-dependent CAP_SYS_ADMIN fallback, while retaining the mount denial. These are contracts to validate against the actual build, kernel, and policy; a restriction earns its place by working with the service I intend to run.

The Node.js Website That Cannot Reach the Internet makes this difference concrete in its documented exceptions. That small application has no reason to contact the network and can have an empty capability set. nginx owns the public listeners and has a different job. Copying either unit onto the other service would ignore the reason those controls were chosen.

Protocol support is a deployment choice

Configurations that leave HTTP/2 unused bother me, too. HTTP/3 belongs in the discussion as well. I want the protocols I serve to reflect an intentional decision about the site.

HTTP/2 can carry multiple request streams on one connection. HTTP/3 runs those streams over QUIC, avoiding TCP's connection-wide blocking between streams when packets are lost. Neither makes every request faster in every environment, but both are capabilities I want to evaluate rather than leave unused by habit. The protocol designs are described in RFC 9113 and RFC 9114.

My http block selects the protocols:

http2 on;
http3 on;

Then the sample Node.js site declares the listeners in its TLS server block:

listen 443 ssl;
listen [::]:443 ssl;
# The default HTTPS server owns reuseport for these shared UDP sockets.
listen 443 quic;
listen [::]:443 quic;

Its includes/http3.conf adds the advertisement:

add_header Alt-Svc 'h3=":443"; ma=86400' always;
HTTP/1.1 and HTTP/2 reach the nginx TLS listener over TCP 443. HTTP/3 reaches the QUIC listener over UDP 443. Both paths require the matching build, listener and firewall rule; Alt-Svc advertises the HTTP/3 path.
Protocol support spans the build, the listener and the network path. TCP 443 and UDP 443 are separate firewall decisions.

The firewalld service includes TCP 80 for redirects and ACME, TCP 443 for HTTPS, and UDP 443 for QUIC. The external firewall needs a corresponding policy. Opening TCP 443 alone cannot make the QUIC listener reachable. The official HTTP/2 module documentation and QUIC guide cover the build and configuration requirements.

My baseline includes HTTP/2 and HTTP/3 support, with the site's listeners and HTTP/3 advertisement completing the setup. As of this writing, nginx still describes its HTTP/3 module as experimental. That is part of the decision I am making when I use it.

This is one reason my definition of configuration includes the binary and the operating system. Package contents vary. I want to know what my installed build supports and verify the behavior that reaches the client.

The client-facing protocol is also a different decision from the transport to the application. nginx can accept HTTP/2 or HTTP/3 publicly and proxy HTTP/1.1 over a local Unix socket. The companion post's separate-hops explanation covers why those choices do not conflict. Enabling HTTP/3 at the edge does not require the Node.js application to become a QUIC server.

Own the public boundary

The same reasoning applies to requests. An nginx server directly facing the internet and one sitting behind a trusted proxy receive different evidence about the client. My configuration makes that distinction an explicit deployment choice.

At a direct edge, the shared proxy configuration establishes the values it sends to the application, including X-Forwarded-Host:

proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header X-Forwarded-Host $host;
proxy_set_header X-Forwarded-Proto $scheme;

The important choice is replacing the forwarding header. A visitor can send an X-Forwarded-For header just as easily as any other request header. If the application trusts a chain that visitor started, it may assign rate limits or log attribution to an address the visitor chose. At my direct edge, nginx sends the address it has established instead.

The hostname deserves the same treatment. An application that trusts its proxy may derive its hostname from X-Forwarded-Host, even when the ordinary Host header is correct. Both the HTTP and WebSocket includes now overwrite it with nginx's $host. This closes a separate input path used by some frameworks for routing, redirects, or generated links; it does not replace the site's allowed-host policy. Express's proxy guidance explicitly calls out forwarded host, address, and scheme together.

When another proxy sits in front, the optional real-IP profile changes how that address is established:

include trusted-proxies/*.conf;

real_ip_header X-Forwarded-For;
real_ip_recursive on;

Those deployment-local files contain set_real_ip_from entries for the immediate trusted peers. The real-IP module provides the mechanism; the deployment supplies the trust decision. The shared proxy include can then continue forwarding the resulting $remote_addr. TLS termination and scheme reporting also need to match the actual topology. Selecting trusted addresses alone does not describe every aspect of an upstream proxy.

The application must trust only the proxy that actually connects to it. Why I Use Unix-Domain Sockets Behind nginx follows that boundary through the local connection. The Node.js Website That Cannot Reach the Internet shows how I combine it with a much narrower application policy:

Internet requests reach nginx, which owns TLS and client identity. nginx serves static files directly and sends application requests through a permission-controlled Unix socket to Node.js. The application can answer through that socket while its service policy denies IP networking.
Requests and responses use the Unix socket. nginx owns the public listeners; the application service separately restricts its own IP networking.

The public sample expresses the connection like this:

upstream sample_node_upstream {
    server unix:/run/sample_node/sample_node.sock;
    keepalive 32;
}

Its application location uses proxy_pass http://sample_node_upstream;. The common include selects HTTP/1.1 and clears the Connection header so the upstream keepalive pool can be used. The 32 is a per-worker idle-connection cache, not a cap on all connections to the application; the upstream documentation is explicit about that distinction.

For a local application, a Unix socket lets filesystem permissions and SELinux participate in deciding who can connect. It also removes the need for an application TCP listener. The socket itself does not prevent outbound networking: the Node.js service's IP restrictions provide that separate control. A remote upstream would naturally need a different transport.

Static files are another useful division of responsibility. nginx can serve them without making Node.js open and stream each asset. My baseline enables sendfile. These excerpts from the sample Node.js site show where a request goes next:

root /var/www/sample_node/wwwroot;

location ~* ^/(css|js|img|fonts|static)/ {
    add_header Cache-Control $cache_control_fallback;
    access_log off;

    try_files $uri @app;
}

location @app {
    limit_conn conn_general 20;
    limit_req zone=req_general burst=30 nodelay;

    proxy_pass http://sample_node_upstream;
}

For /css/app.css, try_files checks the corresponding file under that root. If it exists and can be served, nginx handles the bytes. If it is missing, @app sends the request to Node.js. That fallback is deliberate: the application gets to decide whether the URL is an application route or a missing resource. This excerpt is not a standalone server block; the full sample defines the cache-control map, rate-limit zones through its required profiles, proxy headers, and other policy it depends on.

An asset-only namespace can make a different choice: try_files $uri =404; returns a local 404 instead of involving the application. A missing JavaScript file must not accidentally receive a successful HTML application page, so the fallback's status and content type deserve a test. Choosing @app and choosing =404 express different contracts; neither follows automatically from enabling sendfile.

Serving existing assets directly keeps application capacity available for application work. The immutable, read-only release layout then makes the relationship easier to operate: deployment writes the release; the serving processes read it.

The failure path needs an explicit contract too. The Node sample now reuses the shared maintenance location instead of defining a slightly different copy:

error_page 502 503 504 /maintenance.html;
include includes/maintenance.conf;

The included location is small:

location = /maintenance.html {
    internal;
    try_files /maintenance.html =503;
}

If the socket is down and no document has been provisioned, the response stays a service failure: 503, not a misleading 404. If the document exists, nginx serves it with the original upstream status. A direct request cannot reach the internal location. Three cases, three tests.

Refuse an unknown host, and keep useful logs

I want unknown hosts rejected deliberately. nginx's default server is the fallback for a listener; server_name _ is just a deliberately unused name, not a special catch-all operator. My HTTP default owns default_server, permits the ACME challenge location, and closes other requests with nginx's 444 response behavior. It does not send an arbitrary Host value into a redirect.

The HTTPS default also owns the default listeners and declares:

server_name _;

ssl_reject_handshake on;

location / {
    return 421;
}

For an unknown TLS server name, ssl_reject_handshake rejects the handshake. The 421 covers HTTP requests handled by that default server after a connection has reached the HTTP layer. This avoids accidentally serving the first real site in the directory as everybody's fallback. Each real site names the hosts it serves. The nginx request-processing guide explains how those server selections happen.

I also want enough information to diagnose a problem without collecting every value that passes through it. The logging profile derives the original path without its query:

map $request_uri $access_log_request_path {
    default "";
    ~^(?<access_log_path_capture>[^?]*) $access_log_path_capture;
}

A second map records only whether $args is nonempty:

map $args $access_log_has_query {
    ""      false;
    default true;
}

The JSON access log uses those values in place of the full request URL, and the profile strips query values from the referer too. Paths help identify a broken route; timing and status fields help explain what happened. Search terms, reset tokens, or authorization codes in a query usually add exposure without helping that diagnosis. Request IDs connect the nginx line to the application line without relying on those values.

For example, a request for /reset?token=example-canary produces these two fields in an access-log record that uses this profile, assuming the request's arguments are not rewritten:

{ "request_path": "/reset", "query_present": true }

That is an illustrative excerpt, not the complete log line or a real reset token. The example-canary value has no place in those fields. The boolean comes from $args, which nginx can change during rewriting; it describes the arguments at logging time, while the path is derived from the original $request_uri.

This protection belongs to the configured access log. It does not sanitize nginx's error log, an application's own logs, or telemetry collected elsewhere. Those paths need their own checks; a harmless, distinctive canary value is useful for finding an unexpected copy in a test environment.

This is data minimization, with limits. Client addresses, paths, user agents, and other request information remain in the logs, and sensitive values can appear in a path. Retention and access still matter. JSON escaping protects the structure of a log line; it does not make its contents anonymous.

Test across the boundaries

The public baseline brings these policies together with systemd, SELinux, firewalld, sysctl settings, log rotation, certificate renewal, and deployment tooling. Different deployments select different features. Public files contain reusable policy and examples. Actual client identities, addresses, credentials, and deployment evidence stay in the appropriate private records.

Validation has to follow the same breadth. nginx -t checks syntax and tries to open referenced files, as the command-line documentation explains. That is an essential check with a defined scope. It cannot establish whether the application interprets a forwarded address correctly or whether a browser can reach the QUIC listener through the firewall.

The repository includes stable and mainline syntax validation alongside runtime behavior tests. The one I care about most sends a request with a forged Proxy header, a forged X-Forwarded-For and a forged X-Forwarded-Host at a local echo upstream, and expects it to report:

proxy=
xff=127.0.0.1
xfh=runtime-edge.invalid

The Proxy header is gone, the forwarding chain has been replaced with the actual loopback peer, and the host is the one nginx was configured with rather than the one the visitor supplied. A 200 response alone would prove none of that.

The Linux fixture activates the actual Node sample over a Unix socket alongside the WordPress checks. Its assertions cover forwarding, static-file fallback, maintenance responses, JSON access-log query redaction, foreground mask retention, and log reopening. Portable tests cover certificate fixtures, issuance configuration, and the host-setup control flow. A mocked service operation does not prove Linux sandbox enforcement; a separate booted-systemd test exercises the startup checker under the actual restrictions. None of that proves the deployed machine behaves; it proves the configuration can. The machine gets its own checks after every deploy. The certificate story above is what trusting the intention instead of the result looks like, and once was enough.

That is why I check what an unknown host receives, what identity arrives at an upstream, whether a query value appears in the log, and what happens when a static file is missing. A valid configuration can still implement the wrong answer to any of those questions.

The Unix-socket post's observability section shows how to inspect and exercise that local connection. The Node.js confinement post adds a verification sequence: inspect the effective policy, exercise the allowed request path, and verify the boundaries that should deny access. The configuration is easier to trust when its claims have corresponding checks.

Choose the parts the deployment needs

Brotli, trusted proxies, WordPress caching, post-quantum TLS settings, QUIC BPF, and application-specific proxy behavior are selected deliberately. Their presence in the repository does not mean that every deployment enables them.

That distinction matters when sharing a configuration. A useful baseline gives me a consistent place to express the choices. Each server still needs choices that fit its application and operating environment.

I also want to acknowledge GetPageSpeed. For people running and maintaining nginx on the RHEL-family systems I favor, I strongly recommend considering a subscription. Convenient access to modules, especially Brotli, is something I value. Their Brotli installation guide shows the packaged route. Making a module straightforward to install and maintain is useful work, and I appreciate it.

My preference for nginx is contextual. When all I need is a simple reverse proxy or TLS termination, I also use Caddy: a self-hosted UniFi controller or a webhook-only application can be a good example. The amount of configuration I am prepared to maintain should fit the job.

How I like it, today

nginx first appealed to me because it was fast, efficient, and elegant. I still value those qualities. Over time, I have come to expect the surrounding system to be equally considered.

I want to know where the service writes, whom it trusts, what it exposes, and how it behaves when maintenance only partly succeeds. I want deployment to make those choices explicit and validation to exercise them.

That's my current version of nginx done right: a system whose behavior I can explain, test, and maintain. The “how I like it” part leaves room for experience to change the answer again. It has been doing that since a commit message in 2016, and I see no reason it would stop.

I plan to follow this with a separate Field note on proactive alerting: failed or partially successful certificate renewals, certificates approaching expiration at the endpoint a client actually reaches, rejected nginx reloads, and other actionable failures recorded in nginx and application logs. Recording a failure is only the first step; somebody needs to know about it while there is still time to act. That follow-up will cover useful notifications, repeated-alert suppression, recovery notices, and checks that the monitoring itself is still working, without turning every log line into an interruption or copying sensitive request data into an alert.

Its acceptance check will be a controlled failure that reaches the intended recipient, a repeated failure that respects suppression, and a recovery that produces the expected notice. A stale or broken monitoring path must also become visible. Those checks belong to the follow-up; this article does not claim that notification delivery has been installed or verified.

Sources and further reading