Amirali YaghoutiSenior Software Engineer

python Case study

WireGuard Soft-fix Watchdog

A WireGuard tunnel reports itself as connected long after it has stopped passing traffic. That gap is the whole problem. The interface looks fine, the work does not, and you find out from a failing command rather than from the VPN.

The business problem

WireGuard is connectionless, so an interface being up says nothing about whether packets reach the other end. The symptom is a tunnel that silently stops working, usually after the machine sleeps or the network changes. The fix is always the same manual restart. Doing that by hand means noticing first, and noticing is the part that does not happen reliably.

What I delivered

  • A macOS-side watchdog that checks connection health continuously rather than on demand.
  • Health checks that test whether traffic reaches the far side. Interface state is not evidence of a working tunnel.
  • Automatic recovery of the session when a check fails, so the common case repairs itself before anyone notices it broke.
  • A record of failures over time, which turns a recurring annoyance into a pattern I can diagnose.

Technical approach

  • I define health as traffic reaching the other end. A check based on interface or process state would call the exact failure this tool exists to catch healthy.
  • Recovery restarts the session instead of attempting a partial repair. On a connectionless tunnel, the blunt fix is the reliable one.
  • The watchdog logs failures instead of only acting on them. The log is what separates a flaky network from a configuration problem.
  • It runs on the client machine, the only place that can see whether the tunnel works from where it is actually used.

Result and evidence

The tunnel now repairs itself in the ordinary case. The failure log makes the pattern visible rather than anecdotal, which the manual routine never did.

Commercial value

Small recurring interruptions cost more than they look like they do, because they take attention rather than time.

implementation-brief.readme

Readable implementation brief

implementation_brief {
  project: "WireGuard Soft-Fix Watchdog"
  runs_on: "the client machine (macOS)"
  health_definition: "traffic reaches the far side --
                      NOT interface or process state"
  recovery: "restart the session; blunt and reliable"
  record: "failures logged over time to expose the pattern"
  triggers_it_catches: "sleep/wake, network change, silent stall"
}

What this project shows

Measuring health as traffic rather than as interface state is the entire point of the tool. A watchdog that checks the wrong signal reports green right through the outage.

Logging the failures instead of just fixing them turned this from a workaround into something that told me what was actually wrong.