indidev

joined 1 month ago
 

Been running a handful of scheduled tasks on two VPS instances for about a year now — nightly DB dumps, cache warmers, a couple of data sync scripts. Had everything piped to log files and thought I was covered.

Then last month one of the sync scripts started hanging mid-execution. It never crashed, never threw an error, just... sat there. The log showed the start timestamp but no finish line. Took me four days to notice because I wasn't checking logs daily (who does).

After that I looked into what people use to catch this kind of thing.

What I tried:

Healthchecks.io — solid, does the job. You curl a URL at the end of your cron and if the ping doesn't arrive within the expected window, you get an alert. Simple concept.

I also tested WatchCron which works on the same principle but felt a bit snappier to set up for multiple jobs. Has a dashboard that shows timing patterns across runs, which helped me spot that one of my backup jobs was gradually taking longer each week — would've missed that with just pass/fail alerts.

The pattern that works for me now:

#!/bin/bash
# at the end of each cron script
curl -fsS -m 10 --retry 3 $MONITOR_URL > /dev/null

If the script hangs or exits early, the ping never fires, and I get a Slack notification within minutes.

One thing I changed — I stopped redirecting cron output to /dev/null. Feels obvious in hindsight but I see it everywhere in tutorials. If your task does fail, you want that output in the mail spool or a log, not gone.

What's your setup for catching silent failures? Curious if anyone's doing something beyond the "ping on success" model — like tracking execution duration or exit codes.

32
submitted 1 week ago* (last edited 1 week ago) by indidev@lemmy.world to c/programming@programming.dev
 

Deployed a backup script on a new server, tested it manually — worked fine. Set up crontab, came back next morning, no backup. Turns out /usr/local/bin wasn't in cron's PATH so pg_dump just silently didn't exist.

Switched to absolute path and it worked immediately. Fifteen minutes of debugging for a one-character fix. I keep making this mistake every time I set up a new box. At this point I should just have a checklist taped to my monitor.

 

Ran into a situation last week that got me thinking about how most teams handle visual changes in production.

We pushed a CSS update that was supposed to fix button alignment on one page. Looked fine in staging. Looked fine in the PR preview. Got merged, deployed, and broke the card layout on three other pages because of a shared utility class.

Caught it about four hours later when a customer mentioned it in a support ticket. Not a great look.

What we tried

After that incident we talked about adding visual regression testing to CI. Looked at a few options:

  • Percy — solid, but the pricing gets steep once you have more than a handful of pages and multiple PRs a day

  • Playwright screenshots in CI — free, but you need to maintain baseline images and the diffs are noisy. Every font rendering difference across runners triggers a false positive

  • Manual spot checks after deploy — what we were doing. Obviously not sufficient

We ended up going with Playwright screenshot comparisons but with a pretty generous diff threshold (0.3%) to cut down on false positives. It catches big layout shifts but ignores sub-pixel rendering differences. Good enough for now, but not ideal.

The production monitoring gap

CI-based visual testing only covers what happens before deploy. It doesn't catch things that break after — CDN caching serving old assets, third-party scripts injecting unexpected elements, A/B test variants rendering wrong.

Some teams run synthetic monitoring that takes periodic screenshots of production pages and compares them. That covers the post-deploy gap but it's a separate system to set up and maintain.

Curious what others do

How are you handling visual verification in production? Specifically:

  • Do you run any visual checks post-deploy, or only in CI?

  • If you use screenshot comparison, how do you deal with dynamic content (dates, user-specific data, ads)?

  • Anyone running continuous visual monitoring on production pages on a schedule?

Feels like there should be a simpler answer than "maintain 200 baseline images and pray the CI runner has the same font rendering as last time."

 

Been dealing with this more often lately. Tests pass on my machine, I push, and CI blows up. Usually it's one of these:

  • Different Node/Python/whatever version
  • Missing env vars that exist in my .env but not in CI secrets
  • File system case sensitivity (macOS vs Linux)
  • Some flaky test that depends on timing

My current debugging flow is pretty basic: check the logs, compare versions, run the exact same Docker image locally if I can. But it still eats 20-30 minutes each time before I figure out the actual problem.

Anyone have a more systematic approach? Like a quick checklist you run through before you even look at the logs?

Also curious — do you replicate your CI environment locally with something like act (for GitHub Actions) or just trust the remote runner?

 

Had a backup script running via cron for months. Worked fine until it didn't — turns out the disk filled up three weeks ago and the job started failing silently. Nobody noticed until we actually needed a restore.

The obvious answer is "check your logs" but let's be honest, nobody's reading cron logs daily for 15 different scheduled tasks across 4 servers.

What's your setup for making sure crons are actually completing? Do you just grep logs periodically, or do you have something more structured? Curious how others handle this without turning it into a whole project.

 

Honest question. I've got automated daily pg_dump backups going to S3 and I check that the files are there, but I've never actually tried restoring one on a fresh instance to see if it works.

Feel like this is one of those things where you assume it's fine until you desperately need it and find out the dumps were corrupted for 3 months.

Anyone have a setup where restores get tested automatically? Or is that overkill for side projects?

[–] indidev@lemmy.world 8 points 1 month ago

Yeah, that's the silver lining I guess. Still felt dumb staring at broken builds because of a rename that nobody asked for.

 

Spent an hour today renaming env vars across three services to make them "consistent." Broke staging in the process because one service cached the old values. Should've just left the mess alone — it worked fine before I touched it.

 

Working on a project where I need to grab screenshots of pages on a schedule — mostly for QA and keeping a visual history of what changed and when.

Right now I'm running Puppeteer in a cron job on a cheap VPS. It works, but honestly it's a pain to maintain. Chromium eats RAM like crazy, sometimes it hangs and the whole thing needs a restart, and the screenshots come out wrong on pages that load content dynamically.

A few things I've been struggling with:

  • Pages with lazy-loaded images — half the time the screenshot fires before everything renders
  • Cookie consent banners blocking the actual content
  • Memory usage goes through the roof when I try to do more than ~50 pages in a batch

I've looked into Playwright as a replacement but from what I can tell the resource usage is about the same. Also tried running headless Chrome in Docker which at least makes cleanup easier, but didn't solve the core problems.

Curious what others are using. Are screenshot APIs worth it for this kind of thing, or is self-hosting still the way to go? Anyone running something similar at scale?

 

I run a few side projects and I've gone through different stages of monitoring them. First it was just checking manually if the site loads. Then I added a simple curl ping in cron. Then I started tracking response times, certificate expiry, even visual changes on pages.

At some point I realized I was spending more time building monitoring than the actual product. Classic trap.

Curious what other devs use for keeping an eye on their stuff. Do you go with a hosted service, self-host something like Uptime Kuma, or just wing it with scripts?

 

Hey everyone! Just signed up on Lemmy. I've been running self-hosted services for a while now and looking forward to learning from this community. Glad to be here.