Run signals
A plain heartbeat tells Mast that a job finished. Run signals tell it more: when the job started, whether it failed, how it exited, and which run a signal belongs to. With them a vital can page on a job that hangs instead of waiting out the whole period, and it can let a job that flakes once try again before anyone is woken up.
Every example below uses a vital's URL, https://mast.tissue.dev/mk_8e2a…. Each signal is a path added to the end of it. GET and POST both work.
Start
curl -fsS https://mast.tissue.dev/mk_8e2a…/startA start says a run began. It is not a heartbeat: the vital's state stays where it was, and a job that keeps starting without ever finishing still goes flat. What a start buys you is the longest-run limit. Set Longest run on the vital, and a run that is still open after that long pages right away, without waiting for the period or the schedule to catch up. A backup that hangs at 04:00 pages at 04:30, not the next morning.
The end of the run is the next heartbeat, fail or exit code. Mast records how long the run took, and the time shows in the vital's history.
Anything in the body of the start is kept as the run's label, a git sha or a dataset name. It is shown in the history and never pages.
Fail
curl -fsS -X POST https://mast.tissue.dev/mk_8e2a…/fail -d body="disk full on /var/backups"A fail says the job ran and it went wrong. It flatlines the vital at once and pages, carrying the body as the reason. The next heartbeat closes the page.
Exit code
17 3 * * * curl -fsS https://mast.tissue.dev/mk_8e2a…/start && /usr/local/bin/backup.sh; curl -fsS https://mast.tissue.dev/mk_8e2a…/$?Put the job's exit status on the end of the URL and one line covers both outcomes. /0 is a heartbeat. Any other number, from 1 to 255, is a fail with that code on the card and in the history. The body is kept as the job's output and never becomes a page of its own, so a job that prints a warning on a good run does not page anyone for it.
Run ids
rid=$(date +%s)-$$
curl -fsS "https://mast.tissue.dev/mk_8e2a…/start?rid=$rid"
/usr/local/bin/backup.sh
status=$?
curl -fsS "https://mast.tissue.dev/mk_8e2a…/$status?rid=$rid"When two runs of the same job can overlap, say a slow run that is still going when the next one starts, Mast cannot tell whose end is whose. Add ?rid= with the same value to the start and to the end, and the end is matched to its own start. The duration is measured from that start, and the other run stays open.
A run id is up to 64 letters, digits, dashes, underscores and dots. One that breaks those rules is ignored, not refused, and the signal is still counted. Without a run id, an end is paired with the most recent open start.
Failures before a page
By default the first failure pages. Failures before a page raises that: set it to 3, and it takes three failed runs in a row before anyone is told. The first two are recorded in the history and on the vital, and the phone stays quiet.
A failed run is a /fail, a non-zero exit code, a run that went past its longest-run limit, or a run that started and never reported an end. Any heartbeat resets the count to zero. A job that just goes quiet, with no start and no end, is not a failed run: it pages when its grace runs out, whatever the threshold says.
On the API it is fail_threshold, from 1 to 100.
Early warning
Turn on Early warning and the vital learns how long the job usually takes to report, from its last 20 heartbeats. When a run is well past its usual time but still inside grace, you get one quiet notification saying it is later than usual. It has no sound and it is not a page. It needs at least five heartbeats before it says anything, and it sends at most one notification between two heartbeats.
On the API it is smart_early_warn, true or false.
Unknown
Sometimes the problem is on our side. If the monitor could not accept beats for a stretch, a vital that was due during that stretch has no way to prove it ran. Rather than page you for our outage, Mast marks that vital unknown and says why:
the monitor could not accept beats 03:10–03:24 UTCAn unknown vital pages nobody. It goes back to ok on its next heartbeat. If no heartbeat comes by the time the next one is due, it is judged the normal way and pages if it is late. The reason and the window are on the vital as vital_reason and blind_window, and on a public status page the component shows as unknown with the same sentence.