Spot machines

How SuperCI uses AWS spot machines, what happens when AWS takes one back mid-job, and how to keep a job off spot.

On AWS, jobs run on spot machines unless you say otherwise. Spot machines are AWS’s spare capacity, at a much lower price. The price of that is that AWS can take one back with two minutes’ notice.

When AWS takes a machine back

  1. The machine hears AWS’s notice and stops the runner, so the job fails at once instead of hanging.
  2. The dashboard shows the job as interrupted.
  3. When the job’s run has finished, SuperCI runs the job again. (GitHub cannot restart one job while its run is still going.)
  4. The new attempt does not go to a spot machine: it goes to the next provider in your order below AWS’s spot machines, which is AWS on-demand unless you changed that.

On GitLab the job is retried the same way.

If AWS takes back two machines within a quarter of an hour, new jobs skip spot machines for half an hour.

When spot machines are the only place a job can run (on-demand is off and no other provider fits), it runs on spot again rather than not at all.

When AWS has no spot machine

SuperCI asks each of your regions in order. If none has a spot machine for the job, the job goes to the next provider in your order. With nothing else that can run it, it waits and keeps trying.

Jobs that must not be interrupted

A job that should never run twice, like a deploy, can ask for an on-demand machine in its label:

runs-on: superci-ondemand

You can also make on-demand the default for every job, under Workflows.

What it costs

A spot machine’s price is its zone’s spot price while it ran; an on-demand machine’s is AWS’s list price. The dashboard estimates a job’s cost while it runs and settles it at the prices AWS billed once it has ended.