Beta. Experimental — core APIs are stable, schema may change between minor versions. Contributions and feedback are welcome.
Open source · Beta

CI at any scale.
Your EC2. Spot pricing.
150ms cold start.

Run hundreds of concurrent GitHub Actions jobs on EC2 you control, billed at spot rates — not GitHub’s per-minute price. Each job gets a fresh Firecracker microVM: dedicated kernel, clean filesystem, boots in 150 ms, gone when it’s done. No shared state. No secret bleed.

Try the full scheduler and worker loop on any machine with BURSTGRID_MODE=simulate — no EC2, no KVM needed.

Job lifecycle
GitHub Actions
workflow_job
queued
HMAC-verified webhook POST
Scheduler
verify · queue
tier · route
job assignment SSE push
Worker Agent
bare-metal EC2
KVM host
Firecracker API boot · config
microVM
isolated kernel
run · destroy
<5ms dispatch
~150ms VM boot
0 shared state

Fit

Is BurstGrid right for you?

BurstGrid is not the right tool for everyone. Here is when it makes sense and when it does not.

Good fit
  • Security or compliance requirements
    SOC 2, HIPAA, or FedRAMP workloads that require kernel-level isolation per job. Containers share a kernel; Firecracker microVMs do not. Secrets and filesystem state are gone when the job exits.
  • High concurrency, steady load
    Running 20+ jobs simultaneously most of the day. Packing many microVMs onto bare metal beats per-instance EC2 pricing above roughly 15 concurrent jobs.
  • Short, frequent jobs
    150 ms VM boot versus 60–90 s EC2 cold start. On commit-heavy repos that gap is the CI wait time, not overhead.
Not a good fit
  • Spiky or low-volume load
    BurstGrid requires warm bare-metal capacity at a fixed floor cost. If load is unpredictable or under ~50 jobs per day, terraform-aws-github-runner provisions ephemeral EC2 spot instances and scales to zero — a better fit.
  • Already on Kubernetes
    ARC (Actions Runner Controller) integrates directly with an existing cluster. If you are already on EKS or GKE, ARC is the right call.
  • Windows runners
    Firecracker is Linux-only. Windows jobs are not supported.

The flow

Webhook to running job in four hops

No Kubernetes. No job-path message broker. Just a webhook, a scheduler queue, a persistent SSE connection, and a VM that exists for exactly as long as the job does.

1
GitHub fires a workflow_job webhook
The moment a workflow job is queued, GitHub POSTs a signed workflow_job event to your scheduler. BurstGrid verifies the X-Hub-Signature-256 HMAC before doing anything else. If the signature doesn’t match, the request is dropped. If the queue is full, it returns 503; GitHub will retry the delivery for up to 72 hours, so nothing is lost.
POST /webhook/github  —  X-GitHub-Event: workflow_job
{ "action": "queued", "workflow_job": { "id": 32145678, "run_id": 9821034, "labels": ["self-hosted", "linux", "burstgrid:size=large"], "status": "queued" }, "repository": { "full_name": "acme/monorepo" } }
2
Scheduler mints a runner token and enqueues the job
BurstGrid authenticates as your GitHub App installation and calls the Runners API to generate a one-time registration token for the target repo. The token and job metadata go into a priority queue partitioned by tier. critical jobs always jump ahead of standard ones. Token generation is protected by a circuit breaker: after 5 consecutive GitHub API failures, the scheduler opens the circuit and returns 503 until the API recovers.
POST /repos/acme/monorepo/actions/runners/registration-token
// one-time token injected into the VM via kernel boot args { "token": "AABCDEFGHIJKLMNOPQ", "expires_at": "2026-08-25T11:30:00Z" }
3
Router matches the job to a worker and pushes over SSE
The router checks the job’s size label (burstgrid:size=large → 4 vCPU / 4 GiB) against every connected worker’s free resources and capability labels. The first worker that fits gets the assignment pushed as a JSON event over its persistent SSE connection, the same open HTTP response it’s been holding since it registered. No database read. No poll. Dispatch latency is typically under 5 ms.
SSE → worker-i-0a3f2b  (persistent connection)
data: { "jobId": "3d9f...", "runnerToken": "AABCDEFGHIJKLMNOPQ", "labels": ["self-hosted", "linux", "burstgrid:size=large"], "vcpus": 4, "memoryMiB": 4096 }
4
Worker boots a Firecracker VM; runner picks up the job
The worker agent calls the Firecracker Unix socket API to boot a microVM with exactly the requested vCPU and RAM. Non-secret network hints are passed as boot args; runner/cache secrets are delivered through Firecracker MMDS by default so they do not appear in guest /proc/cmdline. The init script inside the VM reads MMDS, calls ./config.sh --ephemeral to register with GitHub, then ./run.sh to execute the job. When it exits, the VM is killed, the disk is discarded, and the slot is freed. The rootfs is never reused.
Firecracker API → PUT /boot-source
{ "kernel_image_path": "/opt/burstgrid/vmlinux", "boot_args": "console=ttyS0 reboot=k panic=1 init_on_free=1 nomodule MMDS_MODE=1 GUEST_IP=172.20.0.2 GATEWAY=172.20.0.1" }

Get started

From zero to running jobs

Four steps. Steps 1–2 take minutes. Step 3 handles building, uploading, and provisioning in one command.

1
Create a GitHub App
Go to Settings → Developer Settings → GitHub Apps → New GitHub App. Set the webhook URL to https://your-scheduler/webhook/github and generate a secret. Grant Administration: read & write and Actions: read, then subscribe to workflow_job events. Download the private key PEM and install the App on your org.
Scheduler environment variables
# GitHub App (full CI integration, recommended for orgs) GITHUB_APP_ID=123456 GITHUB_PRIVATE_KEY_PATH=/run/secrets/burstgrid.pem # — or — GITHUB_TOKEN=ghp_xxx (personal access token, simpler for single-repo testing) BURSTGRID_WEBHOOK_SECRET=your-webhook-secret BURSTGRID_WORKER_TOKEN=your-shared-secret # workers must present this as Bearer token BURSTGRID_PORT=8080 # default; set BURSTGRID_ADDR to restrict bind # optional: set OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318 to enable telemetry
2
Run burstgrid setup, then burstgrid deploy, then burstgrid init
burstgrid setup detects your default VPC, public subnet, and the latest Ubuntu 24.04 AMI automatically, generates webhook and worker secrets in AWS Systems Manager Parameter Store (SSM), and writes deploy/terraform/terraform.tfvars for you. Then burstgrid deploy builds the binaries, uploads them to S3, and calls terraform apply. Finally, burstgrid init reads the deployed EC2 Launch Templates and writes their IDs back into burstgrid.config.yaml — the autoscaler won’t launch workers without these. No digging around the AWS console.
Production secret source: SSM SecureString (default). The scheduler and workers fetch /burstgrid/webhook-secret and /burstgrid/worker-token with their instance IAM roles after boot. This keeps secrets out of Terraform state and EC2 user data. Add the GitHub credential you use as /burstgrid/github-app-private-key or /burstgrid/github-token. Use npx burstgrid setup --secret-source terraform only for the legacy path, where secrets are written to terraform.tfvars and embedded in launch data.
Step 1 — scaffold terraform.tfvars
$ npx burstgrid setup BurstGrid setup ✓ Account 123456789012 (us-east-1) ✓ Default VPC vpc-0e1f2a3b4c5d6789 ✓ Public subnet subnet-0a2b3c4d5e6f7890 ✓ Ubuntu 24.04 ARM64 AMI ami-0d3e4f5a6b7c8d9e0 ✓ S3 bucket burstgrid-123456789012 ✓ Webhook secret (generated → SSM) ✓ Worker token (generated → SSM) ✓ Secret source SSM Parameter Store (/burstgrid) Wrote deploy/terraform/terraform.tfvars Next steps: 1. Request AWS vCPU quota increase (required for multiple workers): https://us-east-1.console.aws.amazon.com/servicequotas/home/services/ec2/quotas Search "Running On-Demand Standard (A, C, D, H, I, M, R, T, Z) instances" Each c6g.metal worker uses 64 vCPUs. Default account limit is often 32-64. Also request "All Standard (A, C, D, H, I, M, R, T, Z) Spot Instance Requests". 2. Build rootfs images and upload to S3 (run on an ARM64 host with Docker): # Default rootfs — for standard CI jobs: ./scripts/build-rootfs.sh rootfs/template/Dockerfile /tmp/rootfs-arm64.img 4G arm64 --compress aws s3 cp /tmp/rootfs-arm64.img.gz s3://burstgrid-123456789012/rootfs-arm64.img.gz # Docker-heavy rootfs — for jobs with Docker-in-Docker, databases, etc.: ./scripts/build-rootfs.sh rootfs/ubuntu-docker/Dockerfile /tmp/rootfs-arm64-ubuntu-docker.img 16G arm64 --compress aws s3 cp /tmp/rootfs-arm64-ubuntu-docker.img.gz s3://burstgrid-123456789012/rootfs-arm64-ubuntu-docker.img.gz # Download the Firecracker-compatible ARM64 kernel: curl -fsSL https://s3.amazonaws.com/spec.ccfc.min/firecracker-ci/v1.9/aarch64/vmlinux-6.1.102 -o /tmp/vmlinux-aarch64 aws s3 cp /tmp/vmlinux-aarch64 s3://burstgrid-123456789012/vmlinux-aarch64 3. Set GitHub credentials (choose one): GitHub App (recommended): aws ssm put-parameter --name /burstgrid/github-app-private-key --type SecureString --value "$(cat app.pem)" Personal access token: aws ssm put-parameter --name /burstgrid/github-token --type SecureString --value <token> 4. Deploy infrastructure: npx burstgrid deploy 5. Populate burstgrid.config.yaml with the new launch template IDs: npx burstgrid init 6. Register GitHub webhook: GitHub repo → Settings → Webhooks → Add webhook Payload URL: http://<scheduler-ip>:8080/webhook Content type: application/json Secret: <value in SSM /burstgrid/webhook-secret> Events: select "Workflow jobs" (workflow_job)
Step 2 — build, upload, and provision
$ npx burstgrid deploy [deploy] bucket=burstgrid-123456789012 region=us-east-1 [deploy] Building… [deploy] upload dist/scheduler.mjs → s3://burstgrid-123456789012/scheduler.mjs [deploy] upload dist/worker-agent.mjs → s3://burstgrid-123456789012/worker-agent.mjs [deploy] Running terraform apply in deploy/terraform… [deploy] Done. # Options $ npx burstgrid deploy --bucket my-bucket --no-terraform --dry-run
3
Workers boot automatically
Each worker EC2 instance runs the userdata bootstrap script injected by Terraform. On first boot it installs Node, downloads worker-agent.mjs from S3, and starts the agent as a systemd service. The agent connects to the scheduler over SSE, registers its available vCPUs and memory, and begins accepting jobs. No SSH, no manual setup on each host — scale up by increasing max_workers in your fleet config. For Firecracker mode (microVM isolation), you additionally need a Linux kernel binary and a rootfs ext4 image on each host. See the Custom images section for how to build them.
Worker agent — key environment variables
BURSTGRID_SCHEDULER_URL=http://10.0.1.50:8080 BURSTGRID_WORKER_TOKEN=your-worker-token BURSTGRID_MODE=firecracker # firecracker | process (GPU) | simulate (local dev) BURSTGRID_VM_IMAGE=/opt/images/runner.img # Firecracker mode only BURSTGRID_KERNEL=/opt/images/vmlinux # Firecracker mode only BURSTGRID_SLOTS=8 # concurrent VMs / processes per host
Simulate mode — try the full loop without EC2 or KVM

Set BURSTGRID_MODE=simulate on any machine to run the complete scheduler and worker agent loop in-process. No Firecracker binary, no KVM, no AWS account required. Simulate mode exercises the full path: webhook → queue → router → worker dispatch → job execution (in a child process). Per-repo concurrency limits, the S3 cache protocol, and all routing logic work identically. No config file required — every setting can be provided as an environment variable. When you’re ready to go to production, run burstgrid deploy and nothing else changes.

Minimal env vars for simulate mode (no YAML config file needed)
BURSTGRID_MODE=simulate BURSTGRID_WEBHOOK_SECRET=dev-secret BURSTGRID_WORKER_TOKEN=dev-token GITHUB_TOKEN=ghp_xxx # or GITHUB_APP_ID + GITHUB_PRIVATE_KEY_PATH BURSTGRID_REDIS_URL=redis://localhost:6379 # optional — omit for in-memory queue BURSTGRID_REPO_CONCURRENCY=10 # per-repo concurrency cap, applies in simulate too $ npx burstgrid scheduler # terminal 1 $ npx burstgrid worker # terminal 2 (or more)
4
Point your workflows at BurstGrid
Add self-hosted and a burstgrid:size= label to any job’s runs-on. That’s the only change needed in your workflow files. The scheduler reads the label, picks the right worker fleet, and boots a microVM sized to match. GPU jobs add gpu and land on workers that advertise that capability. No routing config required on your end.
.github/workflows/ci.yml
jobs: test: runs-on: [self-hosted, linux, burstgrid:size=large] steps: - uses: actions/checkout@v4 - run: pnpm test train: runs-on: [self-hosted, linux, gpu, burstgrid:size=4xlarge] steps: - run: python train.py

Configuration

Environment variables and config, without guessing

Start with npx burstgrid setup, then use npx burstgrid doctor before a real run. Terraform fills most production variables for you; local and custom deployments can use the reference below.

Recommended operator flow

1. npx burstgrid setup — writes Terraform vars and secrets
2. Build/upload rootfs + kernel artifacts
3. npx burstgrid deploy — uploads binaries and applies infra
4. npx burstgrid init — writes launch template IDs
5. npx burstgrid doctor — preflight safety and blast-radius checks

Legend

RequiredMust be provided unless the listed alternative is used.
AutoUsually generated by Terraform userdata or setup.
OptionalEnables an extra feature or overrides a default.
RecommendedStrongly recommended for production or realistic tests.

Scheduler environment

VariableStatusUsed byWhat it does
GITHUB_APP_ID + GITHUB_PRIVATE_KEY_PATHRequiredschedulerRecommended auth path for minting ephemeral runner registration tokens. Use GITHUB_PRIVATE_KEY instead of path if injecting PEM text directly, or use GITHUB_TOKEN for simpler single-repo tests.
GITHUB_TOKENAlternativeschedulerSimpler PAT-based auth for single-repo tests. Use instead of GitHub App credentials.
BURSTGRID_WEBHOOK_SECRETRecommendedschedulerValidates GitHub webhook HMAC. Empty is allowed for local dev only.
BURSTGRID_WORKER_TOKENRecommendedscheduler + workerShared bearer token for worker registration, heartbeat, stream, status, and evict routes. Empty disables worker auth for local dev.
BURSTGRID_FLEETSAutoschedulerJSON fleet config rendered by Terraform. Prefer autoscaler.fleets in YAML for local/manual config.
BURSTGRID_SPOT_QUEUE_URLAutoschedulerSQS queue receiving EC2 spot interruption warnings. The scheduler consumes it centrally and requeues jobs from the affected worker.
BURSTGRID_ADDR / BURSTGRID_PORTOptionalschedulerBind address and port. Defaults: 0.0.0.0 and 8080.
BURSTGRID_CONFIG / BURSTGRID_CONFIG_PATHOptionalall processesPath to burstgrid.config.yaml. Defaults to project-root burstgrid.config.yaml.
BURSTGRID_MAX_QUEUE_DEPTHOptionalschedulerQueue admission limit before returning 503 to GitHub so it retries webhook delivery.
BURSTGRID_WATCHED_REPOSOptionalschedulerComma-separated owner/repo list for periodic reconcile of queued GitHub jobs missed by webhook delivery.
BURSTGRID_RECONCILE_INTERVAL_MSOptionalschedulerReconciler interval. Defaults to 120 seconds.
BURSTGRID_AUTOSCALEROptionalschedulertrue|false|1|0 override for autoscaler.enabled.

Worker environment

VariableStatusModeWhat it does
BURSTGRID_SCHEDULER_URLRequiredallScheduler URL the worker registers against.
BURSTGRID_WORKER_IDAutoallWorker identity. Defaults to EC2 instance ID when metadata is available.
BURSTGRID_WORKER_TOKENRecommendedallBearer token sent to scheduler. Must match scheduler BURSTGRID_WORKER_TOKEN.
BURSTGRID_MODEOptionalallfirecracker (default), process for GPU/bare-metal, or simulate for local dev.
BURSTGRID_SLOTSAutoallMax concurrent jobs on this worker. Terraform sets this from fleet slots_per_worker; local default is half CPU count.
BURSTGRID_VCPUS / BURSTGRID_MEMORY_MIBOptionalallAdvertised worker capacity. Defaults to host CPU and memory.
BURSTGRID_CAPABILITIESOptionalallComma-separated labels the worker can serve, e.g. linux,arm64,docker,gpu. Defaults are auto-detected.
BURSTGRID_VM_IMAGE / BURSTGRID_KERNELRequiredfirecrackerRootfs image and kernel path for microVM boot. Terraform userdata defaults to /var/lib/burstgrid/rootfs.img and /var/lib/burstgrid/vmlinux.
BURSTGRID_IMAGE_DIROptionalfirecrackerDirectory for burstgrid:image=<name> rootfs resolution.
BURSTGRID_SECRET_DELIVERYRecommendedfirecrackermmds (default) keeps runner/cache tokens out of guest /proc/cmdline. cmdline is emergency rollback for old rootfs images.
BURSTGRID_USE_JAILERRecommendedfirecrackerRuns Firecracker through jailer chroot + uid/gid drop when host is prepared for it.
BURSTGRID_SNAPSHOT_POOL_SIZEOptionalfirecrackerPre-warmed snapshot count per worker for faster first dispatch.
BURSTGRID_RUNNER_PATHOptionalprocessRunner script path for process/GPU mode. Defaults to ./run.sh.
BURSTGRID_SSH_PUBLIC_KEYOptionalfirecrackerDebug SSH public key injected into guests only when rootfs has sshd. No host port is opened.
BURSTGRID_HEALTH_PORTOptionalallWorker health endpoint port. Default 9090.

Shared optional services

VariableStatusSet onWhat it does
OTEL_EXPORTER_OTLP_ENDPOINTRecommendedscheduler + workerEnables OTel metrics, traces, app logs, microVM console logs, and per-VM CPU/memory. Point to collector HTTP endpoint, e.g. http://localhost:4318.
BURSTGRID_S3_CACHE_BUCKET / REGIONOptionalworkerStarts the S3-backed Actions cache server and injects cache env into VM guests.
BURSTGRID_REGISTRY_MIRROROptionalworkerDocker registry mirror passed into Firecracker guests.
BURSTGRID_REDIS_URLOptionalschedulerRedis-backed durable queue and worker registry.
BURSTGRID_SQS_QUEUE_URL / REGIONOptionalschedulerOptional SQS job queue backend. Different from the spot-interruption queue Terraform wires as BURSTGRID_SPOT_QUEUE_URL.
BURSTGRID_DYNAMODB_TABLE / REGIONOptionalschedulerAppend-only job event history for queued/dispatched/running/completed/failed transitions.
BURSTGRID_S3_BUCKETOptionalCLIArtifact bucket fallback for build, deploy, and bake-ami commands.

Runtime safety defaults

worker.secretDelivery: mmds is the default. It keeps runner tokens, cache tokens, repo URLs, registry mirrors, and debug SSH keys out of the guest kernel command line. The guest init fetches them from Firecracker MMDS after networking is up.

worker: secretDelivery: mmds snapshotPool: size: 1

Spot interruption handling

The scheduler consumes the EC2 spot SQS queue centrally and reacts to two signals, not one. A rebalance recommendation (soft, earlier, no guaranteed follow-up) cordons the worker — stop placing new jobs there, leave running jobs alone — and immediately asks the autoscaler for replacement capacity instead of waiting on a timer. The hard interruption warning (~2 min notice) still drains that worker's tracked jobs now and requeues them. maxPackUtilization/maxActiveJobsPerWorker are static density caps on top of that, not a substitute for it.

scheduler: maxPackUtilization: 0.7 maxActiveJobsPerWorker: 8 autoscaler: fleets: - name: critical capacityType: on-demand

Scheduler availability (HA)

Default is a single EC2 instance with a directly-associated EIP — if it dies, webhooks fail until someone re-associates the EIP. Set scheduler_ha_enabled = true to put an ALB (stable DNS name) in front of a self-healing ASG (desired=1) instead. The ASG relaunches the scheduler automatically on an EC2 status-check or ALB /health/ready failure — no manual terraform apply or EIP reassociation. Covers failure recovery, not zero-downtime rolling deploys.

scheduler_ha_enabled = true scheduler_subnet_ids = ["subnet-aaaa", "subnet-bbbb"] # 2+ AZs, required by ALB

BurstGrid does not checkpoint an arbitrary running shell process mid-step. Interrupted spot jobs rerun from the workflow's last durable boundary. Make reruns cheap with actions/cache, artifacts at phase boundaries, and smaller dependent jobs instead of one giant step.

Optional feature examples

Each feature below activates with one environment variable or a small YAML block.

Shape matrix (size × family)

burstgrid:size= picks the vCPU tier; burstgrid:family= picks an independent memory multiplier — compute (0.5×), general (1×, default), or memory (2×). Combine both for a full shape matrix without a dedicated size entry per combination.

runs-on: [self-hosted, burstgrid:size=xlarge, burstgrid:family=memory] # 8 vCPU / 16 GiB instead of the default 8 vCPU / 8 GiB

GitHub Actions cache (S3)

BurstGrid serves the GitHub Actions cache protocol directly over S3. Set the bucket and actions/cache works in every job without any workflow changes. Cache keys are scoped per repo.

BURSTGRID_S3_CACHE_BUCKET=my-cache-bucket BURSTGRID_S3_CACHE_REGION=us-east-1

Per-repo concurrency limits

Cap how many jobs a single repo runs at once. Org wildcards apply fleet-wide; repo-specific entries take priority. Jobs over the cap stay queued — not dropped.

BURSTGRID_REPO_CONCURRENCY=10 # default for any repo # — or in burstgrid.config.yaml — scheduler: concurrencyLimits: myorg/*: 20 myorg/monorepo: 5

DynamoDB deduplication

Store job state in DynamoDB so duplicate webhook deliveries are silently dropped and already-running jobs survive a scheduler restart. The table is created by the Terraform module.

BURSTGRID_DYNAMODB_TABLE=burstgrid-jobs BURSTGRID_DYNAMODB_REGION=us-east-1

Redis queue

Replace the in-memory queue with Redis. Required if you run more than one scheduler replica or want queued jobs to survive process restarts. ElastiCache or any Redis-compatible endpoint works.

BURSTGRID_REDIS_URL=redis://your-elasticache:6379

SQS queue

Use SQS as the durable queue backend — no Redis needed. Jobs survive restarts and can be retried on failure. Useful when you want AWS-native durability without a separate cache cluster.

BURSTGRID_SQS_QUEUE_URL=https://sqs.us-east-1.amazonaws.com/123456/burstgrid BURSTGRID_SQS_REGION=us-east-1

Snapshot pool (fast boot)

Pre-boot N Firecracker VMs and hold them as memory snapshots so the first job on a fresh worker starts instantly instead of waiting 150 ms. Set on workers, not the scheduler. The pool refills in the background after each use.

BURSTGRID_SNAPSHOT_POOL_SIZE=4 # set on the worker agent

OpenTelemetry collector

The bundled collector (deploy/otel-collector/collector.yaml) is wired into both the scheduler and workers but off by default, since it needs exporter credentials. Store those as a multi-line SSM SecureString, then flip one Terraform variable — both roles run otelcol-contrib locally and export to it over loopback, no credentials in user data.

$ aws ssm put-parameter --name /burstgrid/otel-collector-env --type SecureString \ --value $'GRAFANA_OTLP_ENDPOINT=...\nGRAFANA_INSTANCE_ID=123\nGRAFANA_API_KEY=...' otel_collector_enabled = true # terraform.tfvars

Scale-down

Idle workers terminate automatically after 300 s. One warm standby is kept per fleet by default to avoid a cold-start on the next job landing there. Disable the warm standby if you'd rather scale all the way to zero between bursts.

scaleDownAfterIdleSec: 0 # per-fleet — disables the warm standby

Knowing when the scheduler itself is down

Most alerts in deploy/grafana/alerts.yaml describe symptoms (queue backed up, no capable workers) that assume the scheduler process is alive and exporting metrics. Two alerts specifically watch the scheduler's own health: BurstGridSchedulerDown fires on absent_over_time() of a core gauge — a dead process emits nothing, so absence is the only signal. BurstGridSchedulerCrashLooping fires on repeated process starts in a short window, catching a crash loop fast enough to dodge the absence check. Both still depend on the scheduler's own OTel export path working at all — add an external uptime check (Synthetic Monitoring, UptimeRobot, or a CloudWatch alarm on the HA mode's ALB target group) against GET /health for a fully independent signal.


Custom images

Your Dockerfile is your image spec

Every rootfs is built from a Dockerfile you own. Copy the template, add your tools, build once, reference by name. The image catalog in config maps label names to paths with no hardcoded conventions.

1
Start from the template
rootfs/template/Dockerfile is a commented starting point. Every section (base OS, system packages, Docker daemon, language runtimes, pre-warmed caches, runner binary) is labelled and optional. Copy it and delete what you don’t need.
rootfs/my-image/Dockerfile (after copying template)
FROM ubuntu:22.04 RUN apt-get update && apt-get install -y \ ca-certificates curl git libicu70 libssl3 \ python3 python3-pip build-essential # your additions # pre-warm pip cache so jobs don’t wait for installs RUN pip install pytest boto3 requests # runner binary (required) ARG RUNNER_VERSION=2.317.0 RUN mkdir -p /opt/actions-runner && cd /opt/actions-runner \ && curl -fsSL "https://github.com/actions/runner/releases/download/v${RUNNER_VERSION}/actions-runner-linux-x64-${RUNNER_VERSION}.tar.gz" \ | tar -xz && ./bin/installdependencies.sh
2
Build the .img
burstgrid build wraps scripts/build-rootfs.sh — it runs docker export + mkfs.ext4 to turn any Dockerfile into a bootable ext4 image, injects scripts/vm-init.sh as the PID 1 init script, and optionally uploads to S3. The Dockerfile can live anywhere in your repo; rootfs/my-image/ is just the convention used in these examples. The image name is derived from the directory name. Build output goes to $TMPDIR/burstgrid-images/<name>.img by default — use --out to override. Pass --push to upload to S3 immediately; the bucket auto-detects from terraform.tfvars or BURSTGRID_S3_BUCKET.
Build + upload to S3 in one command
$ npx burstgrid build rootfs/my-image/Dockerfile --size 4G --push [build] image=my-image size=4G bucket=burstgrid-123456789012 prefix=rootfs/ [build] Building Docker image from rootfs/my-image/Dockerfile… [build] Exporting filesystem… [build] Creating 4G ext4 image… [build] done → /tmp/burstgrid-images/my-image.img (1.2G on disk) [build] upload my-image.img → s3://burstgrid-123456789012/rootfs/my-image.img [build] Done. # then pull to each worker: $ ssh worker-01 "aws s3 cp s3://burstgrid-123456789012/rootfs/my-image.img /opt/images/"
3
Register and use
Add the image to the worker.images catalog in burstgrid.config.yaml. The name field is what jobs put in their runs-on label. The path is the local path on each worker host where the .img file lives after being pulled from S3 — it can be anything, /opt/images/ is just the convention used in these examples. All other fields are optional metadata.
burstgrid.config.yaml
worker: images: - name: my-image # used in runs-on path: /opt/images/my-image.img description: "Ubuntu 22.04 + Python 3 + pytest" os: ubuntu-22.04 tools: [python3, pip, git, curl]
.github/workflows/ci.yml
jobs: test: runs-on: [self-hosted, burstgrid:image=my-image]

GPU / AI

GPU jobs run directly on the host

Firecracker has no GPU passthrough. GPU and ML jobs are routed to dedicated EC2 GPU instances that run the Actions runner natively, with no microVM in the path. The host AMI is pre-baked with CUDA, ML frameworks, and optionally Docker + NVIDIA Container Toolkit.

1
Define a GPU AMI profile
Build an EC2 AMI with your GPU stack pre-installed (CUDA drivers, PyTorch, model weights, etc.). Register it in burstgrid.config.yaml under gpuAmis. All config keys use camelCase — snake_case keys will fail Zod validation at startup. Docker and the NVIDIA Container Toolkit can be included. docker run --gpus all works natively on the host.
burstgrid.config.yaml
gpuAmis: - name: my-gpu-profile # referenced in runs-on label amiId: ami-xxxxxxxxxxxxxxxxx # your pre-baked GPU AMI region: us-east-1 cudaVersion: "12.x" # informational only instanceFamilies: [g4dn, g5, p3] dockerEnabled: true cachedPackages: [torch, transformers] # pip packages pre-installed in the AMI cachedModels: [org/model-name] # HuggingFace model IDs pre-downloaded env: TORCH_HOME: /opt/model-cache
2
Configure the GPU autoscaler fleet
Add a fleet entry with gpuAmiId and instanceType — these override the Launch Template's default AMI and instance type for this fleet. Set capacityType: spot to use EC2 Spot (significant cost savings for ML workloads). Workers on these instances should register with BURSTGRID_MODE=process and a gpu capability; only GPU-labelled jobs will land on them.
autoscaler fleet + workflow label
# burstgrid.config.yaml — all keys are camelCase autoscaler: fleets: - name: gpu-ai sizeTag: burstgrid:gpu launchTemplateId: lt-gpu-base gpuAmiId: ami-0d3e4f5a6b7c8d9e0 instanceType: g4dn.xlarge capacityType: spot # spot | on-demand maxWorkers: 5 slotsPerWorker: 1 # 1 job at a time on each GPU host scaleUpThreshold: 0 subnetIds: [subnet-gpu-az1] # workflow label that routes to this fleet runs-on: [self-hosted, burstgrid:gpu]

Why BurstGrid

Container isolation isn't enough

When jobs handle secrets, run privileged operations, or modify the Docker daemon, sharing a kernel is a liability. VM-per-job removes the class of problem entirely.

A real kernel per job

Not a container — a Firecracker microVM with its own kernel, booted in under 200 ms and destroyed on exit. Secrets, file descriptors, and kernel state vanish when the job does. No shared surface between jobs.

150 ms to running, not 90 seconds

EC2 cold starts take 60–90 seconds: provision, OS boot, runner registration. A microVM takes 150 ms. On short, frequent jobs that gap is most of the CI wait time — not overhead.

Dense scheduling on bare metal

One host runs 20+ microVMs simultaneously with exact size matching and zero overcommit. Above roughly 15 concurrent jobs, that density makes bare metal cheaper per job than a fresh EC2 instance each time.

One binary. No orchestration layer.

The scheduler is a single process: webhook in, SSE push to workers, done. No Lambda functions to package, no SQS queues to tune, no Terraform state to maintain.