Run hundreds of concurrent GitHub Actions jobs on EC2 you control, billed at spot rates — not GitHub’s per-minute price. Each job gets a fresh Firecracker microVM: dedicated kernel, clean filesystem, boots in 150 ms, gone when it’s done. No shared state. No secret bleed.
Try the full scheduler and worker loop on any machine with BURSTGRID_MODE=simulate — no EC2, no KVM needed.
BurstGrid is not the right tool for everyone. Here is when it makes sense and when it does not.
No Kubernetes. No job-path message broker. Just a webhook, a scheduler queue, a persistent SSE connection, and a VM that exists for exactly as long as the job does.
workflow_job webhookworkflow_job event to your scheduler.
BurstGrid verifies the X-Hub-Signature-256 HMAC before doing anything else.
If the signature doesn’t match, the request is dropped. If the queue is full, it returns 503;
GitHub will retry the delivery for up to 72 hours, so nothing is lost.
critical jobs always jump ahead of standard ones.
Token generation is protected by a circuit breaker: after 5 consecutive GitHub API failures,
the scheduler opens the circuit and returns 503 until the API recovers.
burstgrid:size=large → 4 vCPU / 4 GiB) against every connected
worker’s free resources and capability labels. The first worker that fits gets the assignment pushed
as a JSON event over its persistent SSE connection, the same open HTTP response it’s been holding
since it registered. No database read. No poll. Dispatch latency is typically under 5 ms.
/proc/cmdline. The init script
inside the VM reads MMDS, calls ./config.sh --ephemeral to register with GitHub,
then ./run.sh to execute the job. When it exits, the VM is killed, the disk is discarded,
and the slot is freed. The rootfs is never reused.
Four steps. Steps 1–2 take minutes. Step 3 handles building, uploading, and provisioning in one command.
https://your-scheduler/webhook/github and generate a secret.
Grant Administration: read & write and Actions: read, then subscribe to workflow_job events.
Download the private key PEM and install the App on your org.
burstgrid setup, then burstgrid deploy, then burstgrid initburstgrid setup detects your default VPC, public subnet, and the latest Ubuntu 24.04 AMI automatically,
generates webhook and worker secrets in AWS Systems Manager Parameter Store (SSM), and writes
deploy/terraform/terraform.tfvars for you.
Then burstgrid deploy builds the binaries, uploads them to S3, and calls terraform apply.
Finally, burstgrid init reads the deployed EC2 Launch Templates and writes their IDs back into
burstgrid.config.yaml — the autoscaler won’t launch workers without these.
No digging around the AWS console.
/burstgrid/webhook-secret and
/burstgrid/worker-token with their instance IAM roles after boot. This keeps
secrets out of Terraform state and EC2 user data. Add the GitHub credential you use as
/burstgrid/github-app-private-key or /burstgrid/github-token.
Use npx burstgrid setup --secret-source terraform only for the legacy path,
where secrets are written to terraform.tfvars and embedded in launch data.
worker-agent.mjs from S3, and starts the agent as a systemd service. The agent connects to
the scheduler over SSE, registers its available vCPUs and memory, and begins accepting jobs. No SSH, no manual
setup on each host — scale up by increasing max_workers in your fleet config.
For Firecracker mode (microVM isolation), you additionally need a Linux kernel binary and a rootfs
ext4 image on each host. See the Custom images section for how to build them.
Set BURSTGRID_MODE=simulate on any machine to run the complete scheduler and worker agent loop in-process.
No Firecracker binary, no KVM, no AWS account required. Simulate mode exercises the full path: webhook → queue →
router → worker dispatch → job execution (in a child process). Per-repo concurrency limits, the S3 cache
protocol, and all routing logic work identically. No config file required — every setting can be provided as an
environment variable. When you’re ready to go to production, run burstgrid deploy and nothing else changes.
self-hosted and a burstgrid:size= label to any job’s runs-on.
That’s the only change needed in your workflow files. The scheduler reads the label, picks the right worker fleet,
and boots a microVM sized to match. GPU jobs add gpu and land on workers that advertise that capability.
No routing config required on your end.
Start with npx burstgrid setup, then use npx burstgrid doctor before a real run. Terraform fills most production variables for you; local and custom deployments can use the reference below.
npx burstgrid setup — writes Terraform vars and secretsnpx burstgrid deploy — uploads binaries and applies infranpx burstgrid init — writes launch template IDsnpx burstgrid doctor — preflight safety and blast-radius checks
| Required | Must be provided unless the listed alternative is used. |
| Auto | Usually generated by Terraform userdata or setup. |
| Optional | Enables an extra feature or overrides a default. |
| Recommended | Strongly recommended for production or realistic tests. |
| Variable | Status | Used by | What it does |
|---|---|---|---|
| GITHUB_APP_ID + GITHUB_PRIVATE_KEY_PATH | Required | scheduler | Recommended auth path for minting ephemeral runner registration tokens. Use GITHUB_PRIVATE_KEY instead of path if injecting PEM text directly, or use GITHUB_TOKEN for simpler single-repo tests. |
| GITHUB_TOKEN | Alternative | scheduler | Simpler PAT-based auth for single-repo tests. Use instead of GitHub App credentials. |
| BURSTGRID_WEBHOOK_SECRET | Recommended | scheduler | Validates GitHub webhook HMAC. Empty is allowed for local dev only. |
| BURSTGRID_WORKER_TOKEN | Recommended | scheduler + worker | Shared bearer token for worker registration, heartbeat, stream, status, and evict routes. Empty disables worker auth for local dev. |
| BURSTGRID_FLEETS | Auto | scheduler | JSON fleet config rendered by Terraform. Prefer autoscaler.fleets in YAML for local/manual config. |
| BURSTGRID_SPOT_QUEUE_URL | Auto | scheduler | SQS queue receiving EC2 spot interruption warnings. The scheduler consumes it centrally and requeues jobs from the affected worker. |
| BURSTGRID_ADDR / BURSTGRID_PORT | Optional | scheduler | Bind address and port. Defaults: 0.0.0.0 and 8080. |
| BURSTGRID_CONFIG / BURSTGRID_CONFIG_PATH | Optional | all processes | Path to burstgrid.config.yaml. Defaults to project-root burstgrid.config.yaml. |
| BURSTGRID_MAX_QUEUE_DEPTH | Optional | scheduler | Queue admission limit before returning 503 to GitHub so it retries webhook delivery. |
| BURSTGRID_WATCHED_REPOS | Optional | scheduler | Comma-separated owner/repo list for periodic reconcile of queued GitHub jobs missed by webhook delivery. |
| BURSTGRID_RECONCILE_INTERVAL_MS | Optional | scheduler | Reconciler interval. Defaults to 120 seconds. |
| BURSTGRID_AUTOSCALER | Optional | scheduler | true|false|1|0 override for autoscaler.enabled. |
| Variable | Status | Mode | What it does |
|---|---|---|---|
| BURSTGRID_SCHEDULER_URL | Required | all | Scheduler URL the worker registers against. |
| BURSTGRID_WORKER_ID | Auto | all | Worker identity. Defaults to EC2 instance ID when metadata is available. |
| BURSTGRID_WORKER_TOKEN | Recommended | all | Bearer token sent to scheduler. Must match scheduler BURSTGRID_WORKER_TOKEN. |
| BURSTGRID_MODE | Optional | all | firecracker (default), process for GPU/bare-metal, or simulate for local dev. |
| BURSTGRID_SLOTS | Auto | all | Max concurrent jobs on this worker. Terraform sets this from fleet slots_per_worker; local default is half CPU count. |
| BURSTGRID_VCPUS / BURSTGRID_MEMORY_MIB | Optional | all | Advertised worker capacity. Defaults to host CPU and memory. |
| BURSTGRID_CAPABILITIES | Optional | all | Comma-separated labels the worker can serve, e.g. linux,arm64,docker,gpu. Defaults are auto-detected. |
| BURSTGRID_VM_IMAGE / BURSTGRID_KERNEL | Required | firecracker | Rootfs image and kernel path for microVM boot. Terraform userdata defaults to /var/lib/burstgrid/rootfs.img and /var/lib/burstgrid/vmlinux. |
| BURSTGRID_IMAGE_DIR | Optional | firecracker | Directory for burstgrid:image=<name> rootfs resolution. |
| BURSTGRID_SECRET_DELIVERY | Recommended | firecracker | mmds (default) keeps runner/cache tokens out of guest /proc/cmdline. cmdline is emergency rollback for old rootfs images. |
| BURSTGRID_USE_JAILER | Recommended | firecracker | Runs Firecracker through jailer chroot + uid/gid drop when host is prepared for it. |
| BURSTGRID_SNAPSHOT_POOL_SIZE | Optional | firecracker | Pre-warmed snapshot count per worker for faster first dispatch. |
| BURSTGRID_RUNNER_PATH | Optional | process | Runner script path for process/GPU mode. Defaults to ./run.sh. |
| BURSTGRID_SSH_PUBLIC_KEY | Optional | firecracker | Debug SSH public key injected into guests only when rootfs has sshd. No host port is opened. |
| BURSTGRID_HEALTH_PORT | Optional | all | Worker health endpoint port. Default 9090. |
| Variable | Status | Set on | What it does |
|---|---|---|---|
| OTEL_EXPORTER_OTLP_ENDPOINT | Recommended | scheduler + worker | Enables OTel metrics, traces, app logs, microVM console logs, and per-VM CPU/memory. Point to collector HTTP endpoint, e.g. http://localhost:4318. |
| BURSTGRID_S3_CACHE_BUCKET / REGION | Optional | worker | Starts the S3-backed Actions cache server and injects cache env into VM guests. |
| BURSTGRID_REGISTRY_MIRROR | Optional | worker | Docker registry mirror passed into Firecracker guests. |
| BURSTGRID_REDIS_URL | Optional | scheduler | Redis-backed durable queue and worker registry. |
| BURSTGRID_SQS_QUEUE_URL / REGION | Optional | scheduler | Optional SQS job queue backend. Different from the spot-interruption queue Terraform wires as BURSTGRID_SPOT_QUEUE_URL. |
| BURSTGRID_DYNAMODB_TABLE / REGION | Optional | scheduler | Append-only job event history for queued/dispatched/running/completed/failed transitions. |
| BURSTGRID_S3_BUCKET | Optional | CLI | Artifact bucket fallback for build, deploy, and bake-ami commands. |
worker.secretDelivery: mmds is the default. It keeps runner tokens, cache tokens, repo URLs, registry mirrors, and debug SSH keys out of the guest kernel command line. The guest init fetches them from Firecracker MMDS after networking is up.
The scheduler consumes the EC2 spot SQS queue centrally and reacts to two signals, not one. A rebalance recommendation (soft, earlier, no guaranteed follow-up) cordons the worker — stop placing new jobs there, leave running jobs alone — and immediately asks the autoscaler for replacement capacity instead of waiting on a timer. The hard interruption warning (~2 min notice) still drains that worker's tracked jobs now and requeues them. maxPackUtilization/maxActiveJobsPerWorker are static density caps on top of that, not a substitute for it.
Default is a single EC2 instance with a directly-associated EIP — if it dies, webhooks fail until someone re-associates the EIP. Set scheduler_ha_enabled = true to put an ALB (stable DNS name) in front of a self-healing ASG (desired=1) instead. The ASG relaunches the scheduler automatically on an EC2 status-check or ALB /health/ready failure — no manual terraform apply or EIP reassociation. Covers failure recovery, not zero-downtime rolling deploys.
BurstGrid does not checkpoint an arbitrary running shell process mid-step. Interrupted spot jobs rerun from the workflow's last durable boundary. Make reruns cheap with actions/cache, artifacts at phase boundaries, and smaller dependent jobs instead of one giant step.
Each feature below activates with one environment variable or a small YAML block.
burstgrid:size= picks the vCPU tier; burstgrid:family= picks an independent memory multiplier — compute (0.5×), general (1×, default), or memory (2×). Combine both for a full shape matrix without a dedicated size entry per combination.
BurstGrid serves the GitHub Actions cache protocol directly over S3. Set the bucket and actions/cache works in every job without any workflow changes. Cache keys are scoped per repo.
Cap how many jobs a single repo runs at once. Org wildcards apply fleet-wide; repo-specific entries take priority. Jobs over the cap stay queued — not dropped.
Store job state in DynamoDB so duplicate webhook deliveries are silently dropped and already-running jobs survive a scheduler restart. The table is created by the Terraform module.
Replace the in-memory queue with Redis. Required if you run more than one scheduler replica or want queued jobs to survive process restarts. ElastiCache or any Redis-compatible endpoint works.
Use SQS as the durable queue backend — no Redis needed. Jobs survive restarts and can be retried on failure. Useful when you want AWS-native durability without a separate cache cluster.
Pre-boot N Firecracker VMs and hold them as memory snapshots so the first job on a fresh worker starts instantly instead of waiting 150 ms. Set on workers, not the scheduler. The pool refills in the background after each use.
The bundled collector (deploy/otel-collector/collector.yaml) is wired into both the scheduler and workers but off by default, since it needs exporter credentials. Store those as a multi-line SSM SecureString, then flip one Terraform variable — both roles run otelcol-contrib locally and export to it over loopback, no credentials in user data.
Idle workers terminate automatically after 300 s. One warm standby is kept per fleet by default to avoid a cold-start on the next job landing there. Disable the warm standby if you'd rather scale all the way to zero between bursts.
Most alerts in deploy/grafana/alerts.yaml describe symptoms (queue backed up, no capable workers) that assume the scheduler process is alive and exporting metrics. Two alerts specifically watch the scheduler's own health: BurstGridSchedulerDown fires on absent_over_time() of a core gauge — a dead process emits nothing, so absence is the only signal. BurstGridSchedulerCrashLooping fires on repeated process starts in a short window, catching a crash loop fast enough to dodge the absence check. Both still depend on the scheduler's own OTel export path working at all — add an external uptime check (Synthetic Monitoring, UptimeRobot, or a CloudWatch alarm on the HA mode's ALB target group) against GET /health for a fully independent signal.
Every rootfs is built from a Dockerfile you own. Copy the template, add your tools, build once, reference by name. The image catalog in config maps label names to paths with no hardcoded conventions.
rootfs/template/Dockerfile is a commented starting point. Every section (base OS, system packages,
Docker daemon, language runtimes, pre-warmed caches, runner binary) is labelled and optional.
Copy it and delete what you don’t need.
burstgrid build wraps scripts/build-rootfs.sh — it runs docker export
+ mkfs.ext4 to turn any Dockerfile into a bootable ext4 image, injects
scripts/vm-init.sh as the PID 1 init script, and optionally uploads to S3.
The Dockerfile can live anywhere in your repo; rootfs/my-image/ is just the convention
used in these examples. The image name is derived from the directory name.
Build output goes to $TMPDIR/burstgrid-images/<name>.img by default —
use --out to override. Pass --push to upload to S3 immediately;
the bucket auto-detects from terraform.tfvars or BURSTGRID_S3_BUCKET.
worker.images catalog in burstgrid.config.yaml.
The name field is what jobs put in their runs-on label.
The path is the local path on each worker host where the
.img file lives after being pulled from S3 — it can be anything, /opt/images/
is just the convention used in these examples. All other fields are optional metadata.
Firecracker has no GPU passthrough. GPU and ML jobs are routed to dedicated EC2 GPU instances that run the Actions runner natively, with no microVM in the path. The host AMI is pre-baked with CUDA, ML frameworks, and optionally Docker + NVIDIA Container Toolkit.
burstgrid.config.yaml under gpuAmis.
All config keys use camelCase — snake_case keys will fail Zod validation at startup.
Docker and the NVIDIA Container Toolkit can be included. docker run --gpus all works natively on the host.
gpuAmiId and instanceType — these override the Launch Template's default AMI and instance type for this fleet.
Set capacityType: spot to use EC2 Spot (significant cost savings for ML workloads).
Workers on these instances should register with BURSTGRID_MODE=process and a gpu capability;
only GPU-labelled jobs will land on them.
When jobs handle secrets, run privileged operations, or modify the Docker daemon, sharing a kernel is a liability. VM-per-job removes the class of problem entirely.
Not a container — a Firecracker microVM with its own kernel, booted in under 200 ms and destroyed on exit. Secrets, file descriptors, and kernel state vanish when the job does. No shared surface between jobs.
EC2 cold starts take 60–90 seconds: provision, OS boot, runner registration. A microVM takes 150 ms. On short, frequent jobs that gap is most of the CI wait time — not overhead.
One host runs 20+ microVMs simultaneously with exact size matching and zero overcommit. Above roughly 15 concurrent jobs, that density makes bare metal cheaper per job than a fresh EC2 instance each time.
The scheduler is a single process: webhook in, SSE push to workers, done. No Lambda functions to package, no SQS queues to tune, no Terraform state to maintain.