SwarmOps
Centralized command and control for drone fleets
Guy · Tony · Harel · Final Project 2026
GUY · COMBAT MEDIC, SEARCH & RESCUE
The Mission
The problem · The idea · Live demo
FROM THE FIELD
Every squad, its own drone, its own picture.
We each met that wall on duty, in three different jobs.
THE PROBLEM
Two squads, one sector, zero coordination
THE COST
Uncoordinated control doesn’t scale
More squads
→
More drones per sector
→
Overlap and gaps
→
Degraded missions
+15%
wasted flight time from overlapping and uncoordinated tasking.
THE IDEA
The system flies the fleet.
People command the mission.
1YOU ASK
2IT DECIDES
3THE FLEET FLIES
Jenin Surveillance 6
Priority 1
by 18:00
thermal
→
every drone weighed
against every constraint
→
DR-0258% left
DR-0544% left
DR-0937% left
One row per drone: launch point, the stops it was given, and the battery it lands with — in seconds.
THE SIMULATOR
A virtual drone, like the real thing
It flies, streams video and reacts in real time. No hardware needed.
FPV feed, straight from the virtual camera
Third-person view of the same flight
IN ACTION
A patrol that never blinks
Battery runs low. Another drone continues the patrol without missing a moment.
LIVE DEMO
The operating picture
Every drone, its battery and its mission. One map.
Watch a patrol hand off to a fresh drone, live.
TONY · SNIPER, DRONE OPERATOR
The Cloud
Architecture · Scale · Cost
ARCHITECTURE
Six services, one gateway
Small pieces instead of one big program. Each can fail and recover on its own.
gateway
the only public door, routes every request
auth
log-in, tokens and roles
fleet
the drone registry: state, battery, payload
mission
orders and their lifecycle, creation to completion
planning
the brain: assigns drones to missions, replans in flight
telemetry
live positions and battery, streamed off the bus
notify
pushes alerts to operators the moment they matter
RabbitMQ message bus
services talk through queues, never directly
drone simulator
virtual drones feeding the bus
AWS ACCOUNT
THE CLOUDWhere it actually runs
VPC›EKS CLUSTER
KUBERNETESInside the cluster
VPC›NODE 3›PLANNING-SERVICE
THE OPTIMIZERInside the planner
A LIVE SNAPSHOT
The cluster, by the numbers
5
spot nodes · 10 vCPU · 18.7 Gi
62
running containers
23
deployments
25 Gi
persistent storage · 6 gp3 volumes
9
service images in ECR
2
availability zones
GITHUB›ACTIONS · CI
CONTINUOUS INTEGRATIONHow an image gets built
COST
What all of this cost
Default configuration, per month
After tuning: spot nodes, one NAT, right-sizing
Actually paid: total, entire project
The whole environment shut down at the end of every workday. Infrastructure as code made rebuilding it a one-liner.
HAREL · TANK GUNNER
Operations
Deployment · Monitoring · Security
OPERATIONS
The cluster syncs itself
git
the desired state — the file CI just bumped
→
Argo CD diffs
desired vs. live, over and over
→
Sync
pulls the new tag, applies the chart
→
Live state
the cluster now matches git
↑
drift detected — Argo puts git’s version back, within minutes
Pull, not push. The cluster converges on git on its own. Nobody runs helm upgrade by hand — not once.
DEPLOYMENT
The deployer is a bot
12 synced · 27 healthy · last sync authored by swarmops-ci-bot, not a human.
DEPLOYMENT
A new version, and nobody notices
Blue-green: the old version keeps serving until the new one is ready.
During the rollout: old and new run side by side
Thirty seconds later: only the new version remains
DEPLOYMENT
20% first,
then everyone
The brain of the system, planning, ships as a canary, with an automated metric check between every step.
If a check fails, the rollout reverses itself. No pager, no 3am rollback by hand.
Set weight: 20%
Analysis · automated check
Set weight: 50%
Analysis · automated check
Set weight: 100%
DEPLOYMENT
A rollout, start to finish
Mid-rollout: two revisions live, traffic split
Finished: canary promoted, the old revision at zero
MONITORING
Know before anyone complains
Every service exports metrics. When something drifts, the system raises the alarm itself.
Grafana · Prometheus · Loki · the SwarmOps overview dashboard
MONITORING
Not just alive. Measured.
Requests, failures, and how long it takes to compute a flight plan.
Service targets up, right now
Container restarts, settling to zero
Plan solve time, p50 / p95
RESILIENCE
It breaks. It heals. Nobody gets paged.
Spot node reclaimedAWS takes a node back
→
Kubernetes reschedulespods land on the remaining nodes
→
Full capacity~2 minutes · no human
Canary metrics failnew version underperforms
→
Rollout reverses itselftraffic returns to the stable version
→
Old version servingautomatic · no 3am rollback
Manual change on a serverunauthorized drift
→
Argo CD detects driftcompares cluster to git
→
Git state restored~3 minutes · self-reverting
SECURITY
Secured like a base — outside in
PERIMETER
one public door, everything else private
IDENTITY
least privilege, per service
SECRETS
short-lived, issued at runtime
GIT
the standing order — drift is reverted
NEXT Closed military networks. The doctrine holds; only the perimeter moves.
SECURITY
One command, three checkpoints
Operatorsends a command
1encrypted
→
2Load balancerthe only public door
→
3gatewaychecks the token
✗ no valid token
→
Stopped at the door · 401 · never reaches a service
✓ valid token
→
PRIVATE NETWORK
planningits own limited role
→
droneexecutes
SwarmOps. In one slide.
One operator, a whole fleet
Plans in seconds
Replanning in flight
6 microservices
EKS on AWS
Private VPC · 2 zones
5 spot nodes
62 containers
Deployed by a bot
Blue-green + canary
Self-healing
Zero-trust network
Automatic rollback
Metrics on everything
$43 total spend
Thank you.
Questions welcome.
Guy · Tony · Harel · SwarmOps