Stop development instances at night
This guide builds a schedule that stops your development EC2 instances every weekday evening and starts them again the next weekday morning. EC2 bills a running instance by the second, so a development instance that sits idle overnight and at weekends costs money for nothing. A stopped instance keeps everything on its EBS volumes.
Half an hour before the instances stop, the workflow posts a warning in Slack. If someone is still working, they react to the warning with 📌 (:pushpin:), and the instances keep running that night.
The work is split across these workflows:
| Workflow | Trigger | What it does |
|---|---|---|
| Stop development instances at night | Schedule, weekdays at 18:30 | Posts the warning, waits 30 minutes, and stops the running development instances, unless someone reacted to the warning. |
| Keep development instances running tonight | Slack Reaction added | Records that someone reacted to tonight’s warning. |
| Start development instances in the morning | Schedule, weekdays at 08:00 | Starts the instances that the evening workflow stopped. |
The workflows pass information to each other through Storage keys with Scope set to This workspace:
| Key | Written by | Read by |
|---|---|---|
dev_warning:<message timestamp> | The evening workflow, when it posts the warning | The reaction workflow, to check that a reaction is on tonight’s warning |
keep_dev_running | The reaction workflow | The evening workflow, before it stops anything |
dev_stopped_instances | The evening workflow | The morning workflow |
Before you start
Section titled “Before you start”- You need the Developer role in a KloudMate workspace whose plan includes workflows.
- You need an AWS connection with the Workflows capability, in the account that runs the instances.
- You need a Slack connection with the Workflows capability, and a channel for the warnings. Invite the KloudMate bot to that channel.
- Tag each development instance with the key
Environmentand the valuedevelopment. The workflows don’t touch instances without this tag.
Step 1: Let the AWS connection stop and start the instances
Section titled “Step 1: Let the AWS connection stop and start the instances”Add this policy to the IAM role that your AWS connection uses. Replace ACCOUNT_ID with your AWS account ID:
The condition limits stopping and starting to instances tagged as development, so the connection can’t stop a production instance even if a step asks it to.
If the instances’ EBS volumes are encrypted with a customer managed KMS key, the role also needs kms:CreateGrant on that key. Without it, an instance goes back to Stopped a few seconds after the morning workflow starts it.
Step 2: Import the workflows
Section titled “Step 2: Import the workflows”Import each workflow on its own. For each one, open Workflows, click Import, paste the YAML, click Import, and then click Open workflow to set what the import couldn’t.
Stop development instances at night
Section titled “Stop development instances at night”After you import it:
- On Find running development instances and Stop the instances, pick your AWS connection.
- On the Slack steps, pick your Slack connection. On Warn the channel, also pick the channel.
- On the trigger, set Timezone to your team’s timezone, such as
America/New_York.
Keep development instances running tonight
Section titled “Keep development instances running tonight”After you import it:
- On the trigger, pick your Slack connection under Account, and set Only in channel to the channel that gets the warnings.
- On Confirm in the thread, pick your Slack connection.
Start development instances in the morning
Section titled “Start development instances in the morning”After you import it:
- On Find the ones still stopped and Start the instances, pick your AWS connection.
- On Say they started, pick your Slack connection and the channel.
- On the trigger, set Timezone to the same timezone as the evening workflow.
All three workflows find the instances in the AWS connection’s default region. For another region, set Region on every AWS step.
How the workflows work
Section titled “How the workflows work”Find the running instances
Section titled “Find the running instances”An AWS API Call step calls EC2 → DescribeInstances with two filters, so it returns only running instances tagged as development:
EC2 returns instances grouped into reservations. A Transform step collects the instance IDs from every reservation into one list, with JSONPath set to $.Reservations[*].Instances[*].InstanceId. A Branch then checks that the list is not empty, so on an evening when nothing is running, the workflow stops there without posting anything.
Warn, then wait
Section titled “Warn, then wait”Warn the channel posts the list of instance IDs and asks anyone who needs them to react with 📌.
Slack identifies a message by its timestamp, which the post returns as ts. A Storage: Put step saves a key that contains it, dev_warning:{{ steps.warn.output.ts }}, with Expires after set to 30m. The reaction workflow looks for this key to tell tonight’s warning apart from any other message, and the expiry stops a late reaction from counting after the instances have stopped.
A Wait / Delay step then waits 30 minutes.
Keep them running tonight
Section titled “Keep them running tonight”The reaction workflow starts on Slack’s Reaction added event, with Only this emoji set to pushpin. The event carries the reacted message’s timestamp as trigger.event.item.ts, so a Storage: Get step looks up dev_warning:{{ trigger.event.item.ts }}. The key exists only when the reaction is on tonight’s warning, and only for 30 minutes after it was posted.
When the key exists, a Storage: Put step saves keep_dev_running, holding the person who reacted:
Expires after is 12h, so the key is gone by the next evening. A reply in the warning’s thread confirms the request.
Stop the instances
Section titled “Stop the instances”After the wait, the evening workflow reads keep_dev_running. If the key exists, it replies in the warning’s thread and names who asked, using Slack’s mention syntax, <@{{ steps.keep.output.value.user }}>.
Otherwise, it saves the list of instance IDs to dev_stopped_instances for the morning, and calls EC2 → StopInstances with InstanceIds set to {{ steps.ids.output.value }}. A field that holds a single reference receives the list itself, not text. See How values render.
The list expires after 4 days, so the list from Friday evening is still there on Monday morning.
Start them in the morning
Section titled “Start them in the morning”The morning workflow reads dev_stopped_instances. Before it starts anything, it calls DescribeInstances again, filtered by those instance IDs and by the state stopped:
This skips an instance that someone started by hand or terminated overnight. StartInstances fails for the whole list if any instance in it no longer exists, so checking first keeps one missing instance from stopping the rest.
It then starts the instances, posts which ones started, and removes dev_stopped_instances.
Step 3: Test and publish
Section titled “Step 3: Test and publish”- In Stop development instances at night, test Find running development instances and List their instance IDs from their Test tabs. Neither one changes anything, and the output shows which instances the workflow would stop.
- Publish all three workflows. The first publish also switches each one on.
- On the first evening, check the warning in Slack, and react to it with 📌 to try the reaction workflow. The instances then keep running that night.
To follow what happened, open Workflows → Runs. See Run history.
Adapt the workflow
Section titled “Adapt the workflow”- Use your own tag. Change the tag filter on Find running development instances and the condition in the IAM policy, for example to
Stage=dev. - Change the hours. Set the time on each trigger. To keep the instances off at weekends, leave Run on weekends off on both schedules.
- Give more notice. Change Wait 30 minutes, and set Expires after on Remember the warning message to the same duration.
- Stop databases too. Add AWS API Call steps for RDS → StopDBInstance and StartDBInstance, and add both actions to the IAM policy. AWS starts a stopped RDS instance again after 7 days, which a nightly schedule never reaches.
Related
Section titled “Related”- Storage for workspace-scoped keys and their expiry.
- Triggers for Slack events such as Reaction added.
- Schedule for time zones and weekday schedules.
- AWS API Call for the services a step can call.