extension-aws 2.4.33

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/eks
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticache
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/eventbridge
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/mq
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/rds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/sqs
  • chore(deps): bump github.com/aws/smithy-go from 1.28.1 to 1.28.2

extension-datadog 1.8.33

  • build(deps): bump github.com/DataDog/datadog-api-client-go/v2
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test

extension-gcp 1.0.42

  • chore(deps): bump cloud.google.com/go/compute from 1.70.0 to 1.71.0
  • chore(deps): bump cloud.google.com/go/container from 1.54.0 to 1.55.0
  • chore(deps): bump github.com/googleapis/gax-go/v2 from 2.24.1 to 2.26.2
  • chore(deps): bump google.golang.org/api from 0.299.0 to 0.300.0

extension-host 1.8.3

  • build(deps): bump github.com/beevik/ntp from 1.5.0 to 1.6.0
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test

extension-istio 1.0.35

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump istio.io/api from 1.31.0 to 1.31.1
  • chore(deps): bump istio.io/client-go from 1.31.0 to 1.31.1
  • chore(deps): bump k8s.io/client-go from 0.37.0 to 0.37.1

extension-jvm 1.3.6

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump k8s.io/client-go from 0.37.0 to 0.37.1

extension-k6 1.3.8

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump k8s.io/client-go from 0.37.0 to 0.37.1

Agent 2.4.7

Improvement

  • Experiment traces go only to your own tracing backend. The agent no longer uploads the spans of an experiment run to the Steadybit platform, which has stopped storing them — find a run's trace in your backend from the run's Tracing tab instead. Tracing is on when you set OTEL_JAVA_GLOBAL_AUTOCONFIGURE_ENABLED=true together with the standard OTEL_* variables such as OTEL_EXPORTER_OTLP_ENDPOINT, and off otherwise; every span of a run carries its experiment.execution.id. STEADYBIT_AGENT_SPANS_ENABLED (steadybit.agent.spans.enabled) is no longer used — to switch tracing off, leave auto-configuration disabled or set OTEL_SDK_DISABLED=true.
  • Extension-call bodies are no longer recorded on spans. STEADYBIT_AGENT_OPENTELEMETRY_BODY_SPAN_ENABLED (steadybit.agent.opentelemetry.body-span-enabled) and its older spelling STEADYBIT_AGENT_OPEN_TELEMETRY_BODY_SPAN_ENABLED are removed and ignored. Like standard OpenTelemetry instrumentation, the agent's spans carry the method, URL and status of each extension call, never the request or response body — which held action parameters, credentials among them, and discovered target data.
  • Lower memory peaks during discovery. The agent no longer keeps an extra full copy of every extension response while reading it. Discovery responses grow with the number of targets — tens of megabytes on large clusters — so each discovery cycle now needs noticeably less memory at its peak.
  • The agent's traces no longer carry its command line. When tracing is on, the Java SDK used to attach the agent's full JVM command line (process.command_args) to every span, where a secret passed as a -D flag would end up in your tracing backend. Following the OpenTelemetry conventions, which make the command line opt-in, it is now left out; process id, runtime and host stay. Set OTEL_RESOURCE_DISABLED_KEYS to choose the dropped keys yourself.

Fix

  • Linux packages trust additional CA certificates again. An agent installed from the .deb/.rpm package now imports the certificates from its extra-certificates directory (STEADYBIT_AGENT_EXTRA_CERTS_PATH) instead of failing silently and then rejecting a platform behind an internal CA with PKIX path building failed.
  • A step whose extension cannot be reached says so in plain words. A step whose extension was registered but not answering used to fail with the HTTP client's own wording — I/O error on POST request for "http://…": Connection refused — which says nothing about what to check. It now reports that the extension could not be reached, or did not answer in time, and that it is still registered with the agent, so the thing to check is whether it is running and healthy. The extension's own error message is still used whenever it answers with one.

Dependencies

  • Jackson updated to 3.1.7 and 2.21.7 to address CVE-2026-91776 and CVE-2026-91777.

extension-kubernetes 2.6.37

  • Cause Crash Loop attack: SIGKILL and SIGSTOP are no longer allowed as signal; experiments using them will now fail on start
  • Initialize OpenTelemetry tracing
  • Update dependencies

extension-redis 1.1.18

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • fix(chart): always render the shared extension env (#55)
  • fix: authenticate actions on Redis Cluster node targets (#57)
  • fix: keep cluster nodes of a plain endpoint on plain TCP (#58)

Platform 2.8.9

Fix

  • Environments no longer keep refusing runs while their targets are updated — Since 2.8.4, a target change that could not be applied to an environment's target list triggered a rebuild of every environment the target belongs to, and while rebuilding an environment refused new runs with "The environment is not yet ready. Please retry later." A rebuild also held up the regular target updates of that environment, so on installations with many targets and frequent changes these failures kept producing new rebuilds, and environments flipped to "Updating associated targets…" again and again. A rebuild now only touches the targets that actually joined or left the environment, so it no longer holds up target updates. Only an environment that may still list a target that has left it pauses runs until it is repaired, so no run can reach a target outside its scope; one that is merely missing a new target stays ready and keeps accepting runs while it is updated.

Security

  • Removing a user's access now also ends their open sessions after an OIDC login — When a user signed in through OIDC and the login was matched to an existing account, for example by email, removing their last team membership or downgrading them to the User role could leave their browser session active with the access they had before. Revoking access now ends every session belonging to that user, however they signed in.

Platform 2.8.8

New

  • Experiment templates and reliable editing through the MCP server — Administrators can now create, update and delete experiment templates from an AI assistant such as Claude, for example to turn a proof-of-value scope into a ready-to-use template for your teams. The template tools are marked as admin-only and follow the same rules as the template editor: only users with the Admin role can use them, and templates imported from the Reliability Hub stay read-only. Reading an experiment or template now returns its complete design in the same format the update tools accept, so an assistant can change an existing experiment or template without losing variables, properties, tags or checks it did not touch.

Improvement

  • Consistent optimistic locking in the API — Creating or updating an experiment through the API (POST /api/experiments, POST /api/experiments/{key}) and updating an experiment from its template (POST /api/experiments/templates/{id}/experiment-update/{key}) now accept an optional version, like the other API endpoints already do. Send the version you read to have an update rejected with 409 Conflict instead of silently overwriting a change someone else saved in the meantime; leave it out to keep today's behavior of overwriting the current version. The API specification now documents version the same way everywhere: always present in responses, optional in requests. Creating or updating a team together with its members (POST /api/teams) now returns the version that was saved. Before, it returned the version from before the member change, so sending that version with the next update was rejected as a conflict.
  • Standard Bearer authentication for the platform API — API access tokens can now be sent as Authorization: Bearer <token>, the format HTTP clients, API tools and generated SDKs support out of the box, as the MCP server already does. The existing Authorization: accessToken <token> format keeps working, so existing scripts and pipelines need no change. See Integrate via API.
  • Rancher RKE2 / K3s in the guided Kubernetes agent installation — The agent setup now lists Rancher (RKE2 / K3s) as a Kubernetes distribution. These clusters run an embedded containerd whose socket is not at the usual location, so the generated Helm command points the container extension at it for you, as described in the installation guide.
  • A SteadyBuddy conversation open in several places now keeps up with itself — Messages and answers appear in every window showing the conversation as they happen, instead of only in the one you typed in, and a window that shows a conversation SteadyBuddy is currently answering says so — including when you come back to it from another page while the answer is still being written. Previously such a window either sat on a stale transcript or had to keep asking the server for the answer, and gave up after two minutes.
  • Configurable team key extraction for OIDC and LDAP team sync — The team key derived from an IdP group can now be selected with a regular expression: STEADYBIT_AUTH_OAUTH2_TEAM_KEY_PATTERN for OIDC and STEADYBIT_AUTH_LDAP_SYNC_TEAM_KEY_PATTERN for LDAP. The first capture group becomes the team key, which is upper-cased, cleaned of special characters and cut to 16 characters as before; groups the pattern does not match are not synchronized. This lets you strip a common prefix (^steadybit-(.+)$) so that long group names no longer collide on their first 16 characters, or restrict synchronization to a subset of your groups. The default, ^(.{0,16}), keeps the existing behavior, so existing teams keep their keys. Changing the pattern on an installation that already synchronizes teams changes their keys: the next sync creates the teams anew under the new keys, and LDAP sync deletes the previously synchronized teams as stale, so plan the switch before teams have accumulated experiments. LDAP team key attributes longer than 16 characters are no longer skipped but cut to 16 characters like OIDC groups.

Fix

  • Clear validation errors for an experiment step without an action — Creating or updating an experiment or experiment template through the API or the MCP server with a targeted step that is missing its actionType failed with an internal validation error (HV000028 ... 'value' must not be null) that hid every other problem in the design. The step is now reported as missing its actionType, alongside any other validation errors. A step now also accepts actionId as another name for its actionType, and the MCP server's experiment design guide no longer names the field actionId.

  • Updating integrations through the API without a version — The version of webhooks (Slack, custom and preflight webhooks) and preflight action integrations is documented as optional, but an update that left it out was rejected with 409 Conflict as soon as the integration had been edited once. Leaving it out now updates the current version, as with every other API endpoint; sending a version still protects against overwriting a newer change.

  • Hub sync now uses the configured HTTP proxy — On an on-prem installation that routes outbound traffic through a proxy (STEADYBIT_PROXY_HOST and related settings), hubs you added yourself were fetched without it, so syncing them and checking their connection failed or hung even though webhooks worked. They now go through the proxy like every other outbound request. The built-in Steadybit Reliability Hub and Service Templates keep being read directly from the platform itself, as before.

  • Resyncing a hub too often no longer empties it — Hub resyncs are limited to a few per tenant every ten minutes. A resync turned away by that limit used to discard the hub's templates, actions, extensions and advice until the next successful sync. It now only shows the "Rate limit exceeded" message and leaves the hub as it was.

  • SteadyBuddy no longer loses track of a conversation you have open twice — With the same conversation open in two places (two tabs, or the docked panel beside the full-page chat), two answers could be generated at the same time and the second one to finish would quietly discard what the first had remembered — leaving an exchange visible in the transcript that the assistant then knew nothing about, and dropping a pending question or confirmation raised in the other window. SteadyBuddy now answers one message at a time per conversation: if a reply is already being written, a second message is declined right away with a note, and your text is kept in the composer.

  • Standard-conform environment variables for the database IAM authentication — Tuning the AWS Advanced JDBC Wrapper used to require an environment variable named literally spring.datasource.hikari.data-source-properties.wrapperPlugins. The wrapper reads its settings case-sensitively, so the camel case and the dots could not be avoided — but dots and mixed case are not valid in a POSIX environment variable name, and deployment tooling that enforces the standard refuses to pass such a variable through. Every wrapper setting can now be given in the standard spelling instead, as STEADYBIT_DB_AWS_JDBC_WRAPPER_ followed by the setting name in upper snake case: STEADYBIT_DB_AWS_JDBC_WRAPPER_WRAPPER_PLUGINS sets wrapperPlugins, STEADYBIT_DB_AWS_JDBC_WRAPPER_IAM_HOST sets iamHost. The old spelling keeps working and still wins where both are set, so no existing deployment has to change, and parameters the wrapper merely forwards to the PostgreSQL driver keep using it.

  • Deleted run artifacts and uploaded files no longer take up database space — The content of run artifacts (such as load-test reports) and of files uploaded to experiments is stored in PostgreSQL as large objects, and deleting an artifact, a file, an experiment or a team used to leave that content behind in the database, where nothing could read it again. On long-running installations most of the stored content could be such leftovers. It is now deleted together with its artifact or file, and the upgrade removes the leftovers the platform's database user has accumulated so far. No PostgreSQL extension is needed. The database's large-object table only hands the freed space back to the disk after a VACUUM FULL pg_largeobject by a database administrator. Until then, new artifacts and files reuse it.

Note

  • Experiment run traces are no longer stored in Steadybit — The platform used to keep a copy of the OpenTelemetry spans agents sent for each run, for up to 28 days, downloadable as a zip. That copy is gone: use the run's Tracing tab, which gives you the run's experiment.execution.id and a ready-made query to find its trace in your own tracing backend. See OpenTelemetry Integration. Agents need no change: spans they still send to the platform are discarded.

extension-appdynamics 1.1.23

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump k8s.io/apimachinery from 0.37.0 to 0.37.1

extension-aws 2.4.32

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/config
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigateway
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigatewayv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/dynamodb
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/eks
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticache
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticloadbalancingv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ssm

extension-gcp 1.0.41

  • chore(deps): bump cloud.google.com/go/compute from 1.69.0 to 1.70.0
  • chore(deps): bump cloud.google.com/go/redis from 1.25.0 to 1.26.0
  • chore(deps): bump cloud.google.com/go/run from 1.22.0 to 1.23.0
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump google.golang.org/api from 0.298.0 to 0.299.0

extension-k6 1.3.7

  • chore(deps): bump goreleaser/goreleaser from v2.18.1 to v2.18.2
  • chore(deps): update bundled k6 to v2.3.0 (#212)

extension-kafka 1.2.26

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/twmb/franz-go/pkg/kadm

extension-appdynamics 1.1.22

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-aws 2.4.31

  • Add OpenTelemetry tracing support
  • Tolerate a prepare request without an experiment key in the ALB static-response attack
  • Tolerate stale subnets when mapping NAT gateway AZs
  • Depend on extension-kit v1.12.1
  • Update dependencies

extension-container 1.8.2

  • chore(deps): bump extensionlib to ^1.5.5
  • feat: initialize OpenTelemetry tracing (#517)
  • fix(linuxpkg): require iptables/iproute by name, not by sbin path (#520)
  • fix: stop action reported messages as null while still stopping (#518)

extension-datadog 1.8.31

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-dynatrace 1.0.33

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-grafana 1.1.12

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-host 1.8.2

  • chore(chart): require extensionlib 1.6.0 so otel values take effect (#261)
  • chore(deps): bump extensionlib to ^1.5.5
  • chore: depend on extension-kit v1.12.1 (#262)
  • feat: support otel (#212)
  • fix(linuxpkg): require iptables/iproute by name, not by sbin path (#263)

extension-host-windows 0.3.8

  • Add OpenTelemetry tracing support
  • Match source as well as destination in network attack filters
  • Update dependencies

extension-http 1.0.56

  • Add OpenTelemetry tracing support
  • Require extensionlib 1.6.0 so the otel values take effect
  • Depend on extension-kit v1.12.1
  • Update dependencies

extension-instana 1.1.26

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-istio 1.0.34

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-jenkins 1.0.24

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-jvm 1.3.5

  • chore(deps): bump extensionlib to ^1.5.5
  • chore(deps): bump golang.org/x/net from 0.58.0 to 0.59.0
  • chore: remove fixed cves
  • feat: initialize OpenTelemetry tracing (#461)
  • fix(linuxpkg): require libcap by name, not by sbin path (#460)

extension-newrelic 1.0.28

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-splunk 1.0.20

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

extension-stackstate 1.0.34

  • Add OpenTelemetry tracing support
  • Refuse to start when a required parameter is set but empty
  • Update dependencies

Agent 2.4.6

Improvement

  • The agent reports how it is deployed. Settings -> Agents previously showed the type other for every agent that was neither running in Kubernetes nor on Windows — one label covering every Linux host, Docker, ECS and Nomad install. An agent now reports kubernetes, aws-ecs, nomad, container, windows or linux, detected from its own environment, so the agent list shows how a fleet is actually installed. Where the detection cannot get it right, STEADYBIT_AGENT_TYPE (steadybit.agent.type) overrides it with a label of your choice, up to 64 characters. The type is informational only — nothing about how an agent works or what it may run depends on it.

Fix

  • A step whose action the agent no longer has says so, instead of showing a stack trace. When an experiment reached a step whose action the agent could not find — because the extension providing it is no longer registered — the run page showed a raw Java stack trace whose only real content was the action id. The step now errors with "'<action>' is not available on this agent. The extension providing it is no longer registered - check that the extension is running and registered with the agent, then run the experiment again." Stack traces are still attached for genuinely unexpected failures, where they are what support needs.

Dependencies

  • Jackson updated to 3.1.6 and 2.21.6 to address CVE-2026-68497.

Platform 2.8.7

New

  • Comments in the query language — A query can now carry a // comment explaining itself, so the reason a target was excluded lives next to the exclusion instead of being forgotten. A comment starts at // and runs to the end of the line, either on its own line or after a query on the same line, and works everywhere queries do — environment scopes, service target scopes, experiment blast radii and the API. The query editor greys comments out and toggles them with the usual comment shortcut. A query that has a comment stays in the query editor rather than the visual query builder, because the builder rebuilds a query from its parts and has nowhere to keep one.

  • SteadyBuddy can start an experiment for you, once you confirm it — Ask SteadyBuddy to run one of your team's experiments and it now offers a confirmation card naming the experiment. Nothing starts until you press the button: typing a reply instead withdraws the request, and the card is replaced by the run itself, with the same progress ring, stop button and outcome you get in the activity list. Early access — enabled per tenant on request.

Improvement

  • The MCP server is on by default — Steadybit's MCP endpoint (/mcp) no longer has to be switched on per installation: it is now part of a standard deployment, so an on-prem platform serves it out of the box. Nothing changes for Steadybit Cloud, where it was already on. Access is unchanged and still deliberately narrow — the MCP feature has to be part of your license, and a client has to authenticate with a platform access token (or, where configured, your identity provider). A deployment that would rather not offer it can still turn it off with steadybit.mcp.enabled=false.

  • The agent list names how each agent is deployed — The type column previously showed the raw internal value, so every agent that was neither running in Kubernetes nor on Windows read other. It now shows a proper label and icon for each deployment — Kubernetes, AWS ECS, Nomad, Container, Windows, Linux — matching what recent agents report about themselves. An agent that predates that detection still reports other, now shown as "Other" with a note that updating the agent reveals its actual type, and an agent that has not sent its metadata yet reads "Unknown". A type configured on the agent itself is shown as configured. (requires agent >= 2.4.6)

  • Agent reconnects no longer rewrite every target — When an agent reconnects it re-submits its whole inventory, and until now the platform rewrote every row of it, even the ones whose attributes had not changed at all. On a large installation that is the single heaviest thing the database does, and it happens every time an agent restarts or a network blip drops a connection. Unchanged targets now only have their timestamp refreshed, which the database can do in place.

  • Your queries keep the shape you wrote them in — Resolving a {{variable}} or [[placeholder]] inside a query used to take the query apart and write a fresh one from its parts, so creating an experiment from a template quietly reformatted its blast radius — k8s.namespace = "checkout" came back as k8s.namespace="checkout", and any comment would have been dropped. Only the marker itself is replaced now; the rest of the query, including its spacing, line breaks and comments, is left exactly as written. Deleting a multi-value variable and inlining it into the queries that used it does the same. For API users: because such a query is no longer taken apart, the blast radius of an experiment created from a template is now returned as a query string rather than as a decomposed predicate object. Both forms remain valid on the way in, and they select the same targets.

  • Query autocompletion offers help where you expect it — The query editor — the target selection in the experiment editor, Explorer's query bar and environment and service scopes — used to go quiet at the points you most need it. Typing the opening " of a value no longer empties the list, a finished value like k8s.deployment="gateway" now offers AND/OR instead of "No suggestions." until you typed a space, and accepting a value moves on to what may follow rather than re-offering the value you just picked. Suggestions are also handled by the editor itself now, so they no longer restart from scratch on every keystroke.

Fix

  • A run stopped by a preflight check now reports that, not an agent error — When the agent running the preflight action also ran one of the experiment's own actions, a blocked run ended as an error blamed on the agent ("Preflight action failed") instead of failing with the check's own reason. The same check blocking the same run reported differently depending on which agent happened to pick it up, which made it look as though preflight behaviour had changed when only the agent layout had. The check's verdict now always decides the outcome: the run fails with "Preflight check … failed" and the reason the check gave. A run stopped by a canceled preflight check also names the check it was waiting on rather than leaving the name blank.
  • The experiment editor's search only suggests what your team can actually use — The search bar above the steps sidebar offered actions, tags and target types drawn from every template in the tenant, regardless of what your team is allowed to run. Picking one filtered the sidebar down to nothing, and an action suggested this way could not be dragged onto the canvas. Suggestions are now bound to your team: actions come from those your team may run, and tags and target types from the templates it can use. The same applies when creating an experiment from a template, where the suggestions follow the "show templates I don't have access to" switch — actions stay limited to your team's either way.
  • Accepting a suggestion no longer corrupts the query — Completing a term could damage the text around it: with the cursor just behind a closing quote k8s.deployment != "ga" became k8s.deployment != "ggateway, completing inside COUNT(…) could leave a second ) behind or drop the closing quote of a quoted attribute name. Two query fields on the same page also offered every shared attribute twice, one copy carrying the other field's targets.
  • Applying a query in the target selection — Applying with Cmd/Ctrl+Enter immediately after picking a value applied the previous draft and emptied the selection. The Apply button appeared about a second after you stopped typing, stayed visible after you deleted your way back to the query that was already applied, and applying discarded the editor's undo history. Deleting an experiment variable — which writes its values into the queries that referenced it, so the target selection survives the variable — left an Apply that would undo that and put the deleted variable back.
  • Picking a service's experiment suggestion reads as a suggestion — Choosing one of a service's suggested experiments opened the chat with the raw technical brief as your first message, with nothing to say it came from a suggestion you picked. The chat now records the same "Suggestion used" summary — the suggestion's title and description — that the SteadyBuddy entry page and the in-chat suggestions already show, while SteadyBuddy still receives the full brief. The experiment draft that comes back also names the service as its pill, matching the environment beside it, instead of as a bare link.
  • A schedule that has run for the last time stops showing a next run — A schedule whose cron has an end — a fixed year, or a month and weekday combination that does not come round again — kept advertising the run it had already done after its final execution, because working out the following run failed once there was none left. Such a schedule now simply shows no next execution.
  • The query editor no longer traps the keyboard — Tab belongs to the editor, which left anyone navigating by keyboard unable to leave the field: pressing Escape now releases it, so the following Tab or Shift+Tab moves to the next or previous control. Cmd/Ctrl+F reaches the browser's own find instead of opening a search box over the field, right-click offers the browser menu rather than a code editor's, screen readers announce the field as "Query", and the box sizes itself to the query instead of leaving empty space below it.

extension-container 1.8.1

  • build(deps): bump golang.org/x/sync from 0.22.0 to 0.23.0
  • chore: add CVE-2026-74860 to the ignores
  • fix: refuse to start when a required parameter is set but empty

extension-host 1.8.1

  • build(deps): bump golang.org/x/sync from 0.22.0 to 0.23.0
  • chore: add CVE-2026-74860 to the ignores
  • chore: remove fixed cves

extension-http 1.0.55

  • build(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.1
  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • feat: Response Time Measurement (time to first byte / last byte) for HTTP checks (#187)

extension-postman 2.0.37

  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • fix: bound Postman API calls to the request budget the agent advertises (#154)
  • fix: refuse to start when a required parameter is set but empty

Platform 2.8.6

Improvement

  • Connecting an MCP client with OAuth no longer needs your tenant key — The OAuth endpoint insisted on the tenant in the URL (/mcp/<tenant>), a key nothing ever tells you. Point your client at https://<your-platform>/mcp and the tenant now comes from your membership, the way signing in to the UI already does; if you belong to several tenants the connection says so and lists the tenant-scoped URLs to choose from rather than guessing. Refused connections also explain themselves for the first time — the 401 carried no reason at all, so a client could only report a generic failure. It now says whether you are not a member of that tenant, or have no Steadybit tenant linked to this account and should sign in to the platform once before reconnecting.

extension-appdynamics 1.1.21

  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test from 1.4.9 to 1.4.10
  • chore(deps): bump golang.org/x/oauth2 from 0.36.0 to 0.37.0

extension-auto-registration-ecs 1.0.22

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs

extension-azure 1.3.11

  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • chore(deps): bump goreleaser/goreleaser from v2.18.0 to v2.18.1

extension-datadog 1.8.30

  • build(deps): bump github.com/DataDog/datadog-api-client-go/v2
  • build(deps): bump goreleaser/goreleaser from v2.18.0 to v2.18.1
  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0

extension-gatling 1.0.56

  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • chore: remove fixed CVEs from ignore list

extension-gcp 1.0.39

  • chore(deps): bump cloud.google.com/go/compute from 1.67.0 to 1.68.0
  • chore(deps): bump cloud.google.com/go/spanner from 1.95.0 to 1.95.1
  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • chore(deps): bump google.golang.org/api from 0.294.0 to 0.297.0

extension-istio 1.0.33

  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • chore(deps): bump k8s.io/client-go from 0.36.4 to 0.37.0

extension-jvm 1.3.4

  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • chore(deps): bump github.com/moby/moby/api from 1.55.0 to 1.56.0
  • chore(deps): bump github.com/shirou/gopsutil/v4 from 4.26.7 to 4.26.8
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • chore: remove fixed CVEs from ignore list

extension-k6 1.3.5

  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • chore(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.1
  • chore: remove fixed CVEs from ignore list

extension-kong 2.0.35

  • build(deps): bump goreleaser/goreleaser from v2.18.0 to v2.18.1
  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0

extension-kubernetes 2.6.35

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0
  • chore: remove fixed CVEs from ignore list

extension-prometheus 2.1.32

  • build(deps): bump github.com/moby/moby/api from 1.55.0 to 1.56.0
  • build(deps): bump github.com/prometheus/common from 0.70.1 to 0.71.0
  • build(deps): bump goreleaser/goreleaser from v2.18.0 to v2.18.1
  • chore(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0

extension-container 1.8.0

  • feat: new Fault Dependency Traffic action (transparent-proxy) — transparently proxy a target container's outgoing traffic to a dependency and inject a fault (latency, connection reset, or an HTTP error status), selected by hostname/SNI or HTTP Host header.
  • feat: Intercept Outgoing HTTP Request can now synthesize responses for HTTPS dependencies, not only cleartext HTTP, when an interception CA is configured via tlsIntercept.existingSecret. Off by default; HTTPS is left untouched without it. The statistics widget surfaces tls_intercept_rejected so an untrusted CA is visible rather than a silent no-op.
  • fix: verify the interception CA is a usable certificate/key pair before handing it to the proxy, so a corrupt or mismatched pair fails Prepare with a named error instead of a bare "transparent-proxy failed".
  • chore: refresh the trivy ignore list (systemd, attr, util-linux, zlib; drop fixed CVEs)
  • Update dependencies

extension-host 1.8.0

  • feat: new Fault Dependency Traffic action (transparent-proxy) — transparently proxy the host's outgoing traffic to a dependency and inject a fault (latency, connection reset, or an HTTP error status), selected by hostname/SNI or HTTP Host header.
  • feat: Intercept Outgoing HTTP Request can now synthesize responses for HTTPS dependencies, not only cleartext HTTP, when an interception CA is configured via tlsIntercept.existingSecret. Off by default; HTTPS is left untouched without it. The statistics widget surfaces tls_intercept_rejected so an untrusted CA is visible rather than a silent no-op.
  • chore: refresh the trivy ignore list (systemd, attr, util-linux, zlib; drop fixed CVEs)
  • Update dependencies

extension-auto-registration-ecs 1.0.21

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs

extension-aws 2.4.30

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/autoscaling
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/dynamodb
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/fis
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/rds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/sts
  • chore(deps): bump goreleaser/goreleaser from v2.18.0 to v2.18.1

extension-azure 1.3.10

  • chore(deps): bump github.com/Azure/azure-sdk-for-go/sdk/azidentity
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.0

extension-datadog 1.8.29

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.0

extension-gcp 1.0.38

  • chore(deps): bump cloud.google.com/go/compute from 1.66.0 to 1.67.0
  • chore(deps): bump cloud.google.com/go/container from 1.53.1 to 1.54.0
  • chore(deps): bump github.com/googleapis/gax-go/v2 from 2.24.0 to 2.24.1
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump goreleaser/goreleaser from v2.18.0 to v2.18.1

extension-istio 1.0.32

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump istio.io/api from 1.30.3 to 1.31.0
  • chore(deps): bump istio.io/client-go from 1.30.3 to 1.31.0

extension-jvm 1.3.3

  • chore(deps): bump actions/setup-java from 5 to 6
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump k8s.io/client-go from 0.36.4 to 0.37.0

extension-kong 2.0.34

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.0

extension-kubernetes 2.6.34

  • build(deps): bump github.com/steadybit/advice-kit/go/advice_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump k8s.io/client-go from 0.36.4 to 0.37.0

extension-prometheus 2.1.31

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.0

Agent 2.4.5

Dependencies

  • Apache Tomcat updated to 11.0.25 to address CVE-2026-65182.

Platform 2.8.5

Security

  • Embedded Tomcat updated to 11.0.25 — Fixes CVE-2026-65182 in Apache Tomcat, the HTTP server embedded in the platform. This release contains no other changes.

Agent 2.4.4

Fix

  • Action parameters keep their configured value when an extension changes a parameter's type. A saved experiment step stores each configuration value with the JSON type its parameter had when the step was designed, so an extension changing a parameter's declared type left saved steps carrying a mismatched value — a duration parameter that used to be an integer still carries 30 as a plain number. Such a step failed in the agent before the extension was even called, reporting an internal error ('IntNode' method stringValue() cannot convert value 30 to java.lang.String); older agents instead silently replaced the configured value with the extension's default. A unit-less number is now read as seconds, matching how the agent already times the step, and a value that genuinely cannot be interpreted is reported with a message naming the parameter, its declared type and the value received. Numeric parameters — integer, percentage and stress-ng workers — now follow that same rule: a value that is not a number used to be passed to the extension unchanged, where it turned into a silent 0, so the step ran with zero workers or zero percent instead of failing. Such a step now reports the same named error, while a numeric field left empty still falls back to the extension's default.
  • One unusable target attribute no longer makes all of an extension's targets disappear. If a discovery response contained an attribute whose value was not a string or a list of strings — a nested object, or a nested list — the agent failed to read the entire response and removed every target of that discovery, after retrying twice over 90 seconds. Experiments targeting them then found nothing, with no explanation in the platform. The unusable attribute is now skipped, and its target along with every other target is kept.
  • Restarting an agent no longer wipes and rebuilds its targets on the platform. The first submission after a restart went out before any discovery had run, but still told the platform to delete every target it did not contain — so the agent's whole inventory was deleted and then re-created over the next minute or so. It self-healed, at the cost of re-creating every target: enrichment relationships were invalidated and the full post-processing ran again for each one, as if they had just been discovered. The agent now makes no target submission at all until every registered discovery has reported, so the platform keeps showing what the agent last reported instead of a half-built inventory, and the first submission it does make is complete enough to delete against. If some discovery has not reported after steadybit.agent.target.submission.initial-discovery-timeout (2 minutes by default) it goes ahead with what it has found and logs a warning naming the situation, so a broken extension cannot stop stale targets from ever being cleaned up. If you know how many discoveries an agent should have, set steadybit.agent.target.submission.min-expected-discoveries and it will wait for all of them rather than for whichever registered first.
  • A discovery that fails once no longer holds up the others. The discovery-kit retry backed off for 30 seconds and then 60 before giving up on a failed call, holding one of the agent's concurrent-discovery slots the whole time — so a handful of unreachable extensions could delay the discoveries that were working. It now retries after 1 second and then 2, configurable as steadybit.agent.discoveries.retry-delay.
  • Extensions are available again immediately after an agent restart. Extensions registered through the agent's local API — which is how the Kubernetes and ECS auto-registration sidecars register them — were already stored durably, but the agent waited 30 seconds before reading them back, so for the first half minute after every restart it behaved as if it had no extensions at all. It now reads them at startup.
  • A damaged extension registration file now repairs itself. If the agent was killed mid-write — an OOM kill, a node going away — the file it stores locally-registered extensions in (extensions.yml under the agent's state directory) could be left half-written. The agent then failed to read it on every attempt for as long as the file existed, so extensions registered through its local API, which is how the Kubernetes and ECS auto-registration sidecars register them, never came back without someone deleting the file by hand. An unreadable file is now discarded and logged, and the sidecars re-register their extensions on their next reconcile. The file is also written in a way that a kill can no longer truncate, so it should not become damaged in the first place.
  • A single failing registration source no longer stops extension registration entirely. If one source failed — Redis briefly unreachable, an unreadable extensions.yml — registration stopped for all sources, including Kubernetes annotations, config files and properties, until the agent was restarted. A failure is now contained to the source it came from, and reading the stored registrations retries by itself until it succeeds.

Platform 2.8.4

New

  • SteadyBuddy and the Remote MCP Server are generally available — Both features leave the Steadybit Labs early-access program and are no longer enabled per tenant on request. SteadyBuddy, the AI assistant that designs experiments, explains failed runs and suggests what to test next, is now available to every tenant whose plan includes it. The same goes for the Remote MCP Server, which connects AI agents such as Claude Code, Claude Desktop or GitHub Copilot directly to your Steadybit tenant.

Improvement

  • Hub sync is faster, and survives a throttling hub — Preparing a tenant downloads the Steadybit Reliability Hub file by file, one request at a time, which is most of the wait on the signup screen. Those files are now fetched several at a time, and if the hub answers with a rate limit the sync waits as long as the hub asks and retries instead of importing nothing.

  • SteadyBuddy shows what it is working on — While SteadyBuddy composes an answer the chat now lists each capability it uses as it starts and finishes, with the steps of a specialist capability nested underneath it, so a long-running answer shows progress instead of only a spinner. Once the answer arrives the list collapses to a single "Used 3 tools" line above it.

  • Faster validation in the experiment editor on large environments — The editor re-checks each step's blast radius while you type. On installations with hundreds of thousands of targets those checks took seconds apiece, so an experiment with several steps — especially with advanced blast-radius attributes — felt sluggish to edit. The queries behind them now ask only what the check needs to know: whether an attribute exists at all, and whether a step matches no target, exactly one, or more than one. On an environment of 112,000 containers that takes a check from about 4.5 seconds to under 20 milliseconds. The target counts shown in the editor are unchanged.

Fix

  • SteadyBuddy says when the AI limit is reached, wherever it can be reached — Once a tenant had used up its SteadyBuddy budget for the period, only the chat and the experiment suggestions for an environment said so. A service's experiment suggestions failed with a generic message, their "Regenerate" quietly handed back the previous suggestions, and asking SteadyBuddy to explain a run looked like it did nothing at all. All of them now show the same "your tenant has reached its SteadyBuddy limit" message, and "Regenerate" — on a service and on the SteadyBuddy page alike — is disabled with that reason on hover instead of silently returning what was already there. A run analysis refused for any other reason — a run that has not finished yet, for instance — now shows that reason too, instead of nothing.

  • SteadyBuddy is only offered where it can actually work — The navigation entry, the "Ask SteadyBuddy" bar, a service's experiment suggestion box and the run-analysis panel appeared whenever the plan included them, even on an installation with no AI provider configured or for a tenant that had not opted in — each of them then failed as soon as it was used. Every surface now asks the platform whether SteadyBuddy is actually available; a tenant that has not opted in keeps only the navigation entry, which leads to the opt-in page. Enabling SteadyBuddy also takes effect right away instead of requiring a page reload, and a single failed availability check no longer leaves the session stuck showing SteadyBuddy as unavailable.

  • Deleting a user now reaches the identity provider even if it is briefly unavailable — When a user lost their last team membership, the platform removed them locally before telling the identity provider, so a failure there left the account behind with no way to retry. The provider is now told first, and the account is kept locally until it confirms, so the nightly clean-up picks it up again. One account the provider refuses no longer holds up the others being deleted in the same pass.

  • New tenants no longer come up with a half-imported hub — Preparing a tenant downloads the Steadybit Reliability Hub and, with Services enabled, the bundled service templates. That download held a database transaction open for its whole duration, so a slow hub could pass the database's idle-transaction limit and have its connection dropped mid-import — leaving the tenant without service templates, service profiles or sample services until a later sync pass repaired it. The same could interrupt the nightly hub sync. Hub content is now downloaded without a transaction open, and only the result is written. Changing a hub's repository URL while it is syncing no longer leaves it showing the new URL with content from the old one — the in-flight download is discarded and the hub says so.

  • Malformed requests get a 4xx, not a server error — A request whose query parameter could not be decoded was answered with 500 Internal Server Error, even though the problem was in the request. Such requests now get 400 Bad Request, and one that is simply too large gets 413 Content Too Large — including uploads over the size limit, which previously reported the unrelated 417 Expectation Failed.

  • Explorer no longer waits forever when a target type has no icon — The landscape Explorer waits for the icons of every target type discovered in your tenant before it draws the map, and an icon it could not load was never counted as settled. More than three such target types — an extension that registers target types with no icon, or with one that does not decode, produces them — left the map permanently unfinished: the spinner ran indefinitely for every query and every environment, even though the targets had already been loaded. A target type whose icon is missing or unusable now gets a placeholder glyph, and an icon that fails to load no longer holds up the map.

  • A broken extension icon no longer leaves a blank space — Target types, actions and advice show the icon their extension provides. An icon that was missing, or present but not a valid image, was drawn as nothing: an empty gap in the target type pill, the discovered-targets widget on the dashboard and the advice lists. Each of those now falls back to the matching built-in icon instead.

  • Trend arrows on the dashboard point the right way — The Team Activities widget shows a trend arrow beside each metric, and it was inverted: a rising count showed a downward arrow and a falling one showed an upward arrow. The arrow is the only directional signal on the card — the percentage next to it carries no sign — so a growing number read as a shrinking one.

  • Fix a batch of smaller UI issues — The actions and templates list in the experiment and template editors no longer stops short of the sidebar's bottom edge, the run page's log fills the card it sits in when a neighbouring card is taller, a tooltip no longer stays visible permanently, the rows in SteadyBuddy's entity pop-over keep their spacing and alignment while they are still resolving, the query editor's category tabs line up with the panel below them, and the suggested-extension cards on the onboarding page are no longer clipped.

extension-auto-registration-ecs 1.0.20

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • build(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-aws 2.4.29

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/feature/ec2/imds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigatewayv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticloadbalancingv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/lambda
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/mq
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/sqs
  • chore(deps): bump github.com/aws/smithy-go from 1.27.8 to 1.28.1
  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1
  • chore(deps): bump github.com/testcontainers/testcontainers-go/modules/localstack
  • chore(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.0

extension-azure 1.3.9

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1

extension-cloudfoundry 1.0.8

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-container 1.7.8

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • build(deps): bump golang from 1.26-trixie to 1.27-trixie
  • build(deps): bump k8s.io/client-go from 0.36.3 to 0.36.4
  • chore: bump runc to v1.5.1 and crun to 1.29.1 (#493)
  • fix(linuxpkg): add missing iptables/fallocate rpm deps, drop unused ps (#494)
  • fix: extract dns-inject straight to its arch-suffixed path (#495)

extension-datadog 1.8.28

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • fix: bound all Datadog API calls with a timeout (#232)
  • fix: bound requests to the Datadog API with a timeout (STEADYBIT_EXTENSION_API_TIMEOUT, default
  • fix: bound the monitor discovery refresh with a timeout so a hanging Datadog API cannot stall the
  • fix: prune step execution data that is retained when the matching experiment-completed event is
  • fix: the monitor status check now backs off between retries and stops retrying once the caller's

extension-dynatrace 1.0.31

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.0

extension-gatling 1.0.54

  • build(deps): bump docker/login-action from 3 to 4
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • build(deps): bump golang from 1.26-trixie to 1.27-trixie
  • chore: drop the zip/unzip binaries in favour of archive/zip (#146)
  • feat(ci): prune stale snyk projects after a release (#147)
  • fix: attach the gatling report again
  • refa: drop the dead -Dgatling.resultsFolder flag
  • test(e2e): wait for the run and assert the report artifact

extension-gcp 1.0.37

  • chore(deps): bump cloud.google.com/go/pubsub/v2 from 2.6.2 to 2.7.0
  • chore(deps): bump cloud.google.com/go/spanner from 1.94.0 to 1.95.0
  • chore(deps): bump github.com/googleapis/gax-go/v2 from 2.23.0 to 2.24.0
  • chore(deps): bump google.golang.org/api from 0.293.0 to 0.294.0
  • chore(deps): bump goreleaser/goreleaser from v2.17.1 to v2.18.0

extension-grafana 1.1.9

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine
  • fix: bound requests to the Grafana API with a timeout (STEADYBIT_EXTENSION_API_TIMEOUT, default
  • fix: bound the annotation search to the annotation's own time window. Searching /api/annotations
  • fix: bound the annotation search to the annotation's time window (#105)
  • fix: drain queued annotations on shutdown so a rolling restart does not silently discard them.
  • fix: reject an experiment-step-completed event without step execution data instead of panicking.
  • fix: stop calling the Grafana API on the request goroutine (#104)
  • fix: the experiment event listeners no longer call the Grafana API on the request goroutine.

extension-host 1.7.5

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • build(deps): bump golang from 1.26-trixie to 1.27-trixie
  • chore: bump runc to v1.5.1 and crun to 1.29.1 (#246)
  • fix(linuxpkg): add missing iptables/fallocate rpm deps, drop unused ps (#247)
  • fix: extract dns-inject straight to its arch-suffixed path (#248)

extension-istio 1.0.31

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine
  • chore(deps): bump k8s.io/client-go from 0.36.3 to 0.36.4

extension-jenkins 1.0.21

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-k6 1.3.3

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine
  • chore(deps): bump k8s.io/client-go from 0.36.3 to 0.36.4
  • chore: zip artifacts in-process, drop the zip package dependency (#200)
  • test(e2e): assert the run's output files come back as artifacts

extension-kong 2.0.33

  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1

extension-kubernetes 2.6.33

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • build(deps): bump golang from 1.26-alpine to 1.27-alpine
  • build(deps): bump k8s.io/api from 0.36.3 to 0.36.4
  • build(deps): bump k8s.io/client-go from 0.36.3 to 0.36.4

extension-newrelic 1.0.26

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-postman 2.0.36

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • build(deps): bump golang from 1.26-alpine to 1.27-alpine
  • test(e2e): wait for the run and assert its artifacts

extension-prometheus 2.1.30

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1

extension-rabbitmq 1.1.3

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • fix: keep the publish duration as an integer number of seconds
  • fix: keep the publish duration as an integer number of seconds (#56)
  • fix: publish duration is a duration parameter, and e2e covers the actions (#55)

extension-redis 1.1.14

  • chore(deps): bump github.com/KimMachineGun/automemlimit
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-splunk 1.0.17

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • build(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-splunk-platform 1.0.17

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-jvm 1.3.2

  • chore(deps): bump docker/setup-buildx-action from 3 to 4
  • chore(deps): bump github.com/moby/go-archive from 0.2.0 to 0.3.0
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.12.0 to 1.12.1
  • chore(deps): bump github.com/testcontainers/testcontainers-go
  • chore(deps): bump golang from 1.26-trixie to 1.27-trixie
  • chore(deps): bump k8s.io/client-go from 0.36.3 to 0.36.4
  • chore(deps): bump steadybit kits and drop Go patch pin (#439)
  • chore(deps): float Go on 1.26.x in the standalone setup-go step
  • chore(linuxpkg): drop the unused procps dependency (#441)
  • chore: bump runc to v1.5.1 and crun to 1.29.1 (#440)
  • ci: add --ignore-scripts to npm tool installs (Sonar) (#438)
  • ci: migrate to npm 12 (#434)
  • ci: pin npm to 12.0.2 to satisfy Sonar dependency pinning (#437)

Agent 2.4.3

Improvement

  • New setting steadybit.agent.shutdown.grace-period (default 25s) — the total time the agent may spend rolling back running attacks and reporting their outcomes while it shuts down. Keep it below your container's terminationGracePeriodSeconds (30s by default in Kubernetes), otherwise the pod is killed before rollbacks can finish.

Fix

  • Linux package installs (deb/rpm) now keep their state across restarts. /var/lib/steadybit-agent was never created by the package, so the agent — running as the unprivileged steadybit user — could not write to it and fell back to in-memory state. Attacks that were still running when the agent stopped were therefore not rolled back on the next start, and extensions registered through the agent's local API were lost. Existing installations are fixed by upgrading; no configuration change is needed.
  • The agent now starts on hardened clusters that mount the temp directory noexec. Determining its own host name no longer loads a native library on Linux. Previously, a pod that both ran under an arbitrary user id (as OpenShift's restricted SCC and read-only root filesystems do) and had a noexec /tmp failed to start with UnsatisfiedLinkError: … failed to map segment from shared object.
  • Duration settings without a unit now mean seconds. Previously a unit-less value was read as milliseconds, so 30 meant 30ms. Values with an explicit unit (30s) are unaffected.
  • Re-running an experiment right after cancelling it no longer fails preparation with "Attack … currently active". While the previous execution is still rolling back, the agent now says so and asks you to retry shortly.
  • Cancelling an experiment now completes in seconds instead of scaling with the target count. Teardown bounds the whole execution with a single grace period rather than 30s per target, so the platform no longer waits tens of minutes on large target sets. Every target also gets a terminal outcome: targets whose rollback was still finishing when the experiment channel closed previously ended up with no result at all.
  • Stopping or restarting the agent during an experiment no longer reports every target as "Agent disconnected unexpectedly". Rollbacks get a bounded window to finish and their outcomes are reported; anything still rolling back when that window closes is reported as errored with "The agent shut down before the rollback finished", so the execution always ends instead of appearing to run on.

Dependencies

  • Dependency updates for CVEs.

extension-appdynamics 1.1.19

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine
  • chore(deps): bump k8s.io/apimachinery from 0.36.3 to 0.36.4

extension-auto-registration-kubernetes 1.0.15

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1
  • build(deps): bump golang from 1.26-alpine to 1.27-alpine
  • build(deps): bump k8s.io/client-go from 0.36.3 to 0.36.4

extension-aws 2.4.28

  • chore(deps): bump github.com/aws/aws-sdk-go-v2 from 1.43.6 to 1.43.7
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/feature/ec2/imds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigateway
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/kafka
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/rds
  • chore(deps): bump github.com/moby/go-archive from 0.2.0 to 0.3.0
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/testcontainers/testcontainers-go

extension-cloudfoundry 1.0.7

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1

extension-dynatrace 1.0.30

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1

extension-gcp 1.0.36

  • chore(deps): bump cloud.google.com/go/pubsub/v2 from 2.6.1 to 2.6.2
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1
  • chore(deps): bump google.golang.org/grpc from 1.83.0 to 1.83.1

extension-http 1.0.53

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1

extension-instana 1.1.24

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-kafka 1.2.22

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1

extension-rabbitmq 1.1.2

  • chore(deps): bump github.com/rabbitmq/amqp091-go from 1.13.0 to 1.14.0
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine

extension-redis 1.1.13

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1

extension-stackstate 1.0.31

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_test
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_test
  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1
  • chore(deps): bump golang from 1.26-alpine to 1.27-alpine

Platform 2.8.3

The following changes are part of Steadybit Labs, our program for early-access features. Expect them to evolve based on your feedback.

Labs / Improvement

  • Run analysis only judges the hypothesis you wrote — When an experiment carries no hypothesis, SteadyBuddy's run analysis now leaves the hypothesis out entirely instead of deriving one from the experiment's design and labelling it. Where there is one, the judgement no longer repeats its wording — hover the row to read the hypothesis itself.
  • SteadyBuddy links to the entities it names — When SteadyBuddy mentions an experiment, service, environment, template or action, it now renders it as a pill that opens that thing in the product instead of naming a raw id.
  • Ground experiment suggestions in the target types you actually run — Suggested experiments are checked against the target types actually discovered in your environment, so a suggestion no longer proposes an attack for infrastructure you do not run.

Labs / Fix

  • Run analysis icon sizing and spacing — The verdict icon, the row spacing, and the run analysis panel's header now match the design's sizing exactly.
  • Consistent SteadyBuddy error states — When SteadyBuddy cannot suggest experiments for a service or analyse a run, it now says so in the same centered message it uses everywhere else, with the retry inline in the sentence.
  • Remote MCP reports a bad log type or metric name instead of an empty answer — Asking for a log stream or a metric an execution does not have came back as an empty-but-successful result over MCP, which reads to an AI agent as "this run recorded nothing" and could lead it to report that as a finding. Both now return a clear error naming what can actually be read for that run.
  • Remote MCP search finds what you ask for — Searching the action catalog over MCP had no effect: the search term was ignored and the full catalog came back, so a connected agent could not look an action up by keyword. Searching now works, and all three MCP searches — actions, experiments and templates — match each word of the query separately instead of requiring the whole phrase to appear in one field. So "container logs" finds the container log action, and "spike checkout" finds the "Checkout latency spike" experiment, neither of which matched before.
  • Links in SteadyBuddy answers cannot be disguised as internal — Markdown link forms that could make an external destination look like a product link — reference-style definitions inside reference data, and a backslash in place of the URL's authority — are now resolved as external and marked as such.
  • Starting an experiment from SteadyBuddy while another one is running — Running an experiment from SteadyBuddy while another experiment was in flight silently redirected to the experiment's design without starting anything. It now opens the same parallel-run confirmation used everywhere else, so the run can be started deliberately alongside the ones already in flight.

New

  • Variables and placeholders in metric check values — The threshold a metric check compares against (e.g. in the Prometheus metrics check) can now be a {{variable}} — environment, service, or experiment variable, including per-run overrides — or a [[placeholder]] in templates, instead of only a fixed number. The value is resolved when the run starts and fails fast with a clear error if it doesn't resolve to a number.
  • Validate queries in the environment editor — Environment scope queries are now checked as you type, exactly as in the experiment editor: invalid syntax, attribute keys and values that do not exist, and references to service attributes are flagged per field and summarized below the editor. The validation is advisory and never blocks saving — an attribute that only appears after the next discovery run is a legitimate configuration.
  • Map attribute values Explorer has not discovered yet — The grouping and color-override modals in the landscape view now accept typed attribute values, so a bucket can map a value before discovery has ever seen it.

Improvement

  • Linked highlighting in run charts — The run page's chart widgets now show the total count in the center of their summary donut, and hovering a legend entry, a donut slice, or a point in the chart highlights that category across all three.
  • Missed runs of a scheduled experiment are skipped — If the platform is unavailable at the time a recurring schedule was due, that run is now skipped and the schedule continues with its next occurrence, instead of starting the experiment late once the platform is back. A schedule set for a single point in time still runs as soon as the platform is available again.
  • Show the running experiments in the parallel-run confirmation modal — Starting an experiment while others are in flight now lists the running experiments in the confirmation dialog, each linking to its run, instead of only warning that something else is running.
  • Validate attribute keys and values in Explorer's landscape view — Group by, size by, color by and bucket mappings in the landscape view now flag attribute keys and values that do not exist, with the same info-level hint the experiment query editor uses, so a typo is no longer indistinguishable from a valid setup. It never blocks saving a view.
  • Mark the selected entry consistently in every dropdown — Value, attribute-key, variable and placeholder suggestions all show the same selected state — accent background and checkmark on a full-width row.
  • Link the settings navigation to the public changelog page — The Changelog entry in the settings navigation now links to changelog.steadybit.com like the other external items, replacing the embedded widget.
  • Simplify onboarding for trial tenants — Sample data is loaded automatically for early-access trials, so the welcome page no longer asks you to choose between sample data and installing an agent, and is skipped entirely once sample data is present. On-premises onboarding is unchanged.

Fix

  • Inviting a user says what went wrong — A failed invitation reported only "Failed to invite user", whatever the cause, so an invite that could never succeed looked like something worth retrying. The invitation service now reports a specific reason and the platform turns it into an actionable message: whether the email domain still needs setting up, whether the domain requires Google SSO, whether too many invitations were sent at once, or whether the account was created but the invitation email could not be sent. Every failure is also recorded with a correlation id so support can trace an individual invitation.
  • On-prem Reliability Hub import — Importing templates from the bundled Steadybit Reliability Hub on an on-premises installation no longer silently fails to load reliability advice content.
  • Run data of deleted experiments is cleaned up again — The nightly jobs that remove the metrics, logs and artifacts left behind by deleted experiment runs asked the database for those rows in a way that made it read the whole table on every pass, so on installations with a lot of run data they ran into their query timeout and gave up without deleting anything, leaving that data on disk indefinitely. They now look the affected runs up through an index and delete their data in batches.
  • Experiment run widget legends — Chart-widget legends on the run page no longer grow scrollbars when space is tight (permanently visible in Safari): legend entries now wrap across the card's full width, the chart trades height with the legend, and on narrow screens the run page's widgets stack in a single column instead of cramping side by side.
  • Variables in environment and service scopes — Variables were never resolved in environment scope and service target scope queries, silently matching no targets. They are no longer suggested in the scope editors, using one shows a clear validation message, and saving a scope containing a variable is now rejected — in the UI and in the API
  • Service scopes can no longer reference service attributes via the UI — Saving a service whose target scope references service.id or service.name is now also rejected in the UI, as it already was in the API; such self-references made the service's target set unstable. The scope editors also no longer suggest the service.id/service.name attributes, and a typed-in reference is flagged on the affected query row
  • Report a step that never started as skipped, not canceled — Stopping a run in the window between dispatching a step to the agents and the agents confirming it started showed the step as canceled with nothing but skipped targets and no start time, suggesting an action had been rolled back that never ran. Such a step is now reported as skipped, in agreement with its targets.
  • Open the run properties sidebar from every run subtab — The Properties button did nothing on the Agent Log and Tracing subtabs, the experiment player's layers could cover the sidebar, and the button no longer only opens the sidebar but closes it again too.
  • Fix the experiment editor and run view layout — The editor's sidebars span the full content height, both pages size themselves from their container instead of computing viewport math, so nothing is cut off or overlaps at unusual window sizes, the run step modal's header background is back, and the Metrics header bar matches the other header bars.
  • Keep the query editor stable while you type — The query editor no longer recreates its editor instance and completion provider while you type, which could drop autocompletion or lose focus mid-edit.
  • Searching experiments no longer fails when a match has no environment — A search whose results included an experiment without an environment returned an error instead of the matches.
  • Fix a batch of smaller UI issues — Truncated text now really shows its "read more" ellipsis, a composed single-line title keeps its icon centered, a click that both closes a modal and navigates no longer swallows the navigation, a confirmation dialog containing a link no longer crashes, a tooltip on an agent log line no longer adds page-level scrollbars, the page width stays constant when a scrollbar appears, and clicking into a value input keeps the whole value selected when the click jitters.

Dependencies

  • Dependency updates for CVEs.

extension-gcp 1.0.35

  • feat(extvm): add gcp.region attribute to VM targets
  • fix(extvm): copy VM labels and instance id to enriched targets

extension-auto-registration-ecs 1.0.19

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#222)

extension-aws 2.4.27

  • chore(deps): bump steadybit kits and drop Go patch pin (#961)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#960)

extension-azure 1.3.8

  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#212)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#211)

extension-cloudfoundry 1.0.6

  • chore(deps): bump steadybit kits and drop Go patch pin (#23)
  • chore(deps): bump steadybit kits and drop Go patch pin (#24)

extension-container 1.7.7

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump dns-inject to v0.2.6
  • chore(deps): bump steadybit kits and drop Go patch pin (#490)
  • chore(deps): pin action_kit_commons to the released v1.11.0
  • chore: drop the cgexec dependency from fill memory

extension-datadog 1.8.27

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#230)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#229)

extension-dynatrace 1.0.29

  • chore(deps): bump steadybit kits and drop Go patch pin (#148)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#147)

extension-gatling 1.0.53

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#144)

extension-gcp 1.0.34

  • chore(deps): bump steadybit kits and drop Go patch pin (#390)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#389)

extension-grafana 1.1.8

  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#101)
  • chore(deps): bump steadybit kits and drop Go patch pin (#102)

extension-host 1.7.4

  • Revert "fix: make the rpm installable on Enterprise Linux 9 (#240)" (#241)
  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump dns-inject to v0.2.6
  • chore(deps): bump steadybit kits and drop Go patch pin (#243)
  • chore(deps): pin action_kit_commons to the released v1.11.0
  • chore(deps): update dns-inject to 0.2.5
  • chore: bump action_kit_commons to v1.10.5
  • chore: drop the cgexec dependency from fill memory
  • fix: make the rpm installable on Enterprise Linux 9 (#240)

extension-host-windows 0.3.7

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#198)

extension-http 1.0.52

  • chore(deps): bump steadybit kits and drop Go patch pin (#184)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#183)

extension-instana 1.1.23

  • chore(deps): bump steadybit kits and drop Go patch pin (#91)
  • chore(deps): bump steadybit kits and drop Go patch pin (#92)

extension-istio 1.0.30

  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#315)

extension-jenkins 1.0.20

  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#38)
  • chore(deps): bump steadybit kits and drop Go patch pin (#39)

extension-k6 1.3.2

  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#198)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#197)

extension-kong 2.0.31

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • build(deps): bump github.com/testcontainers/testcontainers-go
  • chore(deps): bump steadybit kits and drop Go patch pin (#263)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#262)

extension-kubernetes 2.6.32

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#343)

extension-newrelic 1.0.25

  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#103)

extension-postman 2.0.35

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#146)

extension-prometheus 2.1.28

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • build(deps): bump github.com/testcontainers/testcontainers-go
  • chore(deps): bump steadybit kits and drop Go patch pin (#293)
  • chore(deps): pin goreleaser build toolchain to go1.26.6
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#292)

extension-rabbitmq 1.1.1

  • chore(deps): bump steadybit kits and drop Go patch pin (#49)
  • chore: deprecate the exchange parameter of the queue publish attacks (#48)

extension-splunk 1.0.16

  • build(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#65)

extension-splunk-platform 1.0.16

  • chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.0
  • chore(deps): bump steadybit kits and drop Go patch pin (#47)
  • chore(deps): use go-version-file, drop patch pin (go 1.26) (#46)

extension-auto-registration-ecs 1.0.18

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs

extension-aws 2.4.26

  • chore(deps): bump github.com/aws/aws-sdk-go-v2 from 1.43.4 to 1.43.5
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/feature/ec2/imds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigateway
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/autoscaling
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticloadbalancingv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/eventbridge
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/kafka
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/sqs

extension-gcp 1.0.33

  • chore(deps): bump cloud.google.com/go/compute from 1.65.0 to 1.66.0
  • chore(deps): bump google.golang.org/api from 0.292.0 to 0.293.0
  • chore(deps): bump google.golang.org/protobuf from 1.36.11 to 1.36.12

extension-rabbitmq 1.1.0

  • feat: discover RabbitMQ exchanges as targets (excluding the default exchange, amq.* built-ins and internal exchanges)
  • feat: fetch exchanges paged and column-filtered so brokers with thousands of exchanges are discovered in bounded chunks
  • feat: new attacks "Publish to Exchange (# of Messages)" and "Publish to Exchange (Messages / s)" — delivery is determined by the exchange type and its bindings; unroutable messages count as failures
  • feat: fail the prepare step of the queue publish attacks when the exchange parameter is set and more than 10 queue targets publish to the same exchange and routing key, preventing accidental load amplification
  • fix: reject numberOfMessages = 0 at prepare instead of completing instantly with a confusing 0% success rate
  • fix: wait for in-flight publish confirmations when an attack stops, so the success-rate verdict no longer misses the last message's confirm (e.g. 119/120 on a fully successful run)
  • fix: log a warning when the cluster name cannot be resolved during discovery instead of silently reporting an empty rabbitmq.cluster.name

extension-auto-registration-ecs 1.0.17

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • build(deps): bump github.com/jarcoal/httpmock from 1.4.1 to 1.4.2
  • build(deps): bump github.com/steadybit/extension-kit

extension-aws 2.4.25

  • chore(deps): bump github.com/aws/aws-sdk-go-v2 from 1.43.2 to 1.43.4
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/config
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigatewayv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/dynamodb
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticloadbalancingv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/fis
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/mq
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/rds

extension-gcp 1.0.31

  • chore(deps): bump cloud.google.com/go/container from 1.51.0 to 1.53.0
  • chore(deps): bump cloud.google.com/go/container from 1.53.0 to 1.53.1
  • chore(deps): bump cloud.google.com/go/pubsub/v2 from 2.6.0 to 2.6.1
  • chore(deps): bump cloud.google.com/go/redis from 1.23.0 to 1.24.0
  • chore(deps): bump cloud.google.com/go/run from 1.21.0 to 1.22.0
  • chore(deps): bump cloud.google.com/go/spanner from 1.91.0 to 1.93.0
  • chore(deps): bump cloud.google.com/go/spanner from 1.93.0 to 1.94.0
  • chore(deps): bump google.golang.org/api from 0.290.0 to 0.291.0
  • chore(deps): bump google.golang.org/grpc from 1.82.0 to 1.82.1
  • chore(deps): bump goreleaser/goreleaser from v2.17.0 to v2.17.1
  • chore(deps): fix go mod tidy
  • chore(deps): update dependencies
  • cleanup: dedupe shared attribute descriptors, single-owner per attribute (#380)
  • cleanup: extract duplicated attribute-name literals per Sonar go:S1192 (#361)
  • cleanup: extract extvm + extnat Sonar leftovers (#364)
  • feat: add discovery + attacks for 11 GCP services (GKE, MIG, Cloud NAT, Cloud SQL, Spanner, Pub/Sub, Memorystore, Cloud Run, Persistent Disk) (#334)
  • feat: support filtering targets out of discovery
  • feat: swap generic placeholder icons for official GCP product icons (#384)
  • fix(attacks): tighten descriptions to one line and align Technology='GCP' (#366)
  • fix(cloud-nat, mig, gke): three attack-implementation fixes (#381)
  • fix(cloudsql): send required FailoverContext with settingsVersion (#383)
  • fix: address correctness findings from PR #334 code review (#363)
  • fix: shorten Pub/Sub topic persistence regions attribute name (#378)

extension-http 1.0.51

  • fix: differentiate transport errors and HTTP status codes in the bandwidth check's metric instead of collapsing them into a bare failure count, and report the status code for every response received
  • fix: report bandwidth check failures the same way the other HTTP checks do (#180)
  • refactor: reduce cognitive complexity of bandwidthChecker.emitWindowMetric (#181)
  • refactor: reduce cognitive complexity of bandwidthChecker.emitWindowMetric (#182)

Agent 2.4.2

Fix

  • Kubernetes API call metrics no longer grow a series per touched object: the uri tag of k8s.api.call now normalizes namespaces and object names for every path shape, not just pods and services under /api/v1.
  • User info and query-string parameters (which can carry endpoint tokens) are dropped from extension-call OTEL spans, keeping only scheme, host, port and path.

Dependencies

  • Upgrade to Spring Boot 4.1 and update the remaining dependencies, including CVE fixes for logback (CVE-2026-13006) and Bouncy Castle.

extension-appdynamics 1.1.17

  • feat: support filtering targets out of discovery
  • fix: emit the health rule state metric immediately on Start (#75)
  • fix: remove 'monitoring' category from suppression action for consistent grouping
  • fix: remove the "monitoring" category from the "Create Action Suppression" action so both actions are grouped consistently in the experiment editor

extension-aws 2.4.24

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/config
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigateway
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigatewayv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/applicationautoscaling
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/lambda
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/rds
  • feat: support filtering targets out of discovery
  • fix: rename aws.zone label to "Zone" for consistency across cloud extensions

extension-azure 1.3.7

  • chore(deps): bump goreleaser/goreleaser from v2.17.0 to v2.17.1
  • feat: support filtering targets out of discovery
  • fix: add missing label for azure.zone attribute

extension-cloudfoundry 1.0.5

  • Add a "Fail early" option to the app state check. When enabled (the default, matching the previous behavior), the check fails as soon as a deviating state is observed. When disabled, the check keeps collecting events for the whole duration and only fails at the end of the step (with a past-tense message, since the state may have recovered by then).
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#20)
  • chore(deps): update dependencies
  • chore: add Claude Code workflows (#15)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • feat(app state check): add fail early option (#16)
  • feat: support filtering targets out of discovery
  • fix: emit the app state metric immediately on Start (#22)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#21)

extension-container 1.7.5

  • build(deps): bump github.com/moby/moby/client from 0.5.0 to 0.5.1
  • build(deps): bump google.golang.org/grpc from 1.82.1 to 1.83.0
  • build(deps): bump k8s.io/client-go from 0.36.2 to 0.36.3
  • chore(deps): update dns-inject
  • feat: support filtering targets out of discovery
  • fix: use gcp.zone instead of google.zone as availability zone fallback attribute (#489)

extension-datadog 1.8.25

  • build(deps): bump goreleaser/goreleaser from v2.17.0 to v2.17.1
  • feat: support filtering targets out of discovery
  • fix: emit the monitor status metric immediately on Start

extension-grafana 1.1.7

  • chore(deps): bump github.com/jarcoal/httpmock from 1.4.1 to 1.4.2
  • feat: support filtering targets out of discovery
  • fix(alert rule discovery): deduplicate targets and exclude recording rules (#99)
  • fix: emit the alert rule state metric immediately on Start (#100)

extension-host 1.7.3

  • chore(deps): update dns-inject
  • feat: support filtering targets out of discovery
  • fix: use gcp.zone instead of google.zone as availability zone fallback attribute (#239)

extension-host-windows 0.3.5

  • feat: support filtering targets out of discovery
  • fix: use gcp.zone instead of google.zone as availability zone fallback attribute (#196)

extension-http 1.0.50

  • chore: consistent title casing for parameter labels
  • feat: support filtering targets out of discovery
  • fix: prevent response time verification value from silently resetting (#179)
  • fix: response time verification value could not be unset and silently reverted to its default; the verification mode and response time are now required fields with clarified labels and tooltips

extension-jvm 1.3.1

  • chore(deps): bump github.com/shirou/gopsutil/v4 from 4.26.6 to 4.26.7
  • feat: support filtering targets out of discovery

extension-k6 1.3.1

  • chore(deps): bump github.com/jarcoal/httpmock from 1.4.1 to 1.4.2
  • chore(deps): bump goreleaser/goreleaser from v2.17.0 to v2.17.1
  • feat: support filtering targets out of discovery

extension-kafka 1.2.19

  • feat: support filtering targets out of discovery
  • fix: emit broker/consumer group check metrics immediately on Start (#132)
  • fix: emit partition/lag check metrics immediately on Start (#133)

extension-kong 2.0.29

  • build(deps): bump goreleaser/goreleaser from v2.17.0 to v2.17.1
  • feat: support filtering targets out of discovery

extension-kubernetes 2.6.31

  • Update CHANGELOG.md
  • chore(deps): update dependencies
  • feat: support filtering targets out of discovery
  • fix: emit the pod count metric immediately on Start (#342)
  • fix: show "Workload owner" as the pod table's last column instead of the joined "Deployment name / StatefulSet name / DaemonSet name" fallback column; adds labels for k8s.workload-owner and k8s.workload-type
  • fix: show workload owner column for pods instead of joined fallback labels

extension-newrelic 1.0.24

  • feat: support filtering targets out of discovery
  • fix: continue workload discovery with the remaining accounts when one account cannot be read
  • fix: discover the organization's managed accounts
  • fix: discover the organization's managed accounts instead of actor.accounts, which also contains the organization's internal storage account and caused permission errors on workload discovery and failed event delivery
  • fix: do not panic when New Relic answers with a null data, workload or aiIssues payload, as it does for queries the API key's user is not permitted to run
  • fix: fail the incident check instead of reporting "no incidents" when New Relic rejects the incidents query, e.g. because the API key's user lacks the required role
  • fix: log the operation, account and rejected field path for New Relic GraphQL errors, which are reported with HTTP 200 and were previously logged without any hint about which query failed
  • fix: lowercase error strings to satisfy ST1005
  • fix: surface New Relic GraphQL errors instead of silently degrading

extension-prometheus 2.1.27

  • build(deps): bump goreleaser/goreleaser from v2.17.0 to v2.17.1
  • feat: support filtering targets out of discovery

extension-rabbitmq 1.0.21

  • feat: support filtering targets out of discovery
  • fix: emit the node check metric immediately on Start (#42)
  • fix: emit the queue backlog metric immediately on Start (#43)

extension-redis 1.1.10

  • feat: support filtering targets out of discovery
  • fix: emit connection/latency/memory/replication metrics on Start

extension-splunk 1.0.15

  • feat: support filtering targets out of discovery
  • fix: emit detector/SLO check metrics immediately on Start

extension-stackstate 1.0.29

  • Add a "Fail early" option to the service status check. When enabled (the default, matching the previous behavior), the "All the time" mode fails as soon as a deviating status is observed. When disabled, the check keeps collecting events for the whole duration and only fails at the end of the step (with a past-tense message, since the status may have recovered by then). Only affects the "All the time" mode.
  • chore(deps): bump go to 1.26.5 (#150)
  • chore(deps): update dependencies
  • feat(service check): add fail early option (#149)
  • feat: support filtering targets out of discovery
  • fix(e2e): return the error as the last argument in runServiceCheck (ST1008)
  • fix: emit the service status metric immediately on Start (#152)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#151)
  • test(service check): cover the fail early option in unit and e2e tests

Platform 2.8.2

The following changes are part of Steadybit Labs, our program for early-access features. Expect them to evolve based on your feedback.

Labs / New

  • SteadyBuddy is available on every page — SteadyBuddy is now reachable from anywhere in the platform through a chat sidebar that knows which page you are on, greets you with a page-specific opener and shares one conversation with the page you came from. Toggle it with Cmd/Ctrl + .; on wide screens the panel docks into the layout instead of overlaying it.
  • SteadyBuddy analyzes your experiment runs — SteadyBuddy's analysis of a run is now persisted and shown on the run page, runs with a finished analysis are marked in the runs sidebar, and the analysis can hand its investigation over to the chat so you can keep asking follow-up questions.
  • Experiment suggestions on the service page — SteadyBuddy suggests experiments directly on a service page, and suggestions in the chat show which service they belong to.

Labs / Improvement

  • More precise run analysis — SteadyBuddy narrows a metric before reading it and ranks metric series by relevance, reports each step's kind and the parameters it actually ran with, names actions by their catalog name instead of their id, and shows timestamps in your own time zone.
  • Remote MCP covers more of a run — The MCP server exposes a template's step design, reports where a large breakdown was capped instead of silently truncating it, and no longer suggests a narrowing it cannot perform.

Labs / Fix

  • Readable dates in run analysis — Run analyses show real dates instead of raw epoch numbers.
  • Correct explanation for missing actions — The experiment designer no longer reports a missing action as a permission problem when the action simply isn't installed.
  • Binary artifacts — A binary artifact is now reported as such instead of being decoded as text.
  • Chat layout polish — Agent responses use the full chat width, the chat-history popup has a single scrollbar and a pinned header, the search modal opens without the chat panel behind it, and the install banner, experiment-editor zoom controls and sidebar callouts stay clear of the SteadyBuddy bar.

New

  • Public API for Explorer landscape views — Saved landscape views can now be listed, read, created, updated and deleted through the public REST API under /api/explore/landscape/views, closing the gap where saved views existed only in the UI.
  • Duplicate a service profile — Service profiles can be duplicated, so a new profile can start from an existing one instead of from scratch.
  • Filter lists by clicking a tag, service or team — Tags, service pills and shared-team icons in list views are now clickable and filter the list by what you clicked.
  • Explain a service's risk score from the gauge — The risk explanation opens directly from the info icon on the Service Risk gauge.

Improvement

  • Selected items stay in view while searching — In lists that combine checkboxes with a search, the selected items sort to the top, and selections are kept when searching in the link-experiments and hub-import dialogs.
  • Service profile categories keep a consistent order — Categories are ordered consistently and can be reordered by drag and drop.
  • Clearer feedback when an extension comes or goes — Experiments are re-validated when extensions register or de-register, and an unavailable action is explained instead of just flagged.
  • Consistent empty states — Empty states across the platform now share one self-centering component with aligned illustrations and calls to action.
  • Smoother run view — The state-over-time widget no longer re-renders its whole card on every progress tick.

Fix

  • _ and % are literal in query language values — A value containing _ or % now matches literally instead of being treated as a wildcard.
  • Reliable long-press to run — The long-press run button stays alive when the button re-lays out, navigates to the run view even when the button unmounts first, and is disabled while the kill switch is active. The kill-switch tooltip now points to the emergency-stop banner.
  • Metric points attributed to the right target — Widget metric points are attributed to the exact target execution that produced them.
  • Reports and tables — The risk-distribution report shows an empty state when there is no data, an unassigned cell renders as empty instead of None, and table ellipsis, the dashboard's target distribution and the target table's columns were corrected.
  • Form focus fixes — Focus stays on a property list input when its value is cleared, and the Explorer attribute configuration keeps its focus and placeholder title.
  • Kill switch icon on hover — The kill switch activity icon stays visible when hovering its row.
  • Clean shutdown — A stuck tenant-sync job can no longer hang platform shutdown indefinitely.
  • Metric/Span tag verbosity — Outbound HTTP request URLs are no longer recorded verbatim in metrics, traces.

Dependencies

  • Routine dependency and Helm chart updates, including postcss 8.5.23 and the Steadybit UI components library 0.5.4.

Note

  • On-prem: index advisor suggestions name the tenant schema — The DDL the target index advisor logs for an operator to run is now fully schema-qualified, so pasting a suggestion into psql can no longer create or drop an index in the wrong schema.

extension-jvm 1.3.0

  • Add on-demand JVM attack support matrix suite (BM-13107) (#424)
  • chore(deps): bump actions/setup-go from 6 to 7
  • chore(deps): bump actions/setup-go from 6 to 7 (#422)
  • chore(deps): bump actions/setup-python from 6 to 7
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • chore(deps): bump k8s.io/client-go from 0.36.2 to 0.36.3
  • chore(deps): update dependencies
  • fix(agent): support HTTP-client status injection on Spring Framework 6+ (#426)
  • fix(matrix): warm the endpoint before the baseline probe (#427)
  • test(jvm): run the minikube e2e against the in-repo sample (#425)

Agent 2.4.1

Improvement

  • Experiment resiliency - A transient drop on the agent↔platform experiment channel no longer fails the experiment — the agent transparently resumes the conversation. Associated prometheus metric is experiment_rsocket_resume_total. (Requires Platform >=2.8.1)

extension-auto-registration-kubernetes 1.0.11

  • build(deps): bump github.com/steadybit/extension-kit
  • build(deps): bump k8s.io/api from 0.36.2 to 0.36.3
  • build(deps): bump k8s.io/apimachinery from 0.36.2 to 0.36.3
  • build(deps): bump k8s.io/client-go from 0.36.2 to 0.36.3

extension-aws 2.4.22

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/feature/ec2/imds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/applicationautoscaling
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/dynamodb
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/eks
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticloadbalancingv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/fis
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/sqs
  • chore(deps): bump github.com/getkin/kin-openapi from 0.138.0 to 0.144.0
  • chore(deps): bump goreleaser/goreleaser from v2.17.0 to v2.17.1

extension-gatling 1.0.51

  • chore(deps): update dependencies
  • fix: force netty 4.2.16 and jackson 2.21.4 to resolve transitive CVEs

extension-istio 1.0.27

  • chore(deps): bump istio.io/api
  • chore(deps): bump istio.io/client-go from 1.30.2 to 1.30.3

extension-k6 1.3.0

  • chore(deps): update dependencies
  • feat: update bundled k6 to v2.1.0 (#194)
  • fix: force grpc 1.82.1 in k6 binary to resolve GHSA-hrxh-6v49-42gf

extension-kafka 1.2.18

  • chore(deps): update dependencies
  • chore: run CI build/test on free ubuntu-latest runner (#131)

extension-postman 2.0.33

  • chore(deps): update dependencies
  • fix: bump bundled npm to 11.18.0 to resolve tar and brace-expansion CVEs

extension-prometheus 2.1.25

  • build(deps): bump github.com/prometheus/client_golang
  • build(deps): bump github.com/sethvargo/go-retry from 0.3.0 to 0.4.0

Platform 2.8.1

The following changes are part of Steadybit Labs, our program for early-access features. Expect them to evolve based on your feedback.

Labs / New

  • Remote MCP — Connect AI agents such as Claude Code or GitHub Copilot directly to Steadybit to explore, design, analyze, and run experiments. Read More

Improvement

  • Experiments survive brief agent disconnects – A transient drop on the agent↔platform experiment channel no longer fails the experiment; the agent resumes the conversation transparently. (Requires agent 2.4.1 or later)
  • Clearer API error for provided experiments – Trying to update a provided experiment's design through the API now returns a message that explains what actually went wrong.

Fix

  • Consistent target preview for dynamic variable values – The experiment editor no longer shows inconsistent target information when dynamic variable values are used. Random values are sampled once when the experiment is opened, so the preview matches what actually runs. Learn more
  • Agents no longer disconnect on large prepare payloads – RSocket frame fragmentation between agent and platform is now enabled, so keep-alive and Pong frames are no longer blocked behind large, unfragmentable prepare-phase payloads (for example, a pod-count check across 100+ targets) on slower connections.
  • Advice is no longer skipped for negated target predicates – Target post-processing now correctly evaluates advice whose applicable, exclude, or action-required predicates use not contains or unequals, even when the referenced attribute isn't reported as changed.
  • Reliable target descriptions after onboarding – Fixed a transient state in which target descriptions were missing right after onboarding an agent with a longer startup time.
  • Administrators can duplicate experiments into any team – When duplicating an experiment, administrators now see every team they can edit, not just the teams they belong to.
  • Various UI fixes – Tooltips in the experiment editor and run view wrap instead of overflowing, pagination in Settings → Integrations is centered, the onboarding footer is aligned correctly, and the agent count is rounded again.

extension-container 1.7.3

  • fix: network attacks now also affect protocols without ports (e.g. ICMP) when no port is specified. Previously an unset port implied the port range 1-65534, so only port-bearing protocols (TCP/UDP/SCTP) were blocked and ICMP traffic slipped through the blackhole attack.

extension-host 1.7.1

  • fix: network attacks now also affect protocols without ports (e.g. ICMP) when no port is specified. Previously an unset port implied the port range 1-65534, so only port-bearing protocols (TCP/UDP/SCTP) were blocked and ICMP traffic slipped through the blackhole attack.

extension-host-windows 0.3.3

  • fix: WinDivert-based network attacks (delay, blackhole, package loss, package corruption) now also affect protocols without ports (e.g. ICMP) when no port is specified. The filter previously matched only tcp/udp packets, so portless traffic such as ping slipped through the attack. Port-scoped excludes now also only spare their tcp/udp port and no longer spare all ICMP traffic to/from the excluded address.
  • fix: the build information log line (and other early startup logs) is no longer dropped from the on-disk log / Windows Event Log — the log writer is now attached synchronously before startup logging instead of in a background goroutine.

Agent 2.4.0

New

  • OOM-killer protection – The agent lowers its own oom_score_adj so the Linux OOM killer targets other processes first during memory-pressure experiments.
  • Compressed platform connection – RSocket permessage-deflate and gzipped HTTP responses reduce the bandwidth used between agent and platform.

Improvement

  • Dedicated heartbeat RSocket stream for experiment connections keeps them alive even under send-path congestion.
  • Extension kit index fetching now uses ETags, avoiding re-downloads of unchanged indexes.
  • Enabled additional discovery and attacks for extension-gcp.

Fix

  • Pending batched messages (e.g. action started/stopped events) are flushed before a pong timeout closes the connection, instead of being silently dropped.
  • Action/preflight state checkpoint saving is now best-effort — a failed save no longer fails the execution.
  • Windows: fixed extension registration and installer versioning, including a registry-watch race that could miss extension changes and a busy-spinning registry watcher.

Security

  • Prevented secret leakage in agent logs; the agent key also no longer appears in verbose MSI install logs.
  • The bundled JRE download is verified against a pinned SHA-256 during the build.

Dependencies

  • Bundled runtime JRE updated to Zulu Java 25.

Platform 2.8.0

The following changes are part of Steadybit Labs, our program for early-access features. Expect them to evolve based on your feedback.

Labs / New

  • SteadyBuddy is now service-aware — SteadyBuddy can suggest experiments for a service directly in the chat, including the service's target scope, validations, and custom properties.

Labs / Fix

  • SteadyBuddy's loading behavior for older messages — Fixed how SteadyBuddy loads and displays older chat messages.
  • SteadyBuddy's Environment selection scales properly — Environment selection now works reliably even with a large number of environments.

New

  • Experiment variables support fixed multi-values — A fixed variable can now hold multiple comma-separated strings, expanding into an attr IN ({{var}}) expression. Read more

Improvement

  • Compressed platform connection — RSocket traffic between the agent and the platform is now compressed.
  • Faster target ingestion — Enrichments are skipped when a target snapshot is unchanged, and attribute removal is skipped when only copied (not joined) attributes are updated.

Fix

  • Heartbeat resilience — Heartbeats are now decoupled from experiment messages so they can't be blocked.
  • Experiment editor's quick switcher preserves step config — Re-selecting the currently selected action in the experiment editor no longer wipes its step parameters.
  • Experiment template placeholders work with CPU core UI control — Fixed an issue with the template placeholder for CPU core UI controls.
  • Service template hub imports — Experiment templates from Steadybit's Service templates now sync automatically and consistently; manual sync is no longer needed.
  • Uploading experiments with properties — Uploaded properties now receive experiment-scoped associations, with association alignment safely skipped if the experiment no longer exists

extension-appdynamics 1.1.14

  • Add a "Fail early" option to the health rule check. When enabled (the default, matching the previous behavior), the check fails as soon as a deviating state is observed. When disabled, the check keeps collecting events for the whole duration and only fails at the end of the step (with a past-tense message, since the state may have recovered by then).
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#71)
  • chore: add Claude Code workflows (#63)
  • chore: normalize dependabot-auto-merge workflow to the standard version (#66)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(health rule check): add fail early option (#65)
  • fix: actually apply and lift AppDynamics action suppression (#64)
  • fix: guard the health-rule check against missing target attributes instead of panicking, and avoid a possible nil-dereference when an AppDynamics API request fails before a response is received
  • fix: the "Create Action Suppression" action now captures the created suppression id, so it is actually deleted again on stop instead of leaving AppDynamics alerting suppressed indefinitely
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#72)

extension-auto-registration-ecs 1.0.15

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • ci: skip build on .trivyignore.yml-only changes [skip ci]

extension-aws 2.4.21

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigatewayv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/autoscaling
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/dynamodb
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticache
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/kafka
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/lambda
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/rds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/sts
  • ci: skip build on .trivyignore.yml-only changes [skip ci]

extension-container 1.7.2

  • build(deps): bump github.com/containerd/containerd from 1.7.33 to 1.7.34
  • chore: run e2e tests on free ubuntu-latest runner (#479)
  • ci: skip build on .trivyignore.yml-only changes [skip ci]

extension-dynatrace 1.0.25

  • Add a "Fail early" option to the problem check. When enabled (the default, matching the previous behavior), the check fails as soon as the condition is violated. When disabled, the check keeps collecting events for the whole duration and only fails at the end of the step. When failing at the end, the message uses past tense ("were found during the step") since the condition may have recovered by then.
  • build(deps): bump github.com/jellydator/ttlcache/v3 from 3.4.0 to 3.4.1
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • build(deps): bump github.com/steadybit/extension-kit
  • build(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore(deps): bump go to 1.26.5 (#144)
  • chore: add Claude Code workflows (#134)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(problem check): add fail early option (#137)
  • fix(problem check): use past tense for fail-at-end message (#138)
  • fix: URL-escape the entitySelector when querying Dynatrace problems, preventing query-parameter injection into the Dynatrace API (matches the existing escaping in the entities query)
  • fix: escape entitySelector in Dynatrace problems query (#135)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#145)

extension-host 1.7.0

  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat: keep the host alive during fill_mem (reserve, adaptive, oom_score_adj) (#234)
  • fix: fill-disk and fill-memory no longer crash the extension during Prepare when a config parameter is missing or has an unexpected type; config values are now read via the tolerant extutil helpers.
  • fix: network attacks no longer crash the extension during Prepare when the request omits executionContext (nil pointer dereference in mapToExecutionContext).
  • fix: prevent stop-process crash when a matched process exits early (#235)
  • fix: the stop-process attack no longer crashes the extension with a nil pointer dereference when a matched process exits before it is stopped (ps.FindProcess returns nil, nil for a vanished PID on Linux).

extension-host-windows 0.3.2

  • build(deps): bump actions/setup-go from 6 to 7
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • ci: skip build on .trivyignore.yml-only changes [skip ci]

extension-http 1.0.47

  • fix: prevent HTTP check stop from deadlocking when prepared but never started
  • fix: stop the HTTP check no longer deadlocks when an action is prepared but stopped without being started; workers now honor context cancellation

extension-jvm 1.2.20

  • chore: run audit job on free ubuntu-latest runner (#420)
  • ci: skip build on .trivyignore.yml-only changes [skip ci]

extension-kong 2.0.27

  • build(deps): bump github.com/kong/go-kong from 0.76.1 to 0.77.0
  • ci: skip build on .trivyignore.yml-only changes [skip ci]

extension-redis 1.1.8

  • Add a "Fail early" option to the connection count, latency, memory and replication checks. When enabled, the check fails as soon as its threshold is exceeded instead of waiting for the end of the step. Disabled by default, matching the previous behavior of only reporting a threshold breach at the end.
  • chore(deps): bump go to 1.26.5 (#40)
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(checks): add fail early option (#39)
  • fix(checks): use internal time control so breaches fail at the end (#42)
  • fix: the connection count, latency, memory and replication checks now use internal time control so a threshold breach is reliably reported at the end of the step. Previously they used external time control without a stop handler, so the end-of-step failure was never emitted and a breach only produced a warning (the check completed successfully even though its threshold was exceeded).
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#41)

extension-splunk 1.0.13

  • Add a "Fail early" option to the detector and SLO checks. When enabled (the default, matching the previous behavior), the check fails as soon as a deviating state is observed. When disabled, the check keeps collecting events for the whole duration and only fails at the end of the step (with a past-tense message, since the state may have recovered by then).
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • build(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#62)
  • chore: add Claude Code workflows (#55)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(detector & SLO checks): add fail early option (#57)
  • fix: guard detector/SLO checks against a missing name attribute
  • fix: guard the detector and SLO checks against targets missing the name attribute instead of panicking
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#63)

extension-splunk-platform 1.0.13

  • Add a "Fail early" option to the alert status check. When enabled (the default, matching the previous behavior), the "All the time" mode fails as soon as a deviating state is observed. When disabled, the check keeps collecting events for the whole duration and only fails at the end of the step. Only affects the "All the time" mode.
  • Fix link in README.md (#43)
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#41)
  • chore: add Claude Code workflows (#36)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(alert check): add fail early option (#44)
  • fix: guard alert check attributes and fix unbounded alert paging
  • fix: guard the alert check against targets missing the name/url attributes instead of panicking
  • fix: terminate alert paging on the returned page size instead of the server-reported total, preventing an infinite request loop (and dropped results) when Splunk reports an inaccurate total
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#42)

extension-datadog 1.8.23

  • build(deps): bump github.com/DataDog/datadog-api-client-go/v2
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • build(deps): bump github.com/steadybit/extension-kit
  • build(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore(deps): bump go to 1.26.5 (#224)
  • chore: add Claude Code workflows (#215)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(monitor status check): add fail early option (#211)
  • fix(monitor status check): fail when monitor status is unknown (#216)
  • fix(monitor status check): use past tense for fail-at-end message (#217)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#225)

extension-grafana 1.1.5

  • Add a "Fail early" option to the alert rule check. When enabled (the default, matching the previous behavior), the check fails as soon as a deviating state is observed. When disabled, the check keeps collecting events for the whole duration and only fails at the end of the step.
  • chore(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • chore(deps): bump go to 1.26.5 (#94)
  • chore(deps): bump go-openapi/swag/loading to fix go mod tidy (#96)
  • chore: add Claude Code workflows (#92)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(alert rule check): add fail early option (#93)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#95)

extension-http 1.0.46

  • Add a "Fail early" option to the HTTP checks (Requests/s, Fixed number of Requests, and Bandwidth). When enabled, the check fails as soon as enough requests (or measurement windows) have failed that the required success rate can no longer be reached, instead of waiting for the end of the step. Disabled by default, matching the previous behavior of evaluating the success rate only at the end.
  • build(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#175)
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(http checks): add fail early option (#170)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#176)

extension-kafka 1.2.17

  • Add a "Fail early" option to the broker, consumer group, partition and topic lag checks. When enabled, the check fails as soon as a deviating event is observed; when disabled, it keeps collecting events for the whole duration and only fails at the end of the step. The broker/consumer-group/partition checks default to fail-early (matching their previous behavior); the topic lag check defaults to fail-at-end (matching its previous behavior). Only affects the "All the time" mode of the mode-based checks.
  • Merge pull request #123 from steadybit/feat/check-fail-early-option
  • build(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • build(deps): bump github.com/twmb/franz-go from 1.21.3 to 1.21.5
  • build(deps): bump golang.org/x/crypto in /test-dataset/dummyconsumer
  • chore(deps): bump go to 1.26.5 (#125)
  • chore(deps): bump go-openapi/swag/loading to fix go mod tidy (#128)
  • chore(deps): update dependencies
  • chore: add Claude Code workflows (#122)
  • chore: normalize dependabot-auto-merge workflow to the standard version (#124)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#127)

extension-kubernetes 2.6.30

  • Add a "Fail early" option to the pod count check. When enabled (the default, matching the previous behavior), the "All the time" mode fails as soon as the pod count condition is violated. When disabled, the check keeps collecting events for the whole duration and only fails at the end of the step. Only affects the "All the time" mode; "At least once" is unaffected.
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump github.com/steadybit/extension-kit
  • build(deps): bump golang.org/x/text from 0.38.0 to 0.39.0
  • build(deps): bump golang.org/x/text from 0.39.0 to 0.40.0
  • chore(deps): bump go to 1.26.5 (#336)
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(pod count check): add fail early option (#335)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#337)

extension-rabbitmq 1.0.18

  • Add a "Fail early" option to the node check and the queue backlog check. When enabled, the check fails as soon as a deviating event is observed (node check: a deviating change; queue backlog check: the backlog exceeding the threshold), instead of waiting for the end of the step. The node check defaults to fail-early (matching its previous "All the time" behavior); the queue backlog check defaults to fail-at-end (matching its previous behavior). The node check option only affects the "All the time" mode.
  • chore(deps): bump go to 1.26.5 (#39)
  • ci: skip build on .trivyignore.yml-only changes [skip ci]
  • feat(checks): add fail early option (#38)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#40)

extension-auto-registration-ecs 1.0.14

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • build(deps): bump github.com/steadybit/extension-kit
  • chore: update Go to 1.26.5
  • refactor: apply go fix modernizations

extension-aws 2.4.20

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigatewayv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/applicationautoscaling
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/fis
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/lambda
  • chore(deps): bump go to 1.26.5 (#907)
  • chore(deps): bump golang.org/x/crypto from 0.51.0 to 0.52.0
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#909)

extension-azure 1.3.4

  • chore(deps): bump go to 1.26.5 (#207)
  • chore(deps): bump go-openapi/swag/loading to fix go mod tidy (#209)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#208)

extension-container 1.7.1

  • build(deps): bump golang.org/x/sync from 0.21.0 to 0.22.0
  • chore: update dns-inject to v0.2.3 (#475)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#474)

extension-gatling 1.0.49

  • chore(deps): bump go to 1.26.5 (#142)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#143)

extension-gcp 1.0.30

  • chore(deps): bump github.com/googleapis/gax-go/v2 from 2.22.0 to 2.23.0
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#357)
  • chore(deps): bump google.golang.org/api from 0.285.0 to 0.286.0
  • chore(deps): bump google.golang.org/api from 0.286.0 to 0.287.0
  • chore(deps): bump google.golang.org/api from 0.287.1 to 0.288.0
  • chore(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore: add Claude Code workflows (#350)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#358)

extension-host 1.6.1

  • build(deps): bump golang.org/x/sync from 0.21.0 to 0.22.0
  • chore: update dns-inject to v0.2.3 (#232)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#231)

extension-host-windows 0.3.1

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • build(deps): bump golang.org/x/sys from 0.46.0 to 0.47.0
  • build(deps): bump softprops/action-gh-release from 2.6.1 to 3.0.1
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#188)

extension-instana 1.1.19

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#89)
  • chore: add Claude Code workflows (#84)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: URL-escape values interpolated into Instana API requests — the maintenance-window id (derived from the experiment key) is path-escaped and the application/event query parameters are query-escaped, preventing path traversal and query-parameter injection
  • fix: escape Instana API URL values and guard handler panics (#85)
  • fix: guard the event check and maintenance-window actions against missing target attributes and non-numeric duration config instead of panicking
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#90)

extension-istio 1.0.25

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#311)
  • chore(deps): bump istio.io/api from 1.30.1 to 1.30.2
  • chore(deps): bump istio.io/client-go from 1.30.1 to 1.30.2
  • chore: add Claude Code workflows (#305)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#312)

extension-jenkins 1.0.16

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump go to 1.26.5 (#35)
  • chore: add Claude Code workflows (#31)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#36)

extension-jmeter 1.0.40

  • chore(deps): bump go to 1.26.5 (#142)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#143)

extension-jvm 1.2.19

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • chore(deps): bump go to 1.26.5 (#415)
  • chore(deps): bump golang.org/x/net from 0.56.0 to 0.57.0
  • chore(deps): bump golang.org/x/sys from 0.46.0 to 0.47.0
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#416)

extension-k6 1.2.8

  • chore(deps): bump go to 1.26.5 (#192)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#193)

extension-kong 2.0.26

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump github.com/steadybit/extension-kit
  • build(deps): bump github.com/testcontainers/testcontainers-go
  • build(deps): bump golang.org/x/crypto from 0.51.0 to 0.52.0
  • build(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore(deps): bump go to 1.26.5 (#254)
  • chore: add Claude Code workflows (#248)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: don't panic when matching Kong routes whose optional name (or id) is unset — FindRoute now nil-guards the comparisons (and the not-found error no longer dereferences an unset service name)
  • fix: guard nil route name/id when matching Kong routes (#249)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#256)

extension-newrelic 1.0.21

  • chore(deps): bump go to 1.26.5 (#99)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#100)

extension-postman 2.0.31

  • chore(deps): bump go to 1.26.5 (#144)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#145)

extension-prometheus 2.1.23

  • build(deps): bump github.com/prometheus/common from 0.69.0 to 0.70.0
  • build(deps): bump golang.org/x/crypto from 0.51.0 to 0.52.0
  • chore(deps): bump go to 1.26.5 (#284)
  • refactor: register extension index via exthttp.RegisterRevisionedHandler (#286)

Agent 2.3.10

Improvement

  • Extension communication now survives DNS outages during network attacks. Running Block Traffic / Block DNS on the node hosting the agent no longer breaks the agent’s calls to extensions once the DNS cache expires (previously caused action-start errors, failed checks, and discovered targets disappearing mid-attack).
  • Added metrics for rollback failures.

Fix

  • Security: Agent HTTP endpoints now default to loopback-only access. Debug, inventory, and other internal endpoints are no longer reachable remotely by default, except /health and /prometheus.
  • A single malformed artifact, metric, or log message no longer fails the whole action — the bad item is skipped, the action continues.
  • Preflight status polls canceled while executions are stopping are no longer reported as errors.

extension-container 1.7.0

  • feat: new Exclude Hostnames (excludeHostname) and Exclude IPs/CIDRs (excludeIp) parameters on all network attacks sharing the hostname/IP/port filters (delay, loss, corruption, bandwidth, blackhole, TCP reset) — affect all traffic except the given hosts or IPs/CIDRs. Excludes always take precedence over the include restrictions. The existing filter parameters are relabeled to Include Hostnames, Include IPs/CIDRs and Include Ports to make the distinction explicit.
  • fix: dedupe network attacks across containers that share a pod's network namespace. In Kubernetes every container in a pod shares the pod's netns, so an experiment targeting multiple containers in the same pod fired multiple tc applies against the same netns — the second apply then collided at the kernel level (e.g. htb.change rejected because the first apply had installed active classes at handle 1:). The first Start on a netns now applies the attack ("primary"), later Starts on the same netns no-op ("shadow"), and Stop mirrors the same split.
  • fix: strip the per-target TargetExecutionId nonce from the serialized opts before the multi-container dedup comparison. Steadybit fires one action per container target, so sibling containers of the same pod running the same experiment had different state.NetworkOpts bytes — the previous byte-identical comparison never matched siblings and every one of them fell through to Passthrough, reproducing the Change operation not supported collision the dedup was supposed to prevent. ExperimentExecutionId is deliberately kept in the comparison so two unrelated experiments running the same attack on the same pod stay independent (otherwise one experiment's Stop would tear down the other's still-active tc state).
  • Update Go to 1.26.5
  • Update dependencies

extension-host 1.6.0

  • feat: new Exclude Hostnames (excludeHostname) and Exclude IPs/CIDRs (excludeIp) parameters on all network attacks sharing the hostname/IP/port filters (delay, loss, corruption, bandwidth, blackhole, TCP reset) — affect all traffic except the given hosts or IPs/CIDRs. Excludes always take precedence over the include restrictions. The existing filter parameters are relabeled to Include Hostnames, Include IPs/CIDRs and Include Ports to make the distinction explicit.
  • fix: update action_kit_commons to v1.10.2 — stopping a network attack no longer fails with restore qdisc fq_codel on <iface>: netlink receive: no such file or directory on stock multi-queue NICs (e.g. AWS ENA ens5), where the kernel attaches the default mq 0: tree anonymously. The qdisc snapshot restore now skips those kernel-default children (handle 0 is unaddressable via RTNETLINK; the kernel re-creates them identically anyway).
  • Update Go to 1.26.5
  • Update dependencies

extension-host-windows 0.3.0

  • feat: new Exclude Hostnames (excludeHostname) and Exclude IPs/CIDRs (excludeIp) parameters on the WinDivert-based network attacks (delay, blackhole, package loss, package corruption) — affect all traffic except the given hosts or IPs/CIDRs. Excludes always take precedence over the include restrictions. The existing filter parameters are relabeled to Include Hostnames, Include IPs/CIDRs and Include Ports to make the distinction explicit. The bandwidth attack is not covered: Windows QoS policies only support include-style match conditions.
  • Update Go to 1.26.5
  • Update dependencies

Platform 2.7.2

The following changes are part of Steadybit Labs, our program for early-access features. Expect them to evolve based on your feedback.

Labs / New

  • SteadyBuddy chat history — SteadyBuddy now supports chat history Learn more

Labs / Improvement

  • Context-aware SteadyBuddy — Improved context-awareness for more natural, time-saving conversations

New

  • Service-scoped dynamic variables — Dynamic variable values on a service can now be scoped to the service's target scope instead of the environment scope Learn more

Improvement

  • API of services now support variables — API upsert and get of services now include service variables
  • Service validations in experiments — Services used in experiments now show their list of service validations
  • Experiment variable overrides in provided experiments — Provided experiments of a service support shadowing variables with experiment variables
  • Advice query validation — Advice queries built with AdviceKit are now validated for query-language correctness before saving

Fix

  • Experiments allow overriding variables for a run — Fixed an issue where variables could not be overridden for a run when the experiment hadn't been executed yet
  • Timezone issue in reporting — Fixed reporting chart date buckets, drill-downs, and the custom date-range picker using the wrong timezone, which caused monthly/daily labels to shift by a day
  • Validation results in run preview — Fixed experiment run previews in variable override or schedule not showing validation results of the overrides
  • Service's optimistic versioning conflict — Fixed an issue where saving a service via the UI could incorrectly report a version conflict when no concurrent change had occurred
  • Service profile category limit — The service profile category limit is now correctly enforced at 5, showing a validation error instead of failing silently
  • Reporting of experiments in a deleted environment — Fixed a reporting issue when an experiment was saved for an environment that was deleted in between

extension-aws 2.4.19

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/config
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/feature/ec2/imds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigateway
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/dynamodb
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/eks
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticache
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticloadbalancingv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/kafka
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/mq
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ssm
  • chore(deps): bump github.com/aws/smithy-go from 1.27.1 to 1.27.3
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore: add Claude Code workflows (#889)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: guard the ECS SSM heartbeat map with a mutex to prevent a concurrent map writes crash when multiple ECS task attacks run concurrently
  • fix: prevent concurrent map writes crash in ECS SSM heartbeat (#890)

extension-azure 1.3.3

  • chore(deps): bump github.com/Azure/azure-sdk-for-go/sdk/azidentity
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore(deps): update dependencies
  • chore: add Claude Code workflows (#204)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: close the timeout-orphan window in the NSG block-traffic attack
  • fix: don't panic in the NSG block-traffic attack when the security-rule creation poll fails (the rule name was dereferenced before the error was checked), and record the created rule (under its deterministic name, before waiting for the operation) so it is cleaned up even if the create times out after the rule was already applied; cleanup now tolerates an already-removed rule
  • fix: prevent panic and orphaned rules in NSG block-traffic attack
  • fix: return a clear error instead of panicking when the block-traffic 'hosts' configuration is missing or not a list

extension-gatling 1.0.48

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump github.com/steadybit/extension-kit
  • chore: add Claude Code workflows (#136)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: resolve data race on the Gatling process exit code (#138)
  • fix: resolve the data race on the Gatling process exit code between the process-reaping goroutine and the status/stop handlers
  • fix: set a timeout on the Gatling Enterprise HTTP client (#137)
  • fix: set a timeout on the Gatling Enterprise HTTP client so a slow or unresponsive API cannot block discovery, status checks or the run lifecycle indefinitely

extension-http 1.0.45

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore: add Claude Code workflows (#168)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: bandwidth check no longer fails when measured throughput is healthy
  • fix: cancel bandwidth requests on stop and reject maxConcurrent of 0
  • fix: cancel in-flight bandwidth-check requests on stop so workers blocked on a slow or stalled endpoint no longer leak their goroutine and connection
  • fix: reject a maxConcurrent of 0 in the HTTP check actions instead of deadlocking the request scheduler

extension-jmeter 1.0.39

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore: add Claude Code workflows (#137)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: resolve data race on the process exit code (#138)
  • fix: resolve the data race on the JMeter process exit code between the process-reaping goroutine and the status/stop handlers (via extcmd.CmdState.Wait/ExitCode)

extension-jvm 1.2.18

  • chore(deps): bump actions/checkout from 6 to 7
  • chore(deps): bump github.com/shirou/gopsutil/v4 from 4.26.5 to 4.26.6
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump golang.org/x/net from 0.55.0 to 0.56.0
  • chore(deps): bump k8s.io/api from 0.36.1 to 0.36.2
  • chore(deps): bump k8s.io/apimachinery from 0.36.1 to 0.36.2
  • chore(deps): bump k8s.io/client-go from 0.36.1 to 0.36.2
  • chore(deps): runc 1.4.3
  • chore: add Claude Code workflows (#406)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • feat: lower oom_score_adj on startup via extension-kit's extruntime.AdjustOOMScoreAdj() to avoid being killed by the node OOM killer. The extension sets it directly using the cap_sys_resource file capability (default -998, configurable via STEADYBIT_EXTENSION_OOM_SCORE_ADJ).
  • feat: lower oom_score_adj on startup via extension-kit (#402)
  • fix(chart): make javaagent jars readable across SELinux MCS boundaries on OpenShift (#409)
  • fix(chart): on OpenShift, run the extension pod with an MCS-category-less SELinux level (s0) so target JVMs can read the mounted javaagent jars. Without this, SELinux denies the agent jar read (each namespace gets distinct MCS categories) and attacks fail with "connection not found". The SCC pins the level via seLinuxContext: MustRunAs; override via podSecurityContext.seLinuxOptions (mirrored into the SCC).
  • fix: missing-circuit-breaker and missing-timeout accumulated wrong entries (#405)
  • fix: propagate the underlying error when stopping an action fails, instead of returning a bare "Failed to stop action"
  • fix: propagate underlying error when stopping an action fails (#408)

extension-k6 1.2.7

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore: add Claude Code workflows (#185)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: avoid panic and token disclosure when starting a K6 cloud run
  • fix: don't panic when starting a K6 cloud run with an API token shorter than 5 characters, and stop logging the last characters of the cloud API token
  • fix: resolve data race on the process exit code (#187)
  • fix: resolve the data race on the K6 process exit code between the process-reaping goroutine and the status/stop handlers (via extcmd.CmdState.Wait/ExitCode)

extension-newrelic 1.0.20

  • Merge pull request #93 from steadybit/feat/add-claude-workflows
  • chore(deps): bump github.com/jellydator/ttlcache/v3 from 3.4.0 to 3.4.1
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: do not panic when delivering events while New Relic accounts could not be loaded, or when an incident has no description
  • fix: escape values interpolated into New Relic GraphQL queries and build the request envelope with a JSON encoder, preventing query/JSON injection via muting-rule name/description, workload/entity guids and incident priorities
  • fix: prevent GraphQL injection and handler panics in New Relic calls (#94)
  • fix: set a timeout on the New Relic HTTP client so a slow or unresponsive API cannot block discovery, status checks or event delivery indefinitely

extension-postman 2.0.30

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump github.com/steadybit/extension-kit
  • chore: add Claude Code workflows (#137)
  • chore: normalize dependabot-auto-merge workflow to the standard version (#140)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: authenticate against the Postman API with the X-API-Key header and download the collection/environment in-process instead of embedding the API key in the newman command line and serialized action state (prevents leaking the key via ps//proc/<pid>/cmdline)
  • fix: resolve data race on the process exit code (#139)
  • fix: resolve the data race on the newman process exit code between the process-reaping goroutine and the status/stop handlers (via extcmd.CmdState.Wait/ExitCode)
  • fix: stop leaking Postman API key via newman command line (#138)
  • fix: write newman reports into a unique per-execution working directory (mode 0700) and remove it on stop, instead of timestamped world-readable files in /tmp that could collide between concurrent runs and were never cleaned up

extension-prometheus 2.1.22

  • build(deps): bump github.com/moby/moby/api from 1.54.2 to 1.55.0
  • build(deps): bump github.com/prometheus/common from 0.68.1 to 0.69.0
  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • build(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • build(deps): bump github.com/steadybit/extension-kit
  • build(deps): bump github.com/testcontainers/testcontainers-go
  • build(deps): bump goreleaser/goreleaser from v2.16.0 to v2.17.0
  • chore: add Claude Code workflows (#278)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: guard non-string PromQL query parameter
  • fix: return a clear error instead of panicking when the PromQL query parameter is not a string

extension-rabbitmq 1.0.17

  • Merge pull request #32 from steadybit/feat/add-claude-workflows
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: guard the publish attack's jobs channel against being closed twice when stop runs concurrently/twice, which could panic the extension
  • fix: prevent double-close panic on the publish jobs channel (#33)

extension-redis 1.1.7

  • Merge pull request #31 from steadybit/feat/add-claude-workflows
  • chore(deps): bump github.com/redis/go-redis/v9 from 9.20.1 to 9.21.0
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/event-kit/go/event_kit_api
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: avoid duplicating the node address in cluster restore errors
  • fix: report a failed stop for the maxmemory-limit attack when the cluster cannot be reached during restore, instead of silently reporting success while the target's maxmemory is left altered
  • fix: report failed maxmemory restore when the cluster is unreachable
  • fix: stop leaking Redis credentials embedded in endpoint URLs (#32)
  • fix: strip credentials from Redis endpoint URLs before publishing them as target attributes/metric labels and before logging them, so a password embedded in a redis://user:pass@host URL is no longer exposed to the platform or logs (the full credentials remain in the endpoint configuration and are still used to connect)
  • refactor: collapse maxmemory restore onto a single error channel

extension-stackstate 1.0.28

  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_sdk
  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore: add Claude Code workflows (#144)
  • chore: silence SonarQube finding on secrets: inherit in Claude workflows
  • fix: escape the service id used in the StackState snapshot query and build the request body with a JSON encoder, preventing STQL/JSON query injection
  • fix: guard the service check and discovery against missing components, identifiers, short base URLs and unexpected identifier formats instead of panicking, and avoid a possible nil-dereference when a StackState request fails before a response is received
  • fix: prevent STQL injection and reachable panics in service check/discovery (#145)

extension-container 1.6.10

  • feat: opt-in qdisc snapshot/restore for network attacks. Set STEADYBIT_EXTENSION_NETWORK_STRICT_ROOT_QDISC=false (e.g. via extraEnv) to make Apply capture the root qdisc tree (qdiscs + filters) of every target interface and Revert replay it after the attack's tc del. Preserves cloud-tuned root qdiscs (e.g. GKE's mq + fq with buckets=32768 horizon=2s) that would otherwise revert to kernel defaults after tc qdisc del root and leave the host network degraded until reboot. Off by default; Linux only.
  • The pre-attack qdisc snapshot lives in the action's per-execution state instead of an in-memory map in the extension process. An extension pod restart between Start and Stop no longer loses the snapshot, so Stop still restores the cloud-tuned root tree.
  • Update dependencies

extension-host 1.5.10

  • feat: opt-in qdisc snapshot/restore for network attacks. Set STEADYBIT_EXTENSION_NETWORK_STRICT_ROOT_QDISC=false to make Apply capture the root qdisc tree (qdiscs + filters) of every target interface and Revert replay it after the attack's tc del. Preserves cloud-tuned root qdiscs (e.g. GKE's mq + fq with buckets=32768 horizon=2s) that would otherwise revert to kernel defaults after tc qdisc del root and leave the host network degraded until reboot. Off by default; Linux only.
  • The pre-attack qdisc snapshot lives in the action's per-execution state instead of an in-memory map in the extension process. An extension pod restart between Start and Stop no longer loses the snapshot, so Stop still restores the cloud-tuned root tree.
  • Update dependencies

extension-kubernetes 2.6.29

  • feat: add flags to independently disable namespace/node/pod label inheritance during discovery
  • Update dependencies

extension-kubernetes 2.6.28

  • feat: Envoy Gateway support (opt-in, disabled by default via discovery.disabled.envoyGateway). Discovers HTTPRoutes served by an Envoy Gateway GatewayClass and adds two attacks that apply an Envoy Gateway BackendTrafficPolicy for the attack duration: "Envoy Delay Traffic" and "Envoy Abort Traffic" (the latter can optionally overwrite the response body). Attacks refuse to run when another BackendTrafficPolicy already targets the route (Envoy Gateway resolves conflicts oldest-wins).
  • fix: sort multi-value discovery attributes (k8s.container.id, k8s.pod.name, k8s.namespace, k8s.replicaset, k8s.deployment, k8s.daemonset, k8s.statefulset, k8s.service.name) before attaching them to targets, so unrelated Go map/list iteration order no longer registers as a spurious attribute change on every discovery cycle
  • fix: validate the crash-loop signal parameter and pass it to the fallback shell as a positional argument, preventing command injection into the target pod (the signal option list is a UI hint only and was not enforced)
  • fix: reject control characters in ingress request-matcher conditions (path/method/header), preventing injection of additional directives into the nginx/HAProxy ingress controller configuration

Agent 2.3.9

Dependencies

  • Dependency Updates with CVE fixes

New

  • Add option to skip verification of TLS certificates - only for testing purposes!

Platform 2.7.0

The following changes are part of Steadybit Labs, our program for early-access features. Expect them to evolve based on your feedback.

Labs / New

  • Introducing SteadyBuddy – your AI companion for automating reliability work: suggesting experiments, creating them with natural language, and analyzing why a run failed. More features to come. SteadyBuddy is part of Steadybit Labs, available upon request. Learn More

Improvement

  • Experiment's tags and variables are now available in event-kit, preflight-kit, and webhooks
  • Kubernetes agent installation now supports OpenShift and cloud environments as well as configuration options for enterprise contexts.

extension-auto-registration-ecs 1.0.12

  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • build(deps): bump github.com/steadybit/extension-kit

extension-kubernetes 2.6.27

  • feat: shrink HPA/PDB rollup to boolean flags only (k8s.specification.has-hpa / has-pdb) — drops the detailed multi-valued attributes that caused per-cycle platform DB churn on large clusters
  • feat: expose liveness/readiness HTTP probe paths as discovery attributes
  • feat: add "Status Check Mode" (at least once / all the time) to the Deployment, StatefulSet, DaemonSet and ReplicaSet Pod Count Check. Defaults to "at least once" to keep the existing behavior (backward-compatible).
  • feat: add "ready count = 0" pod count check mode to validate that pods are scaled down.
  • The "Timeout" parameter of the Pod Count Check is now labeled "Duration" (label-only change, backward-compatible).
  • feat: Deployment and StatefulSet Pod Count Checks now emit pod-count metrics and show the readiness widget alongside the check timeline.
  • fix: always emit first and final metric points per check run to capture initial state and close widget gaps
  • fix: only emit pod-count metrics when values change to avoid tiny bars in the widget
  • Update dependencies

extension-container 1.6.9

  • build(deps): bump actions/checkout from 6 to 7
  • chore(deps): runc 1.4.3 and dns-inject to v0.2.2
  • chore(deps): update dependencies
  • feat: lower oom_score_adj on startup via extension-kit (#456)
  • fix: switch back to use strict root qdisc checks

extension-host 1.5.9

  • chore(deps): runc 1.4.3 and dns-inject to v0.2.2
  • feat: set oom_score_adj directly via extension-kit (drop root subprocess) (#216)
  • fix: switch back to use strict root qdisc checks

extension-auto-registration-ecs 1.0.11

  • build(deps): bump alpine from 3.23 to 3.24
  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs

extension-aws 2.4.18

  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigateway
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/dynamodb
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/eks
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/kafka
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/mq
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump github.com/testcontainers/testcontainers-go/modules/localstack
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#878)

extension-cloudfoundry 1.0.4

  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#11)

extension-datadog 1.8.22

  • build(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#208)

extension-dynatrace 1.0.24

  • build(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#130)

extension-gcp 1.0.29

  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump google.golang.org/api from 0.284.0 to 0.285.0

extension-grafana 1.1.4

  • chore(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#88)
  • chore(deps): update dependencies
  • fix: write ended_time tag on patch and keep timestamp tags untruncated

extension-host-windows 0.2.14

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#169)

extension-http 1.0.44

  • build(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#162)

extension-instana 1.1.18

  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#80)

extension-istio 1.0.24

  • chore(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#298)
  • chore(deps): bump k8s.io/client-go from 0.36.1 to 0.36.2

extension-k6 1.2.6

  • chore(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#179)
  • chore(deps): bump k8s.io/client-go from 0.36.1 to 0.36.2

extension-kafka 1.2.16

  • build(deps): bump actions/checkout from 6 to 7
  • build(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#117)

extension-kong 2.0.25

  • build(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#242)

extension-newrelic 1.0.19

  • chore(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#87)

extension-postman 2.0.29

  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#132)
  • chore(deps): update npm in container (#134)

extension-rabbitmq 1.0.16

  • chore(deps): bump github.com/rabbitmq/amqp091-go from 1.11.0 to 1.12.0
  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#27)

extension-redis 1.1.6

  • chore(deps): bump github.com/steadybit/extension-kit
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#27)

extension-container 1.6.8

  • Network attacks (delay, loss, corruption, bandwidth) on hostNetwork: true pods or on containers whose eth0 already has a kernel-default root qdisc no longer fail with NLM_F_REPLACE needed to override. The root qdisc is now installed via tc qdisc replace; on revert the kernel restores its default (mq, noqueue, fq_codel, pfifo_fast, fq).
  • If the target interface carries a user- or CNI-installed root qdisc (e.g. htb, cake) that cannot be restored afterwards, the attack now fails fast in the prepare step with a clear error instead of silently replacing it.
  • Optional fallback: set STEADYBIT_EXTENSION_NETWORK_STRICT_ROOT_QDISC=true (e.g. via extraEnv) to make network attacks refuse any interface whose root qdisc is not noqueue — including the kernel default mq — instead of replacing it. Off by default.
  • New privileged chart value (default false): runs the extension in privileged mode and switches the managed SecurityContextConstraint to allow it. Needed on hardened nodes (e.g. CIS/STIG) where the container root filesystem is mounted nosuid, which voids the binary's file capabilities and breaks fault injection (nsenter: operation not permitted).
  • Stress CPU with "All cores" on uncapped containers now uses every online CPU on hosts with more than 32 cores (previously capped at 32 due to a Cpus_allowed mask parsing bug).

extension-host 1.5.8

  • feat: opt-in qdisc snapshot/restore for network attacks. Set STEADYBIT_EXTENSION_NETWORK_SNAPSHOT_RESTORE=true (e.g. via extraEnv) to make Apply capture the root qdisc tree (qdiscs + filters) of the target interface and Revert replay it after the attack's tc del. Preserves cloud-tuned root qdiscs (e.g. GKE's mq + fq with buckets=32768 horizon=2s) that would otherwise revert to kernel defaults after tc qdisc del root and leave the host network degraded until reboot. Off by default; Linux only.
  • Network attacks (delay, loss, corruption, bandwidth) now work on hosts where the kernel has already attached a default root qdisc to the target interface (e.g. mq on GKE COS / EKS / AKS / RHCOS). Previously the attack failed to start with NLM_F_REPLACE needed to override. The kernel default (mq, noqueue, fq_codel, pfifo_fast, fq) is restored automatically after the attack ends.
  • If the target interface carries a user- or CNI-installed root qdisc (e.g. htb, cake) that cannot be restored afterwards, the attack now fails fast in the prepare step with a clear error instead of silently replacing it.
  • Optional fallback: set STEADYBIT_EXTENSION_NETWORK_STRICT_ROOT_QDISC=true (e.g. via extraEnv) to make network attacks refuse any interface whose root qdisc is not noqueue — including the kernel default mq — instead of replacing it. Off by default.
  • New privileged chart value (default false): runs the extension in privileged mode and switches the managed SecurityContextConstraint to allow it. Needed on hardened nodes (e.g. CIS/STIG) where the container root filesystem is mounted nosuid, which voids the binary's file capabilities and breaks fault injection (nsenter: operation not permitted).
  • Stress CPU with "All cores" now uses every online CPU on hosts with more than 32 cores (previously capped at 32 due to a Cpus_allowed mask parsing bug).

extension-jenkins 1.0.14

  • chore(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#27)

extension-splunk 1.0.11

  • build(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump golang.org/x/net to v0.55.0 (CVE-2026-39821) (#50)

Platform 2.6.4

New

  • Service Variables: Define variables directly on a service and reuse them across experiments and validations in that service. Service variables sit between environment and experiment variables in the resolution order, so you can set sensible service-wide defaults that still allow per-experiment overrides. You can find more details on the Service Variable section of the public documentation.
  • Variable Expressions & Nested Variables: Variables are no longer limited to static text, they can now reference target values dynamically. For example, you can select a random subset of all reported Kubernetes cluster names when an experiment starts, then use that variable to target your attacks later in the same run. Variables are also composable now: one variable can reference others, whether static or dynamic, so you can build richer values from smaller building blocks. You can find more details on the Variables documentation page.
  • Search in Teams and Environments: Quickly filter long team and environment lists from the settings views.
  • Read-only access to immutable templates: Immutable (provided) templates can now be opened in view-only mode, with a clear indicator in the UI.

Improvement

  • Service-aware template flows: Template variables are pre-populated from the matching service's variables, and the template dialog is skipped entirely when all inputs are already resolved.
  • Service Average Risk chart: New horizontal legend bands and explicit line colors make the risk-over-time chart easier to read.

Fix

  • Experiment execution no longer mutates the saved design when overrides are applied.
  • 401 responses to XHR requests no longer pile up logout-redirect-logins.
  • SSE proxy responses are no longer buffered, removing artificial latency on live-update streams.
  • View-only icon is now only shown for templates that are actually immutable.

Agent 2.3.8

Dependencies

  • Dependency Updates with CVE fixes

extension-aws 2.4.17

  • chore(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/credentials
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigatewayv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticloadbalancingv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/fis
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/lambda

extension-gcp 1.0.28

  • chore(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump google.golang.org/api from 0.283.0 to 0.284.0

extension-auto-registration-ecs 1.0.10

  • build(deps): bump github.com/aws/aws-sdk-go-v2/config
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ec2
  • build(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • chore: update dependencies
  • feat: add weekly auto patch-release workflow

extension-auto-registration-kubernetes 1.0.4

  • build(deps): bump k8s.io/api from 0.36.0 to 0.36.1
  • build(deps): bump k8s.io/apimachinery from 0.36.0 to 0.36.1
  • build(deps): bump k8s.io/client-go from 0.36.0 to 0.36.1
  • chore: update dependencies
  • feat: add weekly auto patch-release workflow

extension-aws 2.4.16

  • chore(deps): bump github.com/aws/aws-sdk-go-v2 from 1.41.7 to 1.41.9
  • chore(deps): bump github.com/aws/aws-sdk-go-v2 from 1.41.9 to 1.41.12
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/feature/ec2/imds
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigateway
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/apigatewayv2
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/ecs
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/elasticache
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/eventbridge
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/kafka
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/lambda
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/mq
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/sqs
  • chore(deps): bump github.com/aws/aws-sdk-go-v2/service/sts
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-cloudfoundry 1.0.2

  • chore(deps): bump github.com/steadybit/discovery-kit/go/discovery_kit_sdk
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-container 1.6.7

  • build(deps): bump golang.org/x/sync from 0.20.0 to 0.21.0
  • chore: update dns-inject v0.2.1
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-datadog 1.8.21

  • build(deps): bump goreleaser/goreleaser from v2.15.4 to v2.16.0
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-dynatrace 1.0.22

  • build(deps): bump goreleaser/goreleaser from v2.15.4 to v2.16.0
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-gatling 1.0.46

  • chore: update gatling to 3.15.1 and ignore new netty CVEs (#129)
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-gcp 1.0.27

  • chore(deps): bump cloud.google.com/go/compute from 1.63.0 to 1.64.0
  • chore(deps): bump google.golang.org/api from 0.279.0 to 0.280.0
  • chore(deps): bump google.golang.org/api from 0.280.0 to 0.282.0
  • chore(deps): bump google.golang.org/api from 0.282.0 to 0.283.0
  • chore(deps): bump goreleaser/goreleaser from v2.15.4 to v2.16.0
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-host 1.5.7

  • chore: update dns-inject v0.2.1
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-host-windows 0.2.13

  • build(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • build(deps): bump golang.org/x/sys from 0.44.0 to 0.45.0
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-http 1.0.42

  • build(deps): bump goreleaser/goreleaser from v2.15.4 to v2.16.0
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow
  • fix(e2e): add untrusted local server for bad-ssl tests
  • fix(e2e): use local self-signed server for insecureSkipVerify test

extension-istio 1.0.23

  • chore(deps): bump istio.io/api from 1.30.0 to 1.30.1
  • chore(deps): bump istio.io/client-go from 1.30.0 to 1.30.1
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-jvm 1.2.17

  • chore(deps): bump github.com/shirou/gopsutil/v4 from 4.26.4 to 4.26.5
  • chore(deps): bump github.com/steadybit/action-kit/go/action_kit_commons
  • chore(deps): bump golang.org/x/net from 0.54.0 to 0.55.0
  • chore(deps): bump golang.org/x/sys from 0.44.0 to 0.45.0
  • chore(deps): bump golang.org/x/sys from 0.45.0 to 0.46.0
  • chore: bump runc/crun and update trivyignore
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-k6 1.2.5

  • chore(deps): bump goreleaser/goreleaser from v2.15.4 to v2.16.0
  • chore: update to go 1.26.4
  • chore: update to k6 v1.8.0 (#177)
  • feat: add weekly auto patch-release workflow

extension-kafka 1.2.15

  • build(deps): bump github.com/twmb/franz-go from 1.21.2 to 1.21.3
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-kong 2.0.24

  • build(deps): bump github.com/kong/go-kong from 0.75.1 to 0.76.0
  • build(deps): bump github.com/kong/go-kong from 0.76.0 to 0.76.1
  • build(deps): bump goreleaser/goreleaser from v2.15.4 to v2.16.0
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-kubernetes 2.6.25

  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow
  • feat: roll up HPA + PDB attributes onto workload targets (#307)
  • fix: more detailed error for kubectl exec failed

extension-postman 2.0.28

  • build(deps): bump node from 25-alpine to 26-alpine
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-prometheus 2.1.20

  • build(deps): bump github.com/prometheus/common from 0.67.5 to 0.68.1
  • build(deps): bump goreleaser/goreleaser from v2.15.4 to v2.16.0
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

extension-redis 1.1.5

  • chore(deps): bump alpine from 3.23 to 3.24
  • chore(deps): bump github.com/redis/go-redis/v9 from 9.19.0 to 9.20.0
  • chore(deps): bump github.com/redis/go-redis/v9 from 9.20.0 to 9.20.1
  • chore: update to go 1.26.4
  • feat: add weekly auto patch-release workflow

Platform 2.6.2

Dependencies

  • Dependency updates, including CVE fixes

Fix

  • Pagination was broken for all tabs of "Settings" -> "Integrations".

extension-container 1.6.6

  • DNS Error Injection: new hostname parameter to restrict injection to DNS queries with matching query names (exact, case-insensitive, IDN-aware); also exposes the new hostname_filtered metric in the live statistics widget
  • DNS Error Injection: clarify labels and descriptions for the port and cidr parameters — they apply to the DNS server, not to the queried domain
  • Bump bundled dns-inject to v0.2.0
  • Update dependencies

extension-host 1.5.6

  • DNS Error Injection: new hostname parameter to restrict injection to DNS queries with matching query names (exact, case-insensitive, IDN-aware); also exposes the new hostname_filtered metric in the live statistics widget
  • DNS Error Injection: clarify labels and descriptions for the port and cidr parameters — they apply to the DNS server, not to the queried domain
  • Bump bundled dns-inject to v0.2.0

Platform 2.6.1

❗Warning: This release contains breaking database changes

Note

This release reworks some internal database structures. Rolling back needs special care.

For on-prem customers, we strongly advise:

Create a database backup before the update, to be able to rollback to a previous version. In a multi-instance setup, all older instances must be stopped before starting any new instances when updating to this release. E.g. by using strategy: Recreate. Be aware that on termination, the platform waits for the experiment and agents to finish. If you force-delete the pod(s) during termination, the process might still be running.

New

  • Service Risk Reporting: New report showing the average risk and risk distribution across services over time, including support for filtering by service profile categories.
  • Service properties as report dimension: Reports can now be sliced and filtered by service properties.
  • SAML team sync on SaaS: SAML attributes can now be used to synchronize team membership on SaaS deployments.

Improvement

  • Line charts in Reporting: Numeric reports (including the new service risk report) now render as line charts with area fill, replacing or complementing the previous bar-chart visualization where it improves readability.
  • Rename OIDC-managed teams: Teams provisioned through OIDC can now be renamed without losing their managed status.
  • Recreating an access token without an expiration date is now possible.

extension-aws 2.4.15

  • Add discovery for 11 new AWS services: API Gateway, ASG, DynamoDB, EBS, EKS cluster + node group, EventBridge Rule, Amazon MQ Broker, NAT Gateway, NLB, SQS
  • Add attacks: Suspend ASG Processes, Trigger EKS Nodegroup Terminate Instance, Trigger MQ Broker Reboot, Disable EventBridge Rule, Throttle API Gateway (REST v1 + HTTP v2), Change Queue Visibility Timeout, Change Read/Write Table Capacity
  • Fix DynamoDB throttle attack to reject no-op capacity changes at Prepare
  • Fix API Gateway HTTP v2 throttle to restore account defaults on stop
  • Rename API Gateway Stage target to API Gateway
  • Update dependencies

extension-redis 1.1.4

  • breaking: the Pause Clients attack now always issues CLIENT PAUSE ... WRITE and the pauseMode parameter has been removed. CLIENT PAUSE ALL could not be aborted early (Redis blocks CLIENT UNPAUSE itself under an active ALL pause) and also stalled the extension's own discovery probes for the entire duration. Pausing only writes keeps the attack fully reversible via CLIENT UNPAUSE and lets discovery (PING/INFO) keep running. The attack has been relabeled "Pause Write Clients".

extension-appdynamics 1.1.9

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • chart: use shared extensionlib.deployment.env helper so standard env vars (logging, TLS, discovery group) flow through consistently
  • Update dependencies

extension-azure 1.2.9

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-container 1.6.4

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-datadog 1.8.20

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-dynatrace 1.0.21

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-gatling 1.0.45

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-gcp 1.0.26

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-grafana 1.1.2

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-host 1.5.5

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-host-windows 0.2.12

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-http 1.0.41

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-instana 1.1.15

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-istio 1.0.22

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-jenkins 1.0.11

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-jmeter 1.0.34

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-jvm 1.2.16

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-k6 1.2.4

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-kafka 1.2.14

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-kong 2.0.23

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-kubernetes 2.6.24

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-loadtest 1.0.10

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-newrelic 1.0.17

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-postman 2.0.27

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-prometheus 2.1.19

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-rabbitmq 1.0.13

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-redis 1.1.3

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-splunk 1.0.9

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-splunk-platform 1.0.9

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-stackstate 1.0.23

  • Support discovery group attribute via STEADYBIT_EXTENSION_DISCOVERY_GROUP env var (or discovery.group Helm value) — when set, the extension adds steadybit.group=<value> to every discovered target
  • Update dependencies

extension-container 1.6.3

  • Bump bundled nsmount to v1.1.1 — lowers the GLIBC requirement from 2.30 to 2.28, restoring .deb/.rpm installation on RHEL 8 / Debian 10
  • Bump bundled memfill to v1.3.1

extension-container 1.6.2

  • Fix Linux package: STEADYBIT_EXTENSION_DNS_INJECT_PATH was unset, causing DNS error injection attacks to fail on .deb/.rpm installations
  • Update dependencies

extension-host 1.5.4

  • Bump bundled nsmount to v1.1.1 — lowers the GLIBC requirement from 2.30 to 2.28, restoring .deb/.rpm installation on RHEL 8 / Debian 10
  • Bump bundled memfill to v1.3.1

extension-host 1.5.3

  • Fix Linux package: binary paths for nsmount, memfill and dns-inject were unset or pointed at the wrong directory, causing memfill and DNS error injection attacks to fail on .deb/.rpm installations
  • Update dependencies

Agent 2.3.5

Dependencies

  • Dependency updates, including CVE fixes.

extension-gcp 1.0.25

  • Allow starting vm instances with the existing VM attack action.
  • Bump Go to 1.26.3

extension-kubernetes 2.6.22

  • Added advanced parameter "Signal" for the "Pod Crash Loop" attack to be able to specify the signal used to kill the container process (default is "SIGKILL")
  • Update dependencies

Platform 2.5.9

New

  • Share experiments with other teams: Experiments can now be shared with other teams to let them run/schedule the same instance of the experiment design. Admin or team owner permissions are needed to share an experiment, and the experiment design always stays with the owning team's responsibility. Learn more in our docs.
  • New API endpoint for filtering experiments, including filters for sharing relationships.

Improvement

  • Property order validation is now resilient when other users add global properties while you're editing: missing associations are merged in along their existing predecessor chain instead of rejecting the save, and unassociated keys are dropped from the property order rather than failing the save outright.
  • Webhook payloads for experiment step execution events now include an executionStepId field referencing the step that triggered the execution event

Fix

  • Required metric query parameters (e.g. an empty PromQL query) are now validated.
  • Fix error in number of CPUs input for stress CPU attack

Platform 2.5.8

Fix

  • Error in the target enrichment pipeline, that could cause the message queue to pile up endlessly.

Platform 2.5.7

Note

⚠️ Update to 2.5.8 because of an error in the target enrichment causing the message queue to fill up.

Fix

  • Resolved an issue in the Stress CPU attack control where the configured number of cores did not reflect the actual number of workers and could not be modified in the UI

Platform 2.5.6

Note

⚠️ Update to 2.5.8 because of an error in the target enrichment causing the message queue to fill up.

New

  • Reliability Risk for Services: Every service now has a risk associated — a single indicator that summarizes its current reliability posture. Risk helps to navigate your chaos engineering rollout by understanding where to invest in reliability work next. Learn more about risk in our docs.

Improvement

  • Sample data for free trial tenants contain an example service now.
  • Externally managed teams and team members via LDAP or OIDC are now tagged as such.

Fix

  • Fixed an issue in some browsers, where tooltips in a context menu caused a Browser crash
  • 'Delete Service' context menu entry wasn't disabled when user wasn't allowed to delete the service. However, deleting a service was still prevented.
  • Users that are removed from OIDC groups are now also removed from the teams accordingly.
  • Fix for inconsistent environments after huge changes of targets due to congestion.

extension-host-windows 0.2.10

  • Fixed version handling - all definitions returned by the extension now return a valid semver version string instead of "unknown". If you had installed the extension before, please make sure to delete existing definitions in the platform after upgrading by visiting "Settings" -> "Extensions" and deleting the existing extension definitions. This is required to make sure that the platform correctly detects new versions of the definitions provided by the extension.

extension-prometheus 2.1.17

  • Bump Go to 1.26.2
  • Make the Prometheus request timeout configurable via STEADYBIT_EXTENSION_REQUEST_TIMEOUT and raise the default from 5s to 10s.

extension-gcp 1.0.24

  • Support discovery across multiple GCP projects via STEADYBIT_EXTENSION_PROJECT_IDS (shared credentials) or STEADYBIT_EXTENSION_PROJECTS_ADVANCED (per-project service-account impersonation). The legacy STEADYBIT_EXTENSION_PROJECT_ID continues to work.

Agent 2.3.3

Fix

  • Windows installer to correctly detect system architecture on older windows versions.

Platform 2.5.4

Dependencies

  • Dependency updates, including CVE fixes.

extension-azure 1.2.6

  • Bump Go to 1.25.9
  • Support if-none-match for the extension list endpoint
  • Make discoveries resilient to per-resource permission errors
  • Update dependencies

extension-istio 1.0.20

  • Bump Go to 1.25.9
  • Support if-none-match for the extension list endpoint
  • Update dependencies

extension-jenkins 1.0.9

  • Bump Go to 1.25.9
  • Support if-none-match for the extension list endpoint
  • Update dependencies

extension-k6 1.2.1

  • Bump Go to 1.25.9
  • Update to k6 v1.7.1
  • Log k6 exit code on warn
  • Update dependencies

extension-kong 2.0.20

  • Bump Go to 1.25.9
  • Support if-none-match for the extension list endpoint
  • Update dependencies

extension-kubernetes 2.6.20

  • Bump Go to 1.25.9
  • Support if-none-match for the extension list endpoint
  • Add k8s.service.name to hosts
  • Update dependencies

extension-splunk 1.0.7

  • Bump Go to 1.25.9
  • Support if-none-match for the extension list endpoint
  • Update dependencies

extension-host-windows 0.2.7

  • The stop process action reports an error if stopping the process is unsuccessful
  • The stop process action now correctly correlates parallel executions by their ID
  • If the extension runs with SYSTEM privileges, external tools are now directly executed, and not via a one-off SYSTEM task
  • Support if-none-match for the extension list endpoint
  • Update dependencies

extension-k6 1.2.0

  • Add same k6 extensions as supported on k6 cloud
  • Support if-none-match for the extension list endpoint

extension-container 1.5.12

  • Support if-none-match for the extension list endpoint
  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update dependencies

extension-gatling 1.0.41

  • fix: move index handler outside conditional block
  • Support if-none-match for the extension list endpoint
  • Update dependencies

extension-host 1.4.11

  • Support if-none-match for the extension list endpoint
  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update dependencies

extension-jmeter 1.0.29

  • Support if-none-match for the extension list endpoint
  • feat(chart): split image.name into image.registry + image.name
  • Align clusterName template to use dig-based nil-safe pattern
  • Support global.priorityClassName
  • Update dependencies

extension-jvm 1.2.12

  • Use target query to narrow down attack targets
  • Change Spring based attack labels
  • Support if-none-match for the extension list endpoint
  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update dependencies

extension-kafka 1.2.9

  • Support if-none-match for the extension list endpoint
  • Fix Kafka admin client connection leak in broker config describe operations
  • Retry on transient connection errors in broker config operations
  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update dependencies

extension-appdynamics 1.1.6

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-aws 2.4.10

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Support enrichment for argo rollouts
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-azure 1.2.5

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-datadog 1.8.16

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-dynatrace 1.0.17

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-gcp 1.0.20

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Support enrichment for argo rollouts
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-grafana 1.0.12

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-http 1.0.36

  • feat: add http bandwidth check
  • feat: add request-aware timeout header
  • feat(chart): split image.name into image.registry + image.name
  • fix: allow less than one http requests per second
  • fix: deadlock on stop when metric channel is fully packed
  • fix: cancel in-flight requests when stopping periodic http checks
  • fix: cancel fixed amount checks when at deadline
  • fix: handle zero completed requests in success rate check
  • fix: prevent worker from dying permanently on request creation failure
  • fix: prevent ticker goroutine from blocking on stop signal
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-instana 1.1.12

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-istio 1.0.19

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-jenkins 1.0.8

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-k6 1.1.6

  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-kong 2.0.19

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-kubernetes 2.6.19

  • Advice support new service validation step to ease experiment creation
  • fix: reference to service.id in case it is an array
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-loadtest 1.0.7

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Allow fixed value for poduid
  • Add label with index to pods and deployment
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-newrelic 1.0.14

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-postman 2.0.23

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-prometheus 2.1.15

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-rabbitmq 1.0.10

  • fix: prevent deadlock in publish stop when AMQP workers die
  • fix: prevent send on closed channel panic and reduce queue discovery overhead
  • fix: less details in logs when workers are involved
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-redis 1.0.1

  • fix: no default value for cache key
  • fix: reduce CPU and memory usage of BigKey attack
  • fix: used memory in MB
  • fix: client close unexpectedly
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-splunk 1.0.6

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-splunk-platform 1.0.6

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-stackstate 1.0.20

  • feat(chart): split image.name into image.registry + image.name
  • Support global.priorityClassName
  • Update alpine packages in Docker image to address CVEs
  • Update dependencies

extension-container 1.5.10

  • Handle OOMs in fill disk attack
  • Show summary message when container is stopped during stress and fill actions
  • Await fill disk attack duration before removing the created file
  • Don't fail resource attacks if target container is gone
  • Fix flaky schedule utils test
  • Update dependencies

extension-kubernetes 2.6.16

  • Copy the pod attributes by ownership to statefulset/daemonsets/deployments/replicasets instead of the selector

extension-kafka 1.2.7

  • Support multi cluster for configuration
  • Breaking change: now Check Brokers needs broker targets.

extension-host 1.4.4

  • feat: Network Delay - add option "TCP Data Packets Only" (PSH heuristic). Uses iptables marks + tc fwmark to delay only TCP data packets; UDP is not delayed. Honors ports/hosts/CIDRs via iptables filtering.
  • Update dependencies

extension-container 1.5.4

  • feat: Network Delay - add option "TCP Data Packets Only" (PSH heuristic). Uses iptables marks + tc fwmark to delay only TCP data packets; UDP is not delayed. Honors ports/hosts/CIDRs via iptables filtering.
  • chore: debug logging prints prepared iptables-restore scripts and tc/ip batch commands (for add and delete)

extension-http 1.0.31

  • Responses contain verification input was changed to a textarea to allow multi-line inputs.

extension-http 1.0.30

  • Fix: Responses contain verification parameters were renamed with v1.0.29. Existing experiment design will not verify the response if a paramter was set. This fix will revert the change and use the old parameter name.
    • If you have designed experiments with HTTP check v1.0.29 and used the Responses contain verification parameter, you need to migrate experiment designs after updating to v1.0.30.
    • Migration for SaaS Customers Please reach out to us.
    • Migration for On-Prem Customers
      • How to check whether you're affected? If the query below returns any rows, you need to migrate after updating to v1.0.30 and having a database backup in place
        SELECT es.experiment_key, es.custom_label, esa.action_id, es.parameters
          FROM sb_onprem.experiment_step es JOIN sb_onprem.experiment_step_attack esa ON es.id = esa.id
          WHERE esa.action_id IN ('com.steadybit.extension_http.check.periodically', 'com.steadybit.extension_http.check.fixed_amount')
          AND parameters ? 'responsesContain';
        
      • How to migrate existing experiments? After you've done a database backup, execute the following SQL
         UPDATE sb_onprem.experiment_step SET parameters = jsonb_set(parameters - 'responsesContain','{responsesContains}', parameters -> 'responsesContain')
           WHERE id IN (
             SELECT es.id
               FROM sb_onprem.experiment_step es JOIN sb_onprem.experiment_step_attack esa ON es.id = esa.id
               WHERE esa.action_id IN ('com.steadybit.extension_http.check.periodically', 'com.steadybit.extension_http.check.fixed_amount')
               AND parameters ? 'responsesContain'
           );
        

extension-azure 1.2.0

  • (Beta) add support for Azure Function
  • (Beta) add support for Network Security Groupy
  • (Beta) add support for .Net container apps

extension-prometheus 2.1.9

  • add option to enable detailed request and response logging
  • add option to add additional request parameters to the prometheus query
  • update dependencies

extension-kafka 1.2.4

  • Support changing IO and network thread count values with huge increments or decrements
  • Update dependencies

extension-kubernetes 2.6.10

  • Changing Advice's experiment templates
    • Schedule Pods Across Zones: attack 100% of the containers in one zone
    • Limit CPU/memory resources: attack 1 random container
    • Probes configured: attack 1 random container
  • fix: nginx ingress delay action (prepare step failed in some cases)

extension-container 1.5.0

  • Run steadybit sidecar containers using crun
  • Support crun on openshift >= 4.18
  • Use stressng --iomix (instead of --io) to stress io

extension-host 1.4.0

  • run steadybit sidecar containers using crun
  • use stressng --iomix (instead of --io) to stress io

extension-kubernetes 2.6.9

  • Discovery, Pod Count Check and Set Scale Action for ReplicaSets
  • Nginx Ingress: Discovery, delay and block traffic attack

extension-kafka 1.2.0

  • Add cluster name to broker target attributes
  • Better target ID for brokers in case of multiple clusters
  • Add min/max validations
  • Update dependencies

extension-host 1.3.2

  • If stress/diskfill/memfill exits unexpetedly report this as error and not as failure

extension-kubernetes 2.6.7

  • Resync internal k8s cache every 10m and increase update debounce to 20s (both values are configurable)
  • Optimize advice generation
  • Updated dependencies

extension-appdynamics 1.1.0

  • Breaking change - The access token is a short-lived token - Authentication is now done via OAUTH2.0 client credentials flow
    • Removed support for setting an access token via STEADYBIT_EXTENSION_ACCESS_TOKEN or appdynamics.accessToken
    • Added parameters client name, client secret and account name to the configuration.

extension-aws 2.4.5

  • Add AWS ECS Fargate network attacks
  • Add Windows host enrichment rule
  • Update dependencies

extension-kubernetes 2.6.6

  • possibility to disable the advice / kubescore feature
  • Updated dependencies
  • Updated go version to 1.24.4

extension-host 1.2.35

  • possibility to set the host.hostname attribute in the discovery by the k8s downward api

extension-http 1.0.27

  • ability to import own certificates for TLS connections
  • ability to ignore TLS errors for http connections
  • Updated dependencies

extension-kubernetes 2.6.4

  • added "Set Image" attack that allows to set the image of a container in a deployment
  • added the namespace to the messages of "Pod Count Check"
  • Updated dependencies

extension-prometheus 2.1.7

  • ability to import own certificates for TLS connections to prometheus
  • ability to ignore TLS errors for prometheus connections

extension-kafka 1.1.1

  • Make extension-kafka compatible with AWS MSK SCRAM-SHA-512 Auth
  • Add TLS for compatibility with SASL_SSL security protocol
  • Update to go 1.24
  • Update dependencies

extension-host 1.2.32

  • fix shutdown/reboot always failing on plain EC2 instances
  • Rename "Shutdown Host" to "Trigger Shutdown Host"

extension-container 1.4.8

  • add more prefill-queries
  • remove dependency to lsns
  • update depdendencies
  • require iproute-tc and libcap instead of /usr/sbin/tc and /usr/sbin/capsh

extension-host 1.2.31

  • remove dependency to lsns
  • update dependencies
  • require iproute-tc and libcap instead of /usr/sbin/tc and /usr/sbin/capsh

extension-container 1.4.7

  • Update dependencies
  • fix: fill disk fails when file permissions disallow write
  • fix: stress io fails when file permissions disallow write

extension-host 1.2.30

  • Updated dependencies
  • fix: fill disk/stress io fails when file permissions disallow write

extension-istio 1.0.14

  • update dependencies
  • Fix: only create faulty route if the sourceLabel is the same as the original route, discard if sourceLabel value is different to not unintentionally create a new faulty route with unwanted destination

extension-jvm 1.2.2

  • Update dependencies
  • Fix: JVM processes are lost in discovery when wall clock changes

extension-aws 2.4.1

  • Added optional tag filtering for all discoveries
  • Added an advanced method to configure role assuming for the extension.
  • include tags in the discovery of MSK clusters (requires new permission tag:GetResources)
  • include tags in the discovery of Elasticache (requires new permission tag:GetResources)
  • Update dependencies

extension-stackstate 1.0.14

  • Provide service status check mode to verify if the given state was observed at least once or all the time.

extension-kubernetes 2.6.0

  • Removed the advice single_aws_zone,single_azure_zone and single_gcp_zone and combined them using the generic attribute k8s.label.topology.kubernetes.io/zone. With the new advice, you are no longer required to install the cloud provider specific extension.
    • If you like to migrate your existing advice state, like created experiments and you are running ON-Premise, you can use the following migration script after installing the new version of the extension:
      update sb_onprem.advice
      set advice_definition_id='com.steadybit.extension_kubernetes.advice.single-zone',
          validation_states   = replace(validation_states::text, 'com.steadybit.extension_kubernetes.single-aws-zone',
                                        'com.steadybit.extension_kubernetes.single-zone')::jsonb
      where advice_definition_id = 'com.steadybit.extension_kubernetes.advice.single-aws-zone';
      
      update sb_onprem.advice
      set advice_definition_id='com.steadybit.extension_kubernetes.advice.single-zone',
          validation_states   = replace(validation_states::text, 'com.steadybit.extension_kubernetes.single-azure-zone',
                                        'com.steadybit.extension_kubernetes.single-zone')::jsonb
      where advice_definition_id = 'com.steadybit.extension_kubernetes.advice.single-azure-zone';
      
      update sb_onprem.advice
      set advice_definition_id='com.steadybit.extension_kubernetes.advice.single-zone',
          validation_states   = replace(validation_states::text, 'com.steadybit.extension_kubernetes.single-gcp-zone',
                                        'com.steadybit.extension_kubernetes.single-zone')::jsonb
      where advice_definition_id = 'com.steadybit.extension_kubernetes.advice.single-gcp-zone';
      

extension-kafka 1.1.0

  • Fix log line for check error
  • Change metric colors behavior
  • Change name of kafka config for certs

extension-kafka 1.0.6

  • Add controller information to target attributes
  • Add new broker check
  • Add TLS connection support
  • Update dependencies

extension-container 1.4.3

  • Rename some network actions to explicitly contain the term "outgoing"
  • Use runc binary from the opencontainers/runc project

extension-host 1.2.27

  • Rename some network actions to explicitly contain the term "outgoing"
  • Use runc binary from the opencontainers/runc project

extension-aws 2.4.0

  • ignore ecs services and tasks with a tag steadybit.com/discovery-disabled set to true
  • don't cache zones forever (for example removed permissions should lead to removed targets in the platform)
  • include tags in the discovery of Lambda functions (requires new permission tag:GetResources)
  • add vpc name to targets (requires new permission ec2:DescribeVpcs, can be disabled by STEADYBIT_EXTENSION_DISCOVERY_DISABLED_VPC)
  • add subnet target discovery
  • Update dependencies

extension-container 1.4.1

  • Respect the container memory limit for stress-ng based actions
  • Add option to disallow containers in certain namespaces
  • Update dependencies

extension-dynatrace 1.0.8

  • Handle event requests asynchronously, to avoid blocking the agent
  • Don't use entitySelector for events if the entity could not be found in Dynatrace
  • Problem check should ignore empty strings as entity selector
  • Don't log every event request
  • Update dependencies

extension-http 1.0.25

  • Add hint describing the behavior of the fixed amount check and lower the default duration to 2 seconds
  • Fix memory lead in the http check

extension-kubernetes 2.5.20

  • Integrated support for experiment templates in Advice to ease service's validation
  • Fixed a bug for Azure and GCP, where DaemonSets aren't considered in an Advice

extension-jvm 1.2.0

  • Breaking Change: Remove unreliable capturing of application context for spring boot applications
  • Fix: more reliable discovery for jvm processes

extension-kafka 1.0.5

  • Use uid instead of name for user statement in Dockerfile
  • Fix data race issue
  • Update dependencies

extension-azure 1.0.15

  • extend enrichment to more kubernetes types
  • update dependencies
  • Use uid instead of name for user statement in Dockerfile

extension-gatling 1.0.18

  • Optional location selection (can be enabled via STEADYBIT_EXTENSION_ENABLE_LOCATION_SELECTION env var, requires platform => 2.1.27)
  • Use uid instead of name for user statement in Dockerfile

extension-gcp 1.0.13

  • extend enrichment to more kubernetes types
  • update dependencies
  • Use uid instead of name for user statement in Dockerfile

extension-http 1.0.23

  • Location selection for http checks (can be enabled via STEADYBIT_EXTENSION_ENABLE_LOCATION_SELECTION env var, requires platform => 2.1.27)
  • Use "error" in the expected HTTP status code field to verify that requests are returning an error
  • Use uid instead of name for user statement in Dockerfile

extension-jmeter 1.0.17

  • Optional location selection (can be enabled via STEADYBIT_EXTENSION_ENABLE_LOCATION_SELECTION env var, requires platform => 2.1.27)
  • Use uid instead of name for user statement in Dockerfile

extension-k6 1.0.19

  • Optional location selection (can be enabled via STEADYBIT_EXTENSION_ENABLE_LOCATION_SELECTION env var, requires platform => 2.1.27)
  • Use uid instead of name for user statement in Dockerfile

extension-kubernetes 2.5.19

  • Avoid unnecessary enrichment rules for node labels, improving performance
  • update dependencies
  • Use uid instead of name for user statement in Dockerfile

extension-aws 2.3.5

  • Multi region support
  • EC2 Instance State Attack allows to start a stopped instance

extension-container 1.3.28

  • fix: Network attack cannot be executed, after a previous attack skipped cleanup for missing container
  • chore: update dependencies

extension-aws 2.3.3

  • Added MSK Support
    • Discovery of MSK Brokers
    • Action to trigger a reboot of a broker
  • Added Elasticache support
    • Discovery of Elasticache Nodegroups
    • Action to trigger a failover of a nodegroup
  • Fix graceful shutdown
  • Fix categroy and technology of ECS Stop Task attack
  • Update dependencies (go 1.23)

extension-gatling 1.0.16

  • Set new Technology property in extension description
  • "Fail" instead of "Error" if the Script is starting but containing some issues.
  • Update dependencies

extension-gcp 1.0.12

  • Set new Technology property in extension description
  • Update dependencies (go 1.23)

extension-grafana 1.0.1

  • Fix for better handling of annotations
  • Fix to handle multiple grafana targets
  • Update dependencies

extension-jvm 1.1.11

  • Set new Technology property in extension description
  • Update dependencies (go 1.23)

extension-k6 1.0.18

  • Set new Technology property in extension description
  • Update dependencies (go 1.23)

extension-kong 2.0.12

  • Set new Technology property in extension description
  • Update dependencies (go 1.23)

extension-container 1.3.24

  • fix: only create network excludes which are necessary for the given includes
  • fix: aggregate excludes to ip ranges if there are too many
  • fix: fail early when too many tc rules are generated for a network attack

extension-host 1.2.22

  • fix: fail block traffic early on hosts with cilium
  • fix: only create network excludes which are necessary for the given includes
  • fix: aggregate excludes to ip ranges if there are too many
  • fix: fail early when too many tc rules are generated for a network attack

extension-jvm 1.1.10

  • JVM excludes via vm arguments (like steadybit.agent.disable-jvm-attachment) are working again
  • Option to validate user provided class and method name for "Java Method Delay" and" "Java Method Exception" attacks
  • Align method parameter of "Controller Exception" and "Controller Delay" to "HTTP Client Status" and accept multiple values
  • Change default value for "jitter" in all "Delay" attacks to false
  • Fix graceful shutdown

extension-host 1.2.21

  • feat: change default value for "jitter" in "Network Delay" attack to false
  • feat: add memfill attack

extension-kubernetes 2.5.16

  • Increased timeout in the experiment for the single zone advice to detect a pod as being down within 45 seconds instead of just 30 seconds

extension-kubernetes 2.5.15

  • Be able to install the extension with a role instead of a service account to be able to work only in one namespace Example installation:
    helm upgrade steadybit-agent --install --namespace <replace-me-with-namespace> \
    --create-namespace \
    --set agent.key="<replace-me>" \
    --set global.clusterName="<replace-me>" \
    --set extension-container.container.runtime="<replace-me>" \
    --set agent.registerUrl="<replace-me>"\
    --set rbac.roleKind="role" \
    --set agent.extensions.autodiscovery.namespace="<replace-me-with-namespace>" \
    --set extension-kubernetes.role.create=true \
    --set extension-kubernetes.roleBinding.create=true \
    --set extension-kubernetes.clusterRole.create=false \
    --set extension-kubernetes.clusterRoleBinding.create=false \
    steadybit/steadybit-agent
    

extension-grafana 1.0.0

  • Add support for Grafana Alert Rules
    • Discovery of Alert rules
    • Check alert rules states
  • Add support for Grafana annotations
    • Send Steadybit events as annotations

extension-container 1.3.19

  • fix: Don't use the priomap defaults for network attacks, this might lead to unexpected behavior when TOS is set in packets

extension-host 1.2.17

  • fix: Don't use the priomap defaults for network attacks, this might lead to unexpected behavior when TOS is set in packets

extension-aws 2.3.0

  • Update dependencies (go 1.22)
  • Revisited AZ Blackhole attack
    • Reduced amount of required API calls
    • Added tests for rollback behaviour
    • Fixed a bug where an unused network acl wasn't deleted
  • Added ECS Support
    • Discovery for Tasks and Services
    • Action to stop a Task
    • Action to scale a Service
    • Actions to inject CPU/IO/memory stress or disk fill for a Task using SSM see README-ecs-ssm-setup.md for the necessary setup
    • Check Service Task count
    • Service Event Log
    • Requires new permissions (or needs to be disabled via STEADYBIT_EXTENSION_DISCOVERY_DISABLED_ECS)
      "ecs:ListClusters",
      "ecs:ListTasks",
      "ecs:DescribeTasks",
      "ecs:ListServices",
      "ecs:DescribeServices",
      "ecs:StopTask",
      "ecs:UpdateService"
      
  • Added ELB Support
    • Discovery for Application Load Balancers
    • Action to return a static response for a Load Balancer Listener
    • Requires new permissions (or needs to be disabled via STEADYBIT_EXTENSION_DISCOVERY_DISABLED_ELB)
      "elasticloadbalancing:DescribeLoadBalancers",
      "elasticloadbalancing:DescribeListeners",
      "elasticloadbalancing:DescribeTags",
      "elasticloadbalancing:DescribeRules",
      "elasticloadbalancing:SetRulePriorities",
      "elasticloadbalancing:CreateRule",
      "elasticloadbalancing:DeleteRule",
      "elasticloadbalancing:AddTags",
      "elasticloadbalancing:RemoveTags"
      

extension-host 1.2.15

  • added fallback attributes for availability zone of AWS to show one of AWS, GCP or Azure

extension-kubernetes 2.5.12

  • Renamed "Pod Count Check" to "(Deployment, StatefulSet, DaemonSet) Pod Count Check"
  • Pod-Targets now have a unique id. (Used by the UI to fetch details for a specific pod)
  • Update dependencies

extension-container 1.3.15

  • fail actions early when cgroup2 nsdelegation is causing problems
  • support cidrs filters for the network attacks

extension-host 1.2.14

  • fail actions early when cgroup2 nsdelegation is causing problems
  • support cidrs filters for the network attacks

extension-container 1.3.14

  • Update dependencies (go 1.22)
  • Added noop mode for diskfill attack to avoid errors when the disk is already full enough

extension-host 1.2.13

  • Update dependencies (go 1.22)
  • Added noop mode for diskfill attack to avoid errors when the disk is already full enough
  • Better logging to host shutdown / reboot

extension-gcp 1.0.9

  • Update dependencies (go 1.22)
  • Refactored config object
  • Refactored helm chart. Breaking changes. Please refer to the README for more information on how to authenticate.

extension-kubernetes 2.5.11

  • Update dependencies (go 1.22)
  • Added "Pod Count Check" for StatefulSets and DaemonSets
  • Improved advice's experiment for multi availability zones (single-azure-zone, single-aws-zone, and single-gcp-zone) to establish a 20s base-line in the beginning of the experiment
  • Add namespace label to container, k8s-container, k8s-deployment, k8s-statefulset and k8s-daemonset
  • Use FreeMarker syntax for advice templates.
  • Ignore Pods not in state "Running" in all discoveries

extension-kubernetes 2.5.10

  • Fixed advice's experiment for multi availability zones (single-azure-zone, single-aws-zone, and single-gcp-zone) to consistently use the same zone in every step
  • Improved instruction text for advice k8s-single-replica to better explain how to increase replicas for deployments and HorizontalPodAutoscaler

extension-aws 2.2.25

  • Update dependencies
  • Add Force Failover parameter to Trigger RDS Instance Reboot

extension-kubernetes 2.5.9

  • Update dependencies
  • Remove some attributes which have been used by the old 'weakspot' feature
  • Clarify the log message, if the extension stops listing pods, containers and hosts for deployments, statefulsets, etc. because of the discovery.maxPodCount configuration

extension-aws 2.2.24

  • Make attribute for EC2 data enrichment configurable via STEADYBIT_EXTENSION_ENRICH_EC2_DATA_MATCHER_ATTRIBUTE
  • Update dependencies

extension-kubernetes 2.5.7

  • Update dependencies
  • fix: update deployments if services/hpas have changes
  • fix: integrate kubescore check horizontalpodautoscaler-replicas

extension-aws 2.2.21

  • Update discovery_kit_sdk to v1.0.5, to resolve error in caching discovery
  • Update dependencies

extension-container 1.3.8

  • Update dependencies
  • Automatically set the GOMEMLIMIT (90% of cgroup limit) and GOMAXPROCS
  • Disallow running mutliple tc configs on the same container

extension-host 1.2.8

  • Automatically set the GOMEMLIMIT (90% of cgroup limit) and GOMAXPROCS
  • Disallow running multiple tc configurations at the same time

extension-kubernetes 2.5.5

  • use TargetEnrichmentRule Matcher Regex for copying k8s.label.* to container (exclude k8s.label.topology.*) (needs platform version >= 2.0.0 and agent version >= 2.0.2)

extension-kubernetes 2.5.4

  • Crash Loop Attack: validate specified container name with spec
  • Crash Loop Attack: ignore when to be killed container is already gone
  • Renamed attribute k8s.deployment.replicas to k8s.specification.replicas
  • Update dependencies
  • Add attributes k8s.label.topology.kubernetes.io/zone, k8s.label.topology.kubernetes.io/region, k8s.label.node.kubernetes.io/instance-type, k8s.label.kubernetes.io/os and k8s.label.kubernetes.io/arch to container, host, k8s-container, k8s-deplyoment, k8s-statefulset and k8s-daemonset

extension-datadog 1.8.4

  • update dependencies
  • Fix warnings Could not find step infos for step execution ... in logs

extension-postman 2.0.0

  • Breaking Changes
  • Configure ApiKey in the extension configuration
  • Discovery of Postman Collections
  • Support for Postman Environments to select the correct environment for the collection by name
  • Use a Postman Collection a target for the action

extension-instana 1.1.0

  • Added a discovery for application perspectives
  • Added an action to create a maintenance window
  • Filter event check based on application perspective(s)
  • Events are shown in a timeline with clickable links to the event details

extension-gcp 1.0.4

  • Update dependencies
  • add enrichment rules for kubernetes entities
  • align attribute naming

extension-datadog 1.8.2

  • Removed link to Steadybit homepage from event messages
  • use discovery_kit_sdk for discoveries
  • update dependencies

extension-container 1.2.0

  • Add disk fill attack
  • Add timeout and recovery for container discovery
  • Rework stress-io "Disk Usage" parameter to "MBytes written"

extension-host 1.2.0

Update to the latest helm chart steadybit-extension-host-1.0.33 needed!

  • add flush, read_write, read_write_and_flush mode to stress io
  • fill disk attack
  • fix stress memory and stress cpu constrained by the cgroup of the extension container

extension-gcp 1.0.3

  • Update dependencies
  • Added linux package
  • refactored to use discovery-kit-sdk

extension-aws 2.2.16

  • use discovery_kit_sdk for discoveries
  • add aws.zone.id to ec2- and rds-instances
  • added aws.zone.id to all enrichment rules

extension-aws 2.2.15

  • Added pprof endpoints for debugging purposes
  • Update dependencies
  • Enrichment for kubernetes-statefulsets, -daemonsets, -nodes and -pods

extension-azure 1.0.4

  • Added pprof endpoints for debugging purposes
  • Update dependencies
  • Enrichment for kubernetes entities

extension-jvm 1.0.11

  • Enrich application.name to container targets
  • Fixed container-to-jvm enrichment
  • Update dependencies
  • Prepared jvm advice

extension-kubernetes 2.5.0

  • Discoveries added
    • pods
    • daemonsets
    • statefulsets
    • nodes
  • Attack 'Delete Pod' added - :exclamation: Requires new permission delete for pods resources
  • Attack 'Drain node' added - :exclamation: Requires new permission create for pods/eviction resources and patch for nodes resources
  • Attack 'Taint node' added - :exclamation: Requires new permission patch for nodes resources
  • Attack 'Scale Deployment' added - :exclamation: Requires new permission get, update and patch for deployments/scale resources
  • Attack 'Scale StatefulSet' added - :exclamation: Requires new permission get, update and patch for statefulsets/scale resources
  • Attack 'Cause Crash Loop' added - :exclamation: Requires new permission create for pod/exec resources
  • Added options to check if a pod count increased or decreased to the existing pod count check action
  • Performance - Add hostnames to kubernetes-deployment during discovery instead of adding it via enrichment rule
  • Performance - Enrich hosts via kubernetes-node instead of frequent enrichments via kubernetes-container
  • Added pprof endpoints for debugging purposes
  • Memory optimizations
  • Removed the attribute k8s.container.ready as this causes unnecessary enrichment noise
  • Added additional attributes to support advice / weakspots - :exclamation: Requires new permission get, list, and watch for horizontalpodautoscalers resources

extension-container 1.1.21

  • fix invalid character 'i' in literal in runc State func. Do not combine stdout and stderr for json parsing

extension-aws 2.2.11

  • Make Discovery Intervals configurable
  • Keep a copy of current targets in discoveries and call aws apis not in the context of the agent request.
  • Allow parallel API calls using a configurable amount of worker threads via STEADYBIT_EXTENSION_WORKER_THREADS

extension-container 1.1.19

  • Use overlayfs for the sidecar containers reducing cpu consumptions drastically by avoiding to extract the sidecar container over and over again

extension-aws 2.2.10

  • Add enrichment rules for kubernetes deployments
  • Make targets to recieve EC2 data configurable via STEADYBIT_EXTENSION_ENRICH_EC2_DATA_FOR_TARGET_TYPES

extension-kubernetes 2.4.0

  • kubernetes-container are handled as enrichment data and not as targets anymore. (This requires at least agent 1.0.92 and platform 1.0.79)

extension-jvm 1.0.7

  • fix application discovery for rolling node deployments
  • refactor spring and datasource discovery

extension-aws 2.2.5

  • migration to new unified steadybit actionIds and targetTypes
  • added hint to aws account lookup of the agent in case of an error

extension-container 1.1.8

  • update dependencies
  • ignore marked containers during discovery
  • migration to new unified steadybit actionIds and targetTypes

extension-kubernetes 2.3.3

  • migration to new unified steadybit actionIds and targetTypes
  • ignore all labeled deployments and containers from discovery

extension-datadog 1.7.4

  • Add DateHappened to submitted DataDog events
  • Correctly select StepExecution for event creation

extension-aws 2.2.2

  • add RDS instance downtime attack
  • add RDS cluster failover attack
  • add RDS cluster discovery

extension-container 1.1.3

  • Exclude pause containers from Kubernetes and ECS in discovery
  • Fix error for runc inspecting containers using the systemd cgroup manager

extension-datadog 1.7.0

  • Links to Datadogs monitors are now using the timeframe of the experiment execution.
  • "Monitor Status Check" has a new parameter Status Check Mode. Supported values are All the time (default) and At least once.
  • New Action to create a Downtime for a monitor during an experiment execution.
  • Details about step executions are sent to Datadog as events.

extension-container 1.0.3

  • Bugfix: Blackhole and DNS container isn't reverted properly when container failed (and not the pod)

extension-container 1.0.2

  • New: new container.image attributes for registry, repository, and tag
  • Improvement: Logging improved when container couldn't stop because it wasn't found
  • Improvement: Error message for failures when starting stress-ng attacks
  • Bugfix: Fixed unique container ids for sidecar containers in same pod
  • Bugfix: Removing trailing / in container.name
  • Bugfix: Datatype for stop-container's graceful parameter
  • Bugfix: Blackhole container isn't reverted properly when container failed (and not the pod)

extension-k6 1.0.1

  • K6 Cloud: Stop running load tests per API if the user stops the steadybit experiment

extension-aws 2.1.0

  • Support Readiness & Liveness probes (requires helm chart version >= 2.0.0)
  • Refactored to use action_kit_sdk and thus use the extended rollback safety while having connection issues
  • Added Lambda discovery & actions (requires new permissions)

extension-aws 2.0.0

  • Renamed attack ec2-instance.state to com.github.steadybit.extension_aws.ec2_instance.state
  • Added EC2-Instance discovery
  • Added Zone-Discovery and Availability Zone Blackhole attack
  • Added AWS FIS-Experiment discovery and AWS FIS-Experiment action

extension-kubernetes 2.1.0

  • Kubernetes Event Log and Pod Metrics will need a cluster-selection to support multiple kubernetes clusters

extension-kubernetes 2.0.0

  • Added Discoveries for Deployments and Container
  • Added Pod Count Check and Node Count check
  • Added Pod Count Metrics and Event Logs

extension-kong 1.6.0

  • Support creation of a TLS server through the environment variables STEADYBIT_EXTENSION_TLS_SERVER_CERT and STEADYBIT_EXTENSION_TLS_SERVER_KEY. Both environment variables must refer to files containing the certificate and key in PEM format.
  • Support mutual TLS through the environment variable STEADYBIT_EXTENSION_TLS_CLIENT_CAS. The environment must refer to a comma-separated list of files containing allowed clients' CA certificates in PEM format.

extension-kubernetes 1.3.0

  • Support creation of a TLS server through the environment variables STEADYBIT_EXTENSION_TLS_SERVER_CERT and STEADYBIT_EXTENSION_TLS_SERVER_KEY. Both environment variables must refer to files containing the certificate and key in PEM format.
  • Support mutual TLS through the environment variable STEADYBIT_EXTENSION_TLS_CLIENT_CAS. The environment must refer to a comma-separated list of files containing allowed clients' CA certificates in PEM format.

extension-postman 1.3.0

  • Support creation of a TLS server through the environment variables STEADYBIT_EXTENSION_TLS_SERVER_CERT and STEADYBIT_EXTENSION_TLS_SERVER_KEY. Both environment variables must refer to files containing the certificate and key in PEM format.
  • Support mutual TLS through the environment variable STEADYBIT_EXTENSION_TLS_CLIENT_CAS. The environment must refer to a comma-separated list of files containing allowed clients' CA certificates in PEM format.

extension-prometheus 1.3.0

  • Support creation of a TLS server through the environment variables STEADYBIT_EXTENSION_TLS_SERVER_CERT and STEADYBIT_EXTENSION_TLS_SERVER_KEY. Both environment variables must refer to files containing the certificate and key in PEM format.
  • Support mutual TLS through the environment variable STEADYBIT_EXTENSION_TLS_CLIENT_CAS. The environment must refer to a comma-separated list of files containing allowed clients' CA certificates in PEM format.

extension-datadog 1.4.0

  • Support creation of a TLS server through the environment variables STEADYBIT_EXTENSION_TLS_SERVER_CERT and STEADYBIT_EXTENSION_TLS_SERVER_KEY. Both environment variables must refer to files containing the certificate and key in PEM format.
  • Support mutual TLS through the environment variable STEADYBIT_EXTENSION_TLS_CLIENT_CAS. The environment must refer to a comma-separated list of files containing allowed clients' CA certificates in PEM format.

extension-aws 1.7.0

  • Support creation of a TLS server through the environment variables STEADYBIT_EXTENSION_TLS_SERVER_CERT and STEADYBIT_EXTENSION_TLS_SERVER_KEY. Both environment variables must refer to files containing the certificate and key in PEM format.
  • Support mutual TLS through the environment variable STEADYBIT_EXTENSION_TLS_CLIENT_CAS. The environment must refer to a comma-separated list of files containing allowed clients' CA certificates in PEM format.

extension-aws 1.6.0

  • Support for AWS role assumption. This permits one extension instance from gathering data from multiple AWS accounts. To configure this, you must set the STEADYBIT_EXTENSION_ASSUME_ROLES environment variable to a comma-separated list of role ARNs. Example: STEADYBIT_EXTENSION_ASSUME_ROLES='arn:aws:iam::1111111111:role/steadybit-extension-aws,arn:aws:iam::22222222:role/steadybit-extension-aws'.

extension-aws 1.5.0

  • Support for the STEADYBIT_LOG_FORMAT env variable. When set to json, extensions will log JSON lines to stderr.

extension-datadog 1.3.0

  • Support for the STEADYBIT_LOG_FORMAT env variable. When set to json, extensions will log JSON lines to stderr.

extension-kong 1.5.0

  • Support for the STEADYBIT_LOG_FORMAT env variable. When set to json, extensions will log JSON lines to stderr.

extension-kubernetes 1.2.0

  • Support for the STEADYBIT_LOG_FORMAT env variable. When set to json, extensions will log JSON lines to stderr.

extension-postman 1.2.0

  • Support for the STEADYBIT_LOG_FORMAT env variable. When set to json, extensions will log JSON lines to stderr.

extension-prometheus 1.2.0

  • Support for the STEADYBIT_LOG_FORMAT env variable. When set to json, extensions will log JSON lines to stderr.

extension-datadog 1.2.1

  • Also observe the events experiment.execution.failed, experiment.execution.canceled and experiment.execution.errored to report all relevant event types to Datadog.

extension-datadog 1.1.0

  • Correctly mark duration parameter for status check action as required.
  • Add monitor status widgets to the execution view.

extension-postman 1.1.3

  • Define language-related environment variables in Docker image for consistency to the original postman/newman Docker image.

extension-kong 1.4.1

  • Use more specific Kong API gateway API endpoints to avoid security issues related to forbidden API endpoints. Contributed by @achoimet.

extension-aws 1.4.0

  • Restrict discovery execution to AWS agents to avoid common issues.
  • The log level can now be configured through the STEADYBIT_LOG_LEVEL environment variable.

extension-postman 1.1.0

  • The log level can now be configured through the STEADYBIT_LOG_LEVEL environment variable.

extension-kong 1.2.0

  • Update go-kong and use the new APIs so that plugin creation, updates and deletions happen using Kong API paths that are specific to services, i.e., located under /services.

extension-kong 1.1.1

  • Raise version of the request termination attack to v1.1.1 to update the configuration within Steadybit.

extension-aws 1.1.0

  • EC2 instance state attacks, i.e., EC2 instance stop, reboot, hibernate and terminate.