Skip to content

GitOps-managed clusters

Argo CD and Flux keep your cluster matched to what’s in git. The KloudMate agent detects the workloads they manage and doesn’t change them, so you can’t turn on SDK tracing for those workloads from KloudMate. You can still trace them in other ways. Start with eBPF mode and work down the list.

ApproachPods restartRepo changeArgo CD or Flux change
eBPF modeNoNoneNone
Namespace annotationOnceUsually noneNone
Committed annotationOnceOne line per workloadNone
Agent patchingOnceNoneTwo settings
Policy engineOnceA policy manifestNone

Why the agent doesn’t patch these workloads

Section titled “Why the agent doesn’t patch these workloads”

To turn on SDK tracing, the agent adds an OpenTelemetry Operator annotation to the workload’s pod template. Your controller removes anything that didn’t come from git, and the agent would add the annotation again on the next check-in. Each change to the pod template restarts the pods, so they would restart once a minute, and nothing would show up as an error.

The agent checks the markers that each controller adds to a workload. Argo CD’s tracking annotation, an Argo-prefixed instance label, and the labels that Flux’s kustomize-controller and helm-controller add all identify the controller directly, so the agent trusts them on any cluster.

Argo CD with label-based resource tracking is the exception. It marks a workload only with the app.kubernetes.io/instance label, and plain Helm charts set the same label. So the agent trusts this label only on a cluster where Argo CD itself runs. On such a cluster, a workload that you installed with Helm and never put under Argo CD also shows as managed by Argo CD. Any of the options below still instruments it.

Check whether a workload is externally managed

Section titled “Check whether a workload is externally managed”

Go to Settings → Agents, open the cluster agent’s actions menu (⋮), and select Configure. On the APM tab, the Application instrumentation card shows Managed by Argo CD or Managed by Flux on each workload that one of them manages. Other workloads work as usual, even on a cluster that runs Argo CD.

Set the workload to eBPF in Application instrumentation and click Apply. The change takes effect on the next agent check-in. The agent doesn’t change your cluster objects or your manifests, and no pods restart, because eBPF attaches to processes that are already running.

You get request rate, errors, duration, and trace spans. You don’t get the in-process detail that an SDK adds, such as spans for database queries, framework internals, and outbound HTTP calls. For Go, Ruby, and PHP, eBPF is the only option on Kubernetes, whatever manages the workload, because KloudMate doesn’t inject an SDK for those runtimes there. Those workloads aren’t traced until you select eBPF for each one.

The OpenTelemetry Operator reads the inject annotation from the Namespace object as well as from a pod template, so one annotation covers every pod in the namespace:

kubectl annotate namespace payments \
  instrumentation.opentelemetry.io/inject-java=km-agent/km-agent-instrumentation-crd

Swap inject-java for the runtime you need: inject-nodejs, inject-python, or inject-dotnet. In a manifest, the same annotation looks like this:

apiVersion: v1
kind: Namespace
metadata:
  name: payments
  annotations:
    instrumentation.opentelemetry.io/inject-java: km-agent/km-agent-instrumentation-crd

Under Argo CD, the CreateNamespace=true sync option often creates the namespace, instead of a manifest in the app repo. If yours is created that way, the Namespace object isn’t part of any Application’s desired state, so annotating it doesn’t count as drift and nothing reverts it.

The annotation doesn’t change running pods. New pods get the injection, so delete the running ones or wait for the next deploy. Deleting pods doesn’t create drift, because your controller manages the Deployment, not the pods it creates.

Keep these limits in mind:

  • Every pod in the namespace gets injected, including short-lived jobs.
  • One runtime per namespace. The annotation sets a single language. If a namespace holds a Java service and a Python service, this covers only one of them.
  • .NET runtime options come from the pod only. If a .NET workload needs one, put it on that workload’s pod template.

Add the annotation to the workload’s pod template in your repo, where it goes through review like any other change:

spec:
  template:
    metadata:
      annotations:
        instrumentation.opentelemetry.io/inject-java: km-agent/km-agent-instrumentation-crd

Your controller applies it on the next sync, and the pods restart once. Argo CD and Flux need no extra configuration, because the annotation is part of the desired state.

Also set that workload to SDK in Application instrumentation. The agent still doesn’t change the workload. The setting keeps the workload tracked in KloudMate, and after your controller syncs the annotation, the workload shows as Instrumented.

If you’d rather keep the per-workload toggle in KloudMate, set up Argo CD to accept the agent’s annotation, then turn the agent’s patching back on. You need both of the Argo CD settings below.

First, ignore the inject annotations in the argocd-cm ConfigMap:

apiVersion: v1
kind: ConfigMap
metadata:
  name: argocd-cm
  namespace: argocd
data:
  resource.customizations.ignoreDifferences.apps_Deployment: |
    jsonPointers:
      - /spec/template/metadata/annotations/instrumentation.opentelemetry.io~1inject-java
      - /spec/template/metadata/annotations/instrumentation.opentelemetry.io~1inject-nodejs
      - /spec/template/metadata/annotations/instrumentation.opentelemetry.io~1inject-python
      - /spec/template/metadata/annotations/instrumentation.opentelemetry.io~1inject-dotnet

Inside a JSON pointer, ~1 stands for the / in the annotation key. Keep only the runtimes you instrument.

Second, add the sync option to every Application whose workloads you instrument:

spec:
  syncPolicy:
    syncOptions:
      - RespectIgnoreDifferences=true

With both settings in place, turn the agent’s patching back on with the patchExternallyManagedWorkloads Helm value:

helm upgrade kloudmate-release kloudmate/km-kube-agent \
  --namespace km-agent --reuse-values \
  --set patchExternallyManagedWorkloads=true

This value applies to the whole cluster. If an Application doesn’t have the settings above, Argo CD reverts the annotation and the Application’s pods keep restarting. With the value on, workloads that Argo CD manages no longer show Managed by Argo CD in Application instrumentation. They show their usual instrumentation status instead.

These steps are for Argo CD. On Flux, exclude the same annotations from drift correction before you turn on patchExternallyManagedWorkloads, or Flux reverts the agent’s changes.

If you already run Kyverno, have it add the annotation to pods when they’re created:

apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: kloudmate-instrumentation
spec:
  rules:
    - name: annotate-java-pods
      match:
        any:
          - resources:
              kinds:
                - Pod
              namespaces:
                - payments
      mutate:
        patchStrategicMerge:
          metadata:
            annotations:
              instrumentation.opentelemetry.io/inject-java: km-agent/km-agent-instrumentation-crd

Argo CD and Flux compare Deployments, not the pods those Deployments create, so a policy that annotates pods never shows up as drift. The OpenTelemetry documentation suggests the same pattern, and KloudMate needs no configuration for it. As with the namespace annotation, running pods get the annotation when they’re recreated.

Install the agent with Helm, as shown in the Kubernetes installation guide.

To manage the agent through Argo CD instead, set these on its Application:

  • ignoreDifferences on the agent’s own ConfigMaps. The agent writes each collector’s live configuration into the /data/agent-daemonset.yaml and /data/agent-deployment.yaml keys of km-agent-configmap-daemonset and km-agent-configmap-deployment. Argo CD renders the chart without reading the cluster, so without this setting, every sync resets the collectors to the configuration in the chart.
  • The destination namespace set to km-agent. The chart always creates the Instrumentation resource in km-agent, and the inject annotation points to it as km-agent/km-agent-instrumentation-crd.
  • SkipDryRunOnMissingResource=true in the sync options. On the first sync, Argo CD applies the Instrumentation resource in the same wave that installs the OpenTelemetry Operator’s CRDs, and the dry run fails without this option.
spec:
  destination:
    namespace: km-agent
  syncPolicy:
    syncOptions:
      - SkipDryRunOnMissingResource=true
  ignoreDifferences:
    - group: ""
      kind: ConfigMap
      name: km-agent-configmap-daemonset
      jsonPointers:
        - /data/agent-daemonset.yaml
    - group: ""
      kind: ConfigMap
      name: km-agent-configmap-deployment
      jsonPointers:
        - /data/agent-deployment.yaml