Skip to content

Troubleshooting

Look in this order: kubectl get aflow <name> -o wide for the phase and its detail, the events on the AnkkaFlow, the sidecar container's log, then its metrics. Each is described in Observe a pipeline. Then find the symptom here.

kubectl -n shop get aflow cart -o wide
kubectl -n shop get events --field-selector involvedObject.kind=AnkkaFlow
kubectl -n shop logs deploy/flow-cart-router -c sidecar

flow verify or flow generate refuses

The CLI lists every problem at once, one per line, and exits 1. Nothing has been deployed. The common ones:

message cause fix
'<outlet>' (…) is not compatible with '<inlet>' (…). the two ports on one topic carry different contracts give both the same schema name; a new version of a contract is a new name, and every reader must move to it
'<path>' does not point to a known streamlet inlet or outlet, please try … a typo in a port path, or a stale descriptor use a suggested path; regenerate the descriptor if the port was renamed
Streamlet '<name>' names descriptor '<d>', which no descriptor declares. the descriptor file is missing from --descriptors, or names the streamlet differently write the descriptor into that directory, or fix the name in blueprint.streamlets
Inlet <streamlet>.<port> is not connected. an inlet is on no topic add it to a topic's consumers
Topic '<id>' is not managed but has producers … the blueprint writes to a topic it does not own make the topic managed, or write to a managed topic instead
Topic '<id>' is not managed and names no bootstrap.servers or cluster. the platform cannot know where an unmanaged topic lives set cluster or bootstrap.servers on it, in the blueprint or with --conf
streamlet '<name>': parameter '<key>' has no default and no value a required parameter is unset set it with flow.streamlets.<name>.config in a --conf file
overrides name unknown topic '<id>' a --conf file names a topic the blueprint does not correct the id
Streamlet '<name>' has no image. generate was given no image for a streamlet add --image <name>=<ref> or an entry in --images
descriptor <file>: … a descriptor fails validation regenerate it with the SDK rather than editing it by hand

Every message is listed in the CLI reference.

The pipeline is Failed with SidecarImageMissing

The operator has no FLOW_SIDECAR_IMAGE, so it refuses every pipeline and applies nothing. Its log says so at start-up: sidecar image (not configured: every pipeline will be refused). Set the variable on the ankka-flow-operator Deployment; the pipelines are reconciled again when the operator restarts. See the operator reference.

The pipeline is Failed with Refused events

The resource cannot run, so nothing was applied. Each Refused event names one problem, and the status DETAIL lists them all:

detail fix
topic '<id>' uses Kafka cluster '<name>', but there is no Secret 'kafka-cluster-<name>' create the Secret in the operator's clusters namespace, ankka-flow by default
Kafka cluster '<name>' has no bootstrap.servers add bootstrap.servers to that Secret
managed topic '<id>' has no partitions: … (or replicas) set it on the topic, or as a default in the cluster's Secret
streamlet '<name>': inlet '<p>' is not bound to a topic a resource edited by hand; regenerate it with flow generate
streamlet '<name>': its descriptor does not parse: … regenerate the resource from a valid descriptor

Fixing a Secret needs no new apply: the operator reads the Secrets again at its next reconcile, within FLOW_OPERATOR_RESYNC_SECONDS (five minutes by default), or at once when the AnkkaFlow changes.

The pipeline is Degraded with TopicMissing

An unmanaged topic, one the pipeline reads but does not own, does not exist. The platform never creates it, and the consumers of that topic stay not ready; the sidecar logs inlet '<name>': topic '<topic>' does not exist; not ready until it does. Create the topic with whatever owns it, or correct topic.name. The streamlets become ready on their own once it exists.

A Kafka unreachable: … detail means the operator could not reach the topic's brokers: check the cluster's bootstrap.servers and connection-config.

TopicDiffers or TopicSettingsIgnored

A managed topic already exists with other partitions or replication than the resource declares, or with a different value for a topicConfig entry. The operator never alters an existing topic, so it left it as it is and warned. The pipeline still runs, on the topic as it exists. To change the topic, change it with Kafka's own tools, or delete it so the operator creates it afresh.

A pod never becomes ready

The pod's readiness is the sidecar's: a conversation with the process is running and every inlet is subscribed to a topic that exists. Read the sidecar's log:

log cause fix
discovery attempt <n> at 127.0.0.1:9010: …; retrying in …, repeating the process is not listening on 127.0.0.1:$FLOW_PROCESS_PORT check the process container's log; it may have crashed, or bind another port or address
refusing to start: the process does not match the deployed streamlet, then the sidecar exits 1 and the pod restarts the image's declaration differs from the descriptor in the resource: a port, contract or parameter was changed without regenerating, or the wrong image tag is deployed regenerate the descriptor from the code in the image, run flow generate with it, and apply; or deploy the image the descriptor was written from
the SDK speaks protocol '<v>' and this sidecar speaks '<v>'; the major versions must match the SDK and the platform speak incompatible protocol versions use an SDK of the platform's major version
the SDK speaks protocol '<v>', later than this sidecar's '<v>'; … the SDK is newer than the platform upgrade the platform, or pin the SDK
refusing to start: before any discovery, then exit 2 streamlet.conf or descriptor.json does not parse, or does not match the descriptor's ports and parameters on a laptop, correct the file; in a cluster, regenerate and apply the resource
inlet '<name>': topic '<topic>' does not exist; not ready until it does an input topic is missing create it; see the TopicMissing entry
stream failed: … then reconnecting in …, repeating the process fails a batch or breaks the protocol every time see the stalled partition entry

When the sidecar refuses a process at discovery, it sends every problem to the process first, so they also appear in the process container's log.

Lag grows and a PartitionStalled event appears

A batch fails every time: the process raises on it, or breaks the protocol. The sidecar never skips a record or sends it to a dead-letter topic; it redelivers the batch from the last commit, indefinitely, so that partition stops advancing while the others carry on. The stall shows as growing lag, ankka_flow_sidecar_stalled_seconds rising, and one PartitionStalled Warning on the pod after FLOW_STALL_WARNING_AFTER, naming the inlet, the partition and the last error.

The fix is in the streamlet: skip the record by acknowledging the batch without emitting for it, or correct the code that fails. Deploying the fixed image resumes from the last committed offset. See Delivery and failure.

The same record arrives twice

Delivery is at least once. A record whose emits were written but whose offsets were not yet committed when a pod stopped, a rebalance moved its partition, or the stream failed, is delivered again. This is expected; streamlet logic must tolerate repeats.

flow reset refuses

message fix
cannot reset offsets while streamlets are running: [<name>] is not scaled to 0 … set flow.streamlets.<name>.replicas = 0 with --conf, regenerate and apply
… [<name>] still has <n> pod(s) … wait for the pods to terminate, then run it again
no pipeline '<p>' in namespace '<ns>' pass -n with the pipeline's namespace
streamlet [<name>] has no inlets, so no consumer groups to reset leave that streamlet out; only streamlets that read have groups

A ResetRefused event with reset <id> waits: … means the request was recorded but a target still runs; the operator carries it out once the pods are gone. A ResetOffsetsFailed event carries Kafka's reason; consumer group [<g>] has <n> active member(s) means something still consumes with that group. See Rebuild from the start.

A streamlet does not roll after a change

A streamlet rolls when its image reference, descriptor or resolved configuration changes, and only then. The resolved configuration includes the Kafka connection taken from a cluster Secret, so changing the Secret rolls the streamlets that use it at the operator's next reconcile. Changing the pipeline's version alone rolls nothing, and neither does a new image pushed under the same tag; deploy a new tag.