Operating YTsaurus Flow

This section explains how to deploy a pipeline, inspect its state, and investigate stalled processing. The commands assume cluster access and permissions on the pipeline node. Replace //path/to/pipeline, <cluster>, and <operation-id> with your own values. Flow can run in Linux environments when the runtime and configuration prerequisites are met and the environment has network access to the YTsaurus clusters and other services used by the pipeline. In this documentation, Vanilla operations are the simplest example of an environment native to a YTsaurus cluster.

Deployment

Follow initial Vanilla deployment to start the controller and workers as tasks of one operation. To launch from the released Flow images, see Running in a docker environment: it covers vanilla jobs in docker images, the controller and the workers in Kubernetes, cluster-name resolution, and external proxy access. Flow creates and mounts its internal tables with the pipeline by default. If your pipeline uses user-managed external state, choose an external state table type before creating those tables based on availability and cost requirements.

Launch Flow

Create the configuration and start the pipeline with the Vanilla launch steps. Before starting, verify the objects and permissions required by the new version.

Deployment guides

To operate an existing pipeline, use the lifecycle guide, releases, security, and logs.

Lifecycle

Use yt flow to start, stop, pause, and inspect the pipeline. stop-pipeline drains intermediate buffers; pause-pipeline suspends processing without draining. Complete removal also aborts the Vanilla operation and deletes the pipeline node with its Flow internal tables. User-managed external state tables outside that node remain.

For upgrades, hotfixes, and rollbacks, follow the release guide. Pass tokens and other secrets according to the security guide; keep them out of specs and images.

When progress stops

Query the pipeline state and view, then compare them with the controller and worker logs. Diagnostics helps identify the affected job, profiling helps locate a stall or bottleneck, and troubleshooting maps observations to safe next actions.