ADR-0049: Node rows own pipeline layout; Pipeline.data keeps only edges
ACCEPTED
Extends: ADR-0046
Context
ADR-0046 made the Node rows the sole source of node content and reduced
Pipeline.data to a layout blob: per node an id, react-flow type and position,
plus the graph's edges. Position was the only per-node value still living in the
blob, and the react-flow type is derivable from Node.type (StartNode → startNode,
EndNode → endNode, everything else → pipelineNode). Keeping the nodes list in the
blob meant every read still reconciled two node lists, and layout could drift from the
rows exactly as content used to.
A prior phase added nullable position_x/position_y columns to Node and shadow-wrote
them on every save, so the rows already carry layout. This ADR completes the move.
Decision
We will remove node information from Pipeline.data entirely: it becomes {edges} with no
nodes key (plus a vestigial errors key, still written on every save). The Node rows are
the sole source of layout as well as content — position from the position_x/position_y
columns, react-flow type derived from Node.type.
Pipeline.flow_datarebuilds the full react-flow graph from the rows (position and type from each row); it reads onlyedgesfromPipeline.data.- Saves supply a complete membership mapping to
update_nodes_from_data(node_data): its keys are every node id in the graph, each value either content (create/update the row, writing its position) orNone(membership only — the row must exist and is left untouched). Removed rows are deleted, or archived when they have versions. - The PATCH engine works off
flow_datasincePipeline.datano longer lists nodes. It returns the node-less blob and the complete mapping. - The flow models carry no unknown top-level keys, and every save rewrites
Pipeline.datafrom the parsed graph, so the stored blob only ever holds what they model. A key left over in an older blob (the editor'sviewport, which nothing has read since ADR-0046) is therefore dropped by the pipeline's next save. duplicate_pipeline_with_new_idstakes the id→type mapping from the rows, generates new ids and rewrites edges — it no longer reads a node list from the blob.- Migration
pipelines.0030_strip_node_databackfillsposition_x/position_yfrom the blob and then drops thenodeskey, so the layout reaches the rows in the same deploy that switches reads to them. Its reverse rebuilds the fullnodeslist from the rows, so a code rollback remains possible. The backfill is the load-bearing half; the strip is housekeeping, since reads already ignore anodeskey and every save drops it. Pipelines whose blob holds content with no backing row are skipped by a drift guard and keep theirnodeskey, so "nonodeskey" holds for every pipeline the migration touched, not literally all of them.
The wire format is unchanged: the editor still sends and receives full nodes.
Consequences
Pipeline.datano longer duplicates any node state; layout/content drift is structurally impossible for migrated rows, andflow_data's old missing-nodeKeyError(a row absent fromdata["nodes"]) is gone because the rows drive the read.- Position columns stay nullable, so migration 0030 is a hard prerequisite for the read
switch rather than a convenience. A row it cannot fill (no usable position in the blob)
renders at the origin, and the first save of that pipeline drops the
nodeskey — after which the pre-migration layout is gone for good, since a read serves the origin and the editor only sends positions it sees change. - Layout drift introduced after the migration (a writer that bypasses the shadow-write) needs a new data migration to heal: 0030 runs once at deploy and there is no command to rerun.
- The strip half constrains deploy ordering. Once a pipeline's
nodeskey is gone, pre-0048 code cannot parse itsPipeline.data(nodesis required there), so the editor's GET/POST/PATCH return 500 for it and the widget page context loses its nodes. If migrations run ahead of the new code, that is the window. Chat and pipeline execution are unaffected, sincePipelineGraphreadsedgesplus the rows. Rolling the code back requires unapplying 0030 so the blobs are rebuilt. update_nodes_from_datachanged signature again (values may beNone); every caller and test passes a complete mapping. Because an incomplete mapping means "delete the rest", it validates the mapping before removing anything and runs in a single transaction.- Making the position columns non-null is deferred to a follow-up, gated on migration 0030 having landed everywhere.
Alternatives considered
- Keep the
nodeslist in the blob for layout: rejected — leaves layout with two owners and the same drift class ADR-0046 removed for content, for no benefit now that the columns exist. - Pass graph membership separately from the content mapping: rejected — a single complete mapping keeps one argument as the authority for both membership and content.
- Make the position columns non-null in this change: rejected — requires the backfill to have run on every environment first; sequenced as a later ADR.
- Ship no migration, relying on the one-off backfill command having already been run:
rejected. The command was only ever invoked by hand, the position columns are days old so
the shadow-write has populated almost nothing organically, and a self-hosted deployment
would upgrade with NULL positions. Reads would then serve the origin and the first save
would drop the
nodeskey holding the real coordinates — permanent layout loss, with no command left to repair it. The migration is idempotent and cheap where the backfill has already run, so the asymmetry is one-sided.