The more work I delegated to agents, the more attention I had to spend protecting my own agency.

I have been testing a more automated way of doing data science. In the morning, I define the larger question, hypotheses, and possible paths. Then the agents run experiments, sometimes on a cluster or GPU, while I go to meetings. I check in, answer questions, and redirect when needed. By evening, I have results. The next morning, I review them and repeat.

For a while, it feels like a superpower. Then the experiments, artifacts, and context begin to accumulate, and every review asks more of me. Agent activity starts to resemble progress even when the direction drifts. Parallel agents dilute my judgment. The familiar advice to “keep a human in the loop” starts to feel insufficient.

The human needs gates, interfaces, and discipline.

Human-Agent Inversion: When Delegation Starts Directing You

An agent proposes a reasonable next experiment. Accepting it takes seconds. Reconstructing the complete context takes much longer.

So you accept it.

Then the agent returns with a plausible result and three follow-up ideas. Each one makes sense locally. You choose one. The cycle repeats. After enough cycles, the agent is no longer executing inside your frame. You are making choices inside the frame the agent created for you.

I call this human-agent inversion.

The intended relationship is simple: the human defines the question and evaluates the evidence; the agent executes bounded work. In human-agent inversion, that relationship reverses. The human becomes subordinate to the agent’s direction.

Nothing forces this reversal. It develops gradually because fast output rewards trust, plausible explanations reduce the urge to inspect, and cognitive fatigue makes the proposed next step feel cheaper than rebuilding your own view of the problem.

The result still feels productive. Jobs run. Metrics move. Reports accumulate. But one week later, nothing may work beyond the basic baseline.

The inversion is dangerous precisely because losing control can feel so much like making progress.

More Agents, Less Judgment

The obvious response to capable agents is to run more of them.

One agent trains a model. Another investigates the data. A third changes the evaluation. Meanwhile, you attend meetings and switch between projects. Every session has its own context, assumptions, and open decisions.

Compute scales quickly, while human attention remains stubbornly limited.

The more sessions I run, the stronger the temptation to trust their summaries. My brain tries to reduce the cost of switching context, so I begin asking whether an experiment finished without first reconstructing why it exists. Proposed options replace the harder work of challenging the reasoning behind them.

Parallelism can therefore increase the risk of human-agent inversion. More agent capacity creates less human capacity to judge each direction.

My current practical limit is one focused session, or at most two. I do not claim this as a universal law, but the pause has become part of my workflow. I walk around, let my brain connect the dots, and ask what the agent is doing, why it is doing it, and whether its reasoning still matches mine.

As execution becomes abundant, the ability to zoom out becomes more valuable.

The first half of this problem is cognitive: attention, trust, and direction. The second half is structural. If agents can generate experiments faster than a human can interpret them, the workflow needs a larger unit of reasoning.

This is where my Agentic Data Science (ADS) framework enters the story. I am using it to encode boundaries that make automation governable, while continuing to refine it through daily work.

The ADS Framework: Experiments Are Too Small for Human Reasoning

Agentic data science needs two levels of hypothesis.

The first is the bundle-level hypothesis. It should be legible to someone outside the immediate implementation. We believe that a certain method or type of data can improve the model, solve a product problem, or change a decision. A peer or stakeholder should be able to understand the claim and question its logic.

The bundle is the human-sized unit of work. It contains the motivation, the intended outcome, the hypotheses, the accumulated evidence, and the final decision. It answers a basic question: what logical chunk of progress are we trying to make?

Inside the bundle are experiment-level hypotheses. These are technical and granular: extend one feature, add another feature, adjust regularization, or test a bounded model tweak. Agents can run many such experiments, and the details can become deep very quickly. All of them remain inside the bundle that gives the work its purpose.

The repository makes this relationship visible. A simplified research program with two successive bundles looks like this:

research/
└── bundles/
    ├── B001_tabular_baseline/
    │   ├── bundle.yaml
    │   ├── synthesis.md
    │   ├── decisions.yaml
    │   ├── publication.md
    │   └── experiments/
    │       ├── E001_recency_windows/
    │       │   ├── experiment.yaml
    │       │   ├── model/
    │       │   ├── runs/
    │       │   ├── evaluations/
    │       │   └── decision.yaml
    │       ├── E002_frequency_features/
    │       └── E003_calibration_tuning/
    └── B002_sequence_model/
        ├── bundle.yaml
        ├── synthesis.md
        ├── decisions.yaml
        ├── publication.md
        └── experiments/
            ├── E001_simple_sequence_encoder/
            ├── E002_time_gap_feature/
            ├── E003_transition_feature/
            └── E004_context_window_tuning/

The files have different jobs. Each bundle.yaml holds the bundle state, policy, anchor, finalist, and final decision. synthesis.md is the evolving human-readable synthesis. Every experiment owns its model, evidence, evaluations, and disposition. After a bundle decision is recorded and the bundle is closed, ADS can render publication.md as that bundle’s overall report.

Consider a concrete churn-model example.

Bundle B001: Establish a reliable tabular baseline asks whether a simple model built from stable behavioral features can provide a trustworthy reference for future work. This is a coherent question for data scientists, peers, and stakeholders: before exploring more complex models, do we have a baseline whose performance and failure modes we understand?

Its experiments are narrow enough to implement and evaluate independently:

  1. E001: Recency windows. Compare seven-day, thirty-day, and ninety-day activity windows for the existing recency feature.
  2. E002: Frequency features. Add counts of sessions and active days to the same simple model.
  3. E003: Calibration tuning. Adjust regularization and probability calibration after the useful feature set is fixed.

The B001 report records which experiment became the finalist, what the baseline can and cannot explain, and which evaluation contract the next bundle must use.

Bundle B002: Test a sequence model asks a new, larger question: does ordered session behavior improve early churn prediction beyond B001 enough to justify the added complexity? It uses the closed baseline bundle as context and comparison.

Its experiments stay technical:

  1. E001: Simple sequence encoder. Establish the smallest sequence model that can run under the same evaluation contract as B001.
  2. E002: Time-gap feature. Extend the event representation with time between sessions.
  3. E003: Transition feature. Add transitions between important event types.
  4. E004: Context-window tuning. Adjust sequence length and regularization without changing the bundle-level question.

Each experiment receives a local disposition such as retain, supersede, or reject. The B002 report then answers the question for the wider audience: whether sequence modeling added enough value, which candidate was selected, what complexity it introduced, and whether the direction should continue.

This boundary keeps technical progress connected to the original purpose. An experiment may improve a detail, but changing what the project is trying to establish requires returning to the bundle. Individual experiments can receive deep review when needed; the bundle remains the primary unit for zooming out.

The bundle is the peer-review surface. It lets a human return after a day, or after several meetings, and recover the direction of the work. It also makes a final decision possible: what did this collection of experiments establish, and what should happen next?

The ADS Workflow: Structure for Agent and Human Context

Agents have a context-management problem. So do humans.

When an agent’s context becomes fragmented or implicit, its behavior drifts. It forgets constraints, overweights recent information, and follows a locally plausible path. Human cognitive load creates a similar failure mode. After several meetings, sessions, and experiment reports, I also overweight what is in front of me. I accept the next plausible action because rebuilding the full reasoning is expensive.

Reminding myself to “stay in the loop” gives me little help when the context is already fragmented. I need the loop to preserve its own structure.

This is what the ADS workflow is trying to connect: agent context management and human context management. The command-line interface and file structure give the agent an exact operating state. The bundle gives the human a durable view of the question, hypotheses, evidence, gates, and decisions.

The same structure serves both sides. Worker agents receive bounded tasks with exact inputs and expected evidence. The coordinating session retains the larger context and owns hypotheses, gates, finalist selection, interpretation, and decisions. When I return between meetings, I can recover what happened, why it happened, and what remains undecided.

The structure forms a sequence:

  1. Preregister the hypothesis. Define the claim, expected direction, target metric, and abort condition before empirical evaluation, while the result is still unknown.

  2. Prove readiness. Require code or component tests, a complete-path smoke check, committed declared inputs, and available compute capacity before the experiment becomes expensive.

  3. Evaluate and select. Run development evaluation and select a finalist deliberately, keeping interpretation of the result with the coordinator.

  4. Seal confirmation. Lock the confirmation evidence before inspecting it, so a surprising result cannot conveniently reshape the hypothesis that produced it.

  5. Decide and close. Ingest the complete evidence, interpret it, record the bundle decision, and close the bundle. Continued tuning belongs in a successor bundle with a new explicit claim.

This structure is deliberately restrictive. Exact commands, readable files, immutable artifacts, and explicit transitions make the context reconstructable for both agent and human. After a meeting, I can answer “what ran?”, “why did it run?”, and “what decision is still open?” The agent can resume from the same state without inventing a new interpretation of the work.

Automation continues while I am away and pauses when the next gate would change the meaning or direction of the work.

The Larger Shift: Every IC Is Becoming a Manager

There is a funny consequence to all of this: every individual contributor using agents is becoming a manager of delegated reasoning and execution, even without managing people.

The required skills look familiar. Define the objective. Break work into bounded pieces. Inspect evidence. Ask basic questions. Zoom into a suspicious detail, then zoom out and reconnect it to the larger goal. Keep ownership of the decision.

These skills are often associated with executives because executives cannot inspect every implementation detail. They have to identify the assumptions that matter. Agentic work gives the same problem to everyone. Execution becomes cheap, while direction becomes scarce.

This changes the question I ask from “Can the agent do this?” to “What would I need to understand before I let the agent continue?”

Conclusion: Retain Authorship of Direction

I am still testing this workflow. It is fresh from the oven. The current setup is useful, and I intend to automate more of it.

I still want AFK execution and useful parallelism. The interface around them has to preserve human authority as the work becomes more automated.

Before continuing a piece of agentic data-science work, I want to be able to answer four questions:

  1. What is the bundle-level hypothesis?
  2. Why does the current experiment exist?
  3. What evidence would change my mind?
  4. Which gate comes next, and who is allowed to cross it?

If I cannot answer them, human-agent inversion may already be happening.

Agents can run the experiments. The question remains mine.