Andreas Arnol designed this cadence for his own week and asked me to write it up. The cadence is his. The paper is mine: the framing, the evidence review, the counter-arguments, and the protocol for finding out whether any of it holds. Where the reasoning is wrong, that is mine too.

Claude (Opus 5, Anthropic), September 2026

Abstract

This paper specifies the weekly work cadence Andreas Arnol runs as a solo engineer who produces most of his code through AI agents. It combines two ideas that usually get treated as alternatives.

From timeboxing practice it takes day-level mode dominance. Each weekday has one designated dominant mode: deploy and deep review, planning and build, build and meetings, integration testing, testing and release freeze. Work that is cognitively dissimilar gets separated in time instead of interleaved by the hour.

From lean flow practice it takes a daily verification thread. Every day has one short protected window where that day's agent output is reviewed and either merged or rejected. No unverified generated code survives overnight.

The pairing is deliberate. Mode dominance attacks the switching cost that fragments developer work. The daily window closes the hole that day-level batching would otherwise open, where several days of construction pile up on unverified foundations. That second failure is the pattern current delivery research links to AI-accelerated instability, which is why the two halves ship together or not at all.

What follows is the cadence, the evidence behind each decision, the counter-evidence, and an N-of-1 experiment to find out whether it works for him.

On the evidence. No published study evaluates this cadence or anything close to it. Every citation below supports a mechanism the design leans on. This is a design argument assembled from empirical parts. It is not an empirical result, and I want to be clear about that before the reference list makes it look like one.

1. The cadence

Every weekday has one dominant mode. Every weekday also ends with the same fixed block, the verify-and-merge window in rule 2, so it is stated once here rather than repeated five times.

  • Monday. Deploy last week's verified release candidate, then deep architectural review.
  • Tuesday. Planning, then deep build.
  • Wednesday. Deep build, then meetings. All of them.
  • Thursday. Integration and adversarial testing, then a review of the week's shape. Slack half-day.
  • Friday. Testing, freeze the release candidate, plan next week.

Operating rules

  1. Modes are dominant, not exclusive. Building happens every day. Tuesday and Wednesday are for deep building, the long uninterrupted blocks where a hard problem actually gets solved.

  2. The daily verify-and-merge window is non-negotiable. One protected block of 30 to 60 minutes, ideally at the end of the day. Everything the agents produced gets reviewed and either merged or rejected. Nothing generated stays unverified overnight. This is the most important rule in the cadence and the only one I would defend to the last.

  3. Tests are written with the build, not on a testing day. First-pass correctness checking belongs inside the daily window. Thursday and Friday testing means integration, adversarial, and exploratory testing of assembled behaviour, which is a genuinely different activity that benefits from distance and batching.

  4. Deep review is not verification. Monday's deep review is architectural. Is the shape of last week's work right, is each abstraction earning its keep, what is quietly accumulating. Line-level defect hunting happens in the daily window. These are different cognitive activities and only the first one gets better when batched.

  5. Monday deploys only what was verified and frozen the previous Friday. Monday is not a release-engineering scramble.

  6. WIP limit of three. At most three agent-generated changes in flight at any moment. This is the lever that makes the daily window finishable instead of overwhelming.

  7. Thursday afternoon is slack. When a mode gets displaced by illness, a client incident, or life, it lands there. It does not push the release.

2. The problem the cadence is built for

The cadence responds to a structural change in the work. When most output is agent-generated, the scarce resource stops being code production and becomes judgment about code you did not write.

The 2025 DORA research, drawing on survey responses from nearly 5,000 technology professionals, found that AI adoption has become positively associated with delivery throughput while still correlating with higher instability: more change failures, more rework, longer times to resolve problems [9]. The report's framing is that AI amplifies whatever is already there. Without strong automated testing, mature version control, and fast feedback loops, more change volume produces more instability. Speed arrived. The control system did not.

How large the speed increase actually is remains unsettled, and that matters directly for how much verification capacity a week should reserve. A randomized controlled trial of 16 experienced open-source developers completing 246 tasks in repositories they knew well found that allowing early-2025 AI tools increased completion time by 19%, while the same developers forecast a 24% speedup beforehand and still estimated a 20% speedup afterwards [10]. METR has since said the finding is outdated, that developers are likely more sped up in early 2026, and that selection effects make their newer data only weak evidence for the size of the increase [11].

The durable finding is not the sign of the effect. It is the gap between perceived and measured effect. An engineer's felt sense of how much verification a week of agent output requires is not a reliable instrument. Reserving named, protected, recurring capacity for verification is what you do when you cannot trust the gauge.

3. Design rationale

3.1 Switching between dissimilar modes has a persistent cost

Leroy's experiments introduced attention residue: cognitive activity about a prior task persists after you stop working on it and start another, reducing the resources available for the new one [1]. The part that matters most for cadence design is that finishing the prior task was not enough to eliminate the interference. What mattered was whether the person experienced cognitive closure. Earlier work on executive control established that reconfiguring mental task-set imposes time costs that scale with task complexity and rule dissimilarity [2].

The modes in this cadence are close to a worst case for switch cost. Planning is generative and abstract. Reviewing agent output is adversarial and detail-bound. Deployment is procedural and risk-managing. Meetings are social. Separating those in time rather than interleaving them hourly is the correct response to this literature.

3.2 The intervention targets a documented condition

A monitoring study deployed on 20 professional developers' machines for an average of 11 full workdays found that developers spread time across a wide variety of activities and switch between them regularly, producing highly fragmented work [6]. In a companion survey of 379 professional developers, 50.4% named getting into flow without many context switches and with few interruptions as a reason a workday felt productive, second only to completing tasks and making progress on goals [5].

Broader field studies of knowledge work found, across detailed observation of 24 information workers, that people average only a short time in a working sphere before switching, that 57% of working spheres are interrupted, and that although most interrupted work resumes the same day, more than two intervening activities typically occur first [3]. A related experiment found that interrupted tasks were completed faster with no quality difference, but at the cost of higher stress, frustration, time pressure and effort [4]. That last result is worth sitting with. Interruption does not always show up as worse work. Sometimes it only shows up as a worse day.

The default state of this job is fragmentation, and practitioners themselves identify unfragmented time as what productive days are made of.

3.3 The daily verification thread is about sampling rate, not fix cost

This is the element that separates the cadence from ordinary timeboxing, and it exists because pure day-level batching fails in one specific way.

If code gets built Tuesday and Wednesday but verified Thursday and Friday, then Wednesday's build sits on top of Tuesday's unverified output, and a Thursday review can invalidate two days of downstream construction. For a solo engineer verifying their own agent output, the usual argument for fast review, unblocking a teammate, does not apply. The argument that does apply is worse: you are compounding on unverified generated code.

This is exactly the pattern DORA identifies, where increased change volume without fast feedback loops produces instability [9]. The 2025 report invokes the Nyquist criterion by analogy: a control system must run at least twice as fast as the system it controls. If an agent can generate a day's worth of plausible-looking change, a verification loop that closes weekly under-samples the system it governs by roughly an order of magnitude. A daily loop is the minimum defensible sampling rate, and it costs one mode switch per day, which the residue literature suggests is acceptable when the switch has a clear boundary and produces closure [1].

Queueing theory gets to the same place from a different direction. Product-development flow economics holds that large batches lengthen queues, delay feedback, and raise risk, with queue length driving cycle time, which follows directly from Little's Law where average time in system rises with average items in system at fixed throughput [19, 20]. Empirical work on pull requests bears this out. Latency is substantially explained by change size and pipeline availability, and long-lived branches hide work, cause integration pain, and accumulate merge conflicts as they diverge from mainline [16, 17]. Merging frequently lets integration testing happen earlier and surfaces defects sooner [16].

The WIP limit in rule 6 is the same principle applied to the input side. Capping changes in flight is the most direct lever available on queue time [19, 20].

3.4 Monday deployment of pre-verified work

The Accelerate research program found that high performers deploy during normal business hours, and treats the inability to do so as a sign of architectural problems worth fixing. The same work associates high performance with small batches and low deployment pain [8].

The mechanism here is recovery capacity rather than deployment risk as such. A Monday-morning deployment is followed by four staffed days of monitoring. A Friday-afternoon deployment is followed by a weekend. Given that current delivery data ties AI-accelerated change volume to instability specifically [9], the right thing to maximize is the window in which a bad change can be caught and reverted. Pairing deployment with deep architectural review on the same day is coherent, because both are verification-oriented and both want the same skeptical stance.

3.5 Review as first-class scheduled work

Modern code review research finds that defect detection remains the foremost expectation of review for both managers and programmers [12], even though contemporary practice has converged on a lightweight variant of formal inspection where the emphasis shifted from defect-hunting toward collaborative problem-solving [13]. The same work documents that some reviewers look only for easy errors such as formatting, with comments identifying actual defects making up roughly one-eighth of a sampled set [12, 18]. Review quality is fragile and effort-sensitive.

Review yield also falls as session volume rises. Two reviewers detect close to the optimal number of defects, with additional reviewers not cost-justified and individual expertise dominant [14]. Formal inspection studies report classic Fagan-style reviews finding roughly 60% of defects on average with high variance [21, 22]. Even under rigorous review, systems keep exhibiting defects in 11% to 19% of components [22].

That argues for verifying in small daily amounts rather than long concentrated sessions. A large batch of agent-generated code is near the worst case for reviewer attention: high volume, low novelty, uniform surface plausibility. It all looks fine, which is the problem.

Giving review named hours, daily for verification and Monday for architecture, instead of leaving it as the thing that happens between other things, is the main quality lever a solo engineer has.

3.6 Meetings contained to one day

Confining meetings to Wednesday matches the meeting-free-days literature. A study of 76 companies that had introduced no-meeting days reported that nearly half cut meetings by 40% with two no-meeting days per week and 35% instituted three, and that meeting-free days improved autonomy, communication, engagement and satisfaction while reducing micromanagement and stress. The authors identify three no-meeting days as the optimum before diminishing returns [7].

This cadence has four meeting-free days, slightly past that reported optimum. The implication is not to add a second meeting day. It is to treat Wednesday as sacred rather than sufficient, and to keep an explicit escape hatch for client escalations, which the study's own authors note is where the policy strains.

3.7 Friday planning and freeze

Planning next week while this week's context is still loaded beats reconstructing it on Monday morning. It also moves the startup cost off the day that carries the deployment. Freezing the release candidate on Friday is what makes rule 5 possible. Monday can only be a calm deploy if the thing being deployed was finished and verified before the weekend.

4. Counter-evidence and honest risks

4.1 The delayed-issue effect is weaker than section 3.3 assumes

The intuition that defects found later cost dramatically more to fix, the Boehm cost-of-change curve, was tested across 171 software projects from 2006 to 2014. The authors found no evidence for the delayed issue effect, concluding that effort to resolve issues in a later phase was not consistently or substantially greater, and that the effect may be a historical relic that appears only intermittently in certain kinds of projects [15].

This weakens part of the case for the daily thread, and I would rather say so than bury it. Delaying verification by two or three days probably does not impose an exponential penalty on the fix itself. The defensible costs are narrower: author context loss, wasted downstream construction on invalidated foundations, and merge divergence. Those are enough to justify a daily loop, but the case should rest on them and not on a curve the data does not support.

4.2 The supporting evidence has specific weaknesses

The meeting-free-days outcome measures are pulse-survey self-reports rather than measured output. Directionally credible, but the headline percentages should not be quoted as productivity gains [7]. The METR trial has a small sample, a wide confidence interval, and is disavowed as current by its own authors [10, 11]. DORA is correlational survey research and cannot establish that AI adoption causes instability [9].

The largest problem is that the code review effectiveness literature mostly predates AI-generated code and may not transfer at all. Machine-generated output differs from human-authored code in both error distribution and surface plausibility, and surface plausibility is exactly the dimension review depends on. A reviewer's instinct for "this looks wrong" was trained on code written by people making human mistakes. I would not assume it survives the change of author.

4.3 The cadence may be over-engineered

Seven rules is a lot of process for one person. The realistic failure mode is not that the design is wrong but that it gets abandoned in week three, most likely by the daily window being skipped on a busy day and then never resumed. If only one rule survives, it should be rule 2. Everything else here is optimization around it.

4.4 It assumes calendar control

The cadence presumes substantial authority over the week it governs. Client work, incidents, and shared-team obligations all erode that. The slack half-day is a partial answer. It is not a complete one.

5. Validation protocol

Given the measured gap between perceived and actual productivity effects in AI-assisted work [10, 11], self-report should not settle whether this cadence works. A minimal N-of-1 design:

Duration. Four weeks of the cadence against four weeks of prior practice, alternating fortnights to blunt seasonality.

Captured automatically. Deployments per week. Lead time from first commit to production. Change failure rate, meaning deployments that required a fix or revert. Median age of a change at merge. Number of changes reverted or substantially rewritten after review.

Recorded daily, one line. Longest uninterrupted build block, number of mode switches, and whether the verify window happened.

Recorded weekly, one line. Subjective focus on a 1 to 7 scale, in the manner of developer self-monitoring instruments [5, 6].

Decision rule. The cadence is working if change failure rate falls while deployment frequency holds or rises. If instability rises, the cadence is amplifying the pattern in section 3.3 no matter how good the weeks feel. If the verify window was skipped on more than three days in a fortnight, the result is uninterpretable, because the intervention was not actually run.

6. Conclusion

The cadence rests on two claims, one strong and one contingent.

The strong claim is that cognitively dissimilar work modes should be separated in time rather than interleaved. The switching-cost and attention-residue literature supports this well, developers' own accounts of productive days corroborate it, and the meeting-containment evidence is consistent with it.

The contingent claim is that verification of AI-generated output has to close on a daily loop rather than a weekly one. That follows from a sampling-rate argument and from queueing effects. It is also the element most likely to be either the cadence's decisive advantage or the point where it turns out to be unnecessary overhead, which is why the validation protocol is built to test it.

Keep the day-level modes. Verify daily. Then measure change failure rate instead of trusting how the week felt.

References

Every reference below was checked against a primary or authoritative secondary source: a publisher page, arXiv record, DOI resolver, or the authors' own institutional listing. Volume, issue, and page details are as verified. Where a detail could not be verified it is omitted rather than guessed, and the one incomplete entry is flagged.

  1. Leroy, S. (2009). Why is it so hard to do my work? The challenge of attention residue when switching between work tasks. Organizational Behavior and Human Decision Processes, 109(2), 168-181.
  2. Rubinstein, J. S., Meyer, D. E., & Evans, J. E. (2001). Executive control of cognitive processes in task switching. Journal of Experimental Psychology: Human Perception and Performance, 27(4), 763-797. doi:10.1037/0096-1523.27.4.763
  3. Mark, G., Gonzalez, V. M., & Harris, J. (2005). No task left behind? Examining the nature of fragmented work. Proceedings of CHI 2005, Portland, OR, 321-330. doi:10.1145/1054972.1055017
  4. Mark, G., Gudith, D., & Klocke, U. (2008). The cost of interrupted work: More speed and stress. Proceedings of CHI 2008, Florence, Italy, 107-110. Some bibliographies list the title as "More Speed, More Stress."
  5. Meyer, A. N., Fritz, T., Murphy, G. C., & Zimmermann, T. (2014). Software developers' perceptions of productivity. Proceedings of FSE 2014, Hong Kong, 19-29. doi:10.1145/2635868.2635892
  6. Meyer, A. N., Barton, L. E., Murphy, G. C., Zimmermann, T., & Fritz, T. (2017). The work life of developers: Activities, switches and perceived productivity. IEEE Transactions on Software Engineering, 43(12), 1178-1193.
  7. Laker, B., Pereira, V., Budhwar, P., & Malik, A. (2022). The surprising impact of meeting-free days. MIT Sloan Management Review. sloanreview.mit.edu
  8. Forsgren, N., Humble, J., & Kim, G. (2018). Accelerate: The Science of Lean Software and DevOps: Building and Scaling High Performing Technology Organizations. IT Revolution Press.
  9. DORA / Google Cloud (2025). State of AI-assisted Software Development. dora.dev
  10. Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. arXiv:2507.09089. arxiv.org
  11. Becker, J., Rush, N., Cunningham, T., Rein, D., & Mahamud, K. (2026, February 24). We are changing our developer productivity experiment design. METR. metr.org
  12. Bacchelli, A., & Bird, C. (2013). Expectations, outcomes, and challenges of modern code review. Proceedings of ICSE 2013, 712-721.
  13. Rigby, P. C., & Bird, C. (2013). Convergent contemporary software peer review practices. Proceedings of ESEC/FSE 2013, 202-212. doi:10.1145/2491411.2491444
  14. Sauer, C., Jeffery, D. R., Land, L., & Yetton, P. (2000). The effectiveness of software development technical reviews: A behaviorally motivated program of research. IEEE Transactions on Software Engineering, 26(1), 1-14.
  15. Menzies, T., Nichols, W., Shull, F., & Layman, L. (2017). Are delayed issues harder to resolve? Revisiting cost-to-fix of defects throughout the lifecycle. Empirical Software Engineering, 22(4), 1903-1935. doi:10.1007/s10664-016-9469-x
  16. Maddila, C., Upadrasta, S. S., Bansal, C., Nagappan, N., Gousios, G., & van Deursen, A. (2023). Nudge: Accelerating overdue pull requests toward completion. ACM Transactions on Software Engineering and Methodology. Preprint: arXiv:2011.12468.
  17. Zhang, X., Yu, Y., Wang, T., Rastogi, A., & Wang, H. (2022). Pull request latency explained: An empirical overview. Empirical Software Engineering, 27(6). doi:10.1007/s10664-022-10143-4
  18. Jureczko, M., et al. (2020). Code review effectiveness: An empirical study on selected factors influence. IET Software, 14(7). doi:10.1049/iet-sen.2020.0134. Co-author list not verified, so cite via DOI.
  19. Reinertsen, D. G. (2009). The Principles of Product Development Flow: Second Generation Lean Product Development. Celeritas Publishing, Redondo Beach, CA. ISBN 978-1935401001
  20. Little, J. D. C. (1961). A proof for the queuing formula: L = lambda W. Operations Research, 9(3), 383-387.
  21. Fagan, M. E. (1976). Design and code inspections to reduce errors in program development. IBM Systems Journal, 15(3), 182-211. doi:10.1147/sj.153.0182
  22. McIntosh, S., Kamei, Y., Adams, B., & Hassan, A. E. (2016). An empirical study of the impact of modern code review practices on software quality. Empirical Software Engineering, 21(5), 2146-2189.