Weekend Reading #97

Weekend Reading: A weekly roundup of interesting Software Architecture and Programming articles from tech companies. Find fresh ideas and insights every weekend.

This week: bool.dev examines the controls needed before coding agents receive write access; Meta turns resource placement into explicit constraints and objectives; Uber addresses training-serving skew by capturing real inference features; and LinkedIn replaces blanket capacity buffers with demand modeling and backtesting.

AI Across the SDLC, Part 3: What to Build Before You Give Agents Write Access

👉 For platform engineers and tech leads introducing coding agents

AI Across the SDLC, Part 3: What to Build Before You Give Agents Write Access

Giving an agent repository access changes the problem from generating useful code to controlling consequential actions. This article lays out six foundations: short-lived task identities, policy enforcement outside the model, isolated execution, reviewed MCP dependencies, deterministic verification, and action tracing. The distinction is important: an agent can propose an operation, but it should not control the rules that authorize it or approve its own changes.

The practical starting point is a narrow credential and an independently enforced boundary, not a longer system prompt. Sandboxing limits execution, while authorization limits what execution can accomplish; neither replaces the other. Expand autonomy only when denied calls, reversions, and review effort show that the existing scope is manageable.

Open-Sourcing Rebalancer: A Generic, High-Performance Library for Solving Assignment Problems

👉 For infrastructure engineers building schedulers and placement systems

Open-Sourcing Rebalancer: A Generic, High-Performance Library for Solving Assignment Problems

Meta’s Rebalancer expresses placement problems through objects, destinations, resource dimensions, constraints, and optimization objectives. After nine years of internal use, it handles roughly 40 million assignment problems daily. The library supports both mixed-integer programming and optimized local search: the former helps establish optimal baselines, while the latter makes much larger production problems tractable.

The reusable lesson is to separate the placement model from the solving strategy. Capacity and fault-domain requirements are hard boundaries; balance and movement costs are competing preferences. Making that distinction explicit lets teams change algorithms without silently changing correctness. Rebalancer’s inspection tools also highlight an operational requirement: engineers need to understand which constraint prevents a move, not merely receive a new assignment.

Taming the ML Firehose: Scaling Feature Consistency

👉 For ML platform engineers building online inference and training pipelines

Taming the ML Firehose: Scaling Feature Consistency

Uber tackles training-serving skew by logging the feature values actually used during inference instead of reconstructing them later. At eight million requests per second, naive logging would generate around 1.7 PB daily. Feature allowlists and compact identifiers reduce payloads; Kafka and Flink connect those records with impression events. Uber reports that mismatches on tracked key features fell from over 10% to zero.

The architectural gain comes from preserving serving-time evidence, not simply accelerating an offline pipeline. But filtering for observed impressions also determines which examples reach training. A useful adoption check is therefore twofold: verify feature equality and examine whether delayed or missing impressions distort the retained population. Smaller teams can apply those checks before investing in comparable streaming infrastructure.

Infrastructure efficiency: How a classic math problem helped reduce our capacity buffers

👉 For SREs and platform teams balancing capacity costs against shortage risk

LinkedIn adapts the newsvendor inventory model to infrastructure capacity: reserve enough resources to balance the cost of shortages against unused supply. Its approach separates planned demand changes from uncertainty, accounts for provisioning lead time, and backtests forecasts. Demand headroom fell by about 3.5%; combined improvements reduced platform buffers by roughly 8.5%.

This does not justify cutting every service’s reserve by the same percentage. Streaming workloads, deferrable batch jobs, migration overlap, and correlated failures have different consequences. The useful first step is to replay historical demand and supply losses against a proposed buffer. A smaller reserve is defensible only when the model covers the time needed to acquire or recover capacity.

Related reading on bool.dev: What is the Model Context Protocol (MCP) explains the host, client, and server boundaries behind the agent write-access discussion.


Tags:


Comments:

Please log in to be able add comments.