starl3xx.fun / algo / 004
Issue № 004 (commit 8b25829)
A holdout that hides a post from a fixed share of its readers once it collects likes, the dwell-regret scorer deleted whole, a cold-start window doubled to 48 hours and now enforced for every viewer, twelve new measures of how repetitive a feed is, and legal withholding surfaced in Under the Hood
Every Sunday I diff the open-source X ranking code, xai-org/x-algorithm, against last week’s mirror and write down what actually moved. This issue covers 6bb4594 to 8b25829, which is 5 commits, dated 15 September, 16 September, 17 September, 18 September.
Everything below is a default in the code, not a claim about production. The repo’s own README says the two can differ. Every number is copied from the diff; where I’m inferring what a change means for posting, I say so.
If you change one thing this week
- Change nothing for the new holdout, and expect to see it called throttling when someone notices it.
EnableFavHoldouthides a post from a fixed share of the people who would otherwise have seen it, and the share grows with the likes the post has already collected:2percent just above one,15percent above three thousand. It defaults tofalse, so nothing is held out today. The held-out readers are picked by hashing the post against the reader, so if it is ever turned on a given person never sees that post, rather than sometimes seeing it. - Give a new account’s post two days, and stop counting after that.
ColdStartMaxPostAgeSecsdoubled to172800seconds, and the exemption that let every experiment arm but one ignore the deadline was deleted, so a cold-start candidate past two days is now dropped for every viewer rather than for one arm only.
Everything below is the evidence for those two, plus what else moved.
What moved
1. A new filter hides a post from a fixed share of the readers it would have reached. fav_holdout_filter.rs is new, 166 lines with its tests. holdout_percent reads a post’s like count against a twenty-three rung ladder and returns the rung for the highest threshold that count clears: 0 at one like or fewer, 2 above one, 5 above ten, 7 above seventy-five, 10 above three hundred, 13 above one thousand, 14 above two thousand, and 15 above three thousand, where it stops climbing. Which readers land in that share is not a coin flip per feed build: inventory_holdout_filter.rs already carried is_held_out, which mixes the post id and the viewer id into one hash, takes it mod 100, and drops the post when the bucket sits under the percentage. Both that and holdout_bucket went from private to pub(crate) this week so the new filter could call them. EnableFavHoldout is new in param.rs and defaults to false, and one of the filter’s own tests is named for a treatment and control table, which is what a holdout of this shape is usually for. Takeaway nothing to do this week, and if it is ever switched on it works as a ceiling on a post’s audience rather than as a penalty on the post, because the held-out readers are chosen before anyone has seen it
2. The dwell-regret scorer was deleted, and “weighted” stopped being a setting. value_model_gate.rs is gone, all 537 lines of it, ranking_scorer.rs shed 497, and mod.rs no longer declares the module. Twenty-three parameters went with it: the whole DwellRegret family, including a nineteen-feature logistic gate whose coefficients, bias 1.033918 and threshold -0.634264 were spelled out in param.rs as one string, and which read your posting history rather than the post, down to account_age_years and days_since_last. ValueModelMode is removed too, and reranking_kafka_side_effect.rs now reports the mode as a constant instead of reading the switch. The two other values it could hold, dwell_regret_sigmoid and gated_dwell_regret, have no code behind them any more. Takeaway the weighted sum of predictions everyone reads off this repo was one of three possible scoring paths a week ago and is the only one now, so the published ladder describes more of the feed than it used to
3. The clickbait penalty was swapped for a per-impression one, and neither is on. In ranking_scorer.rs, low_fav_penalized_click_dwell is gone and click_dwell_term replaces it. The old one took the predicted click-dwell, divided the predicted like rate by a baseline of 0.01, raised that to 0.5, clamped it between 0.01 and 1.0 and multiplied: a post that held a reader but was not going to be liked lost most of its dwell credit. The new one multiplies the predicted click-dwell by the predicted click instead, which states dwell per impression rather than per click. Five parameters went out with the old shape and EnableCdwellOnImpr came in at false, so the ladder itself did not move. The stats key renamed alongside it, gate.click_dwell_low_fav_rate_penalty becoming gate.cdwell_on_impr. Takeaway a low like rate was not costing anyone dwell credit before this week and still is not, so the folklore about writing for the like rate has no default behind it either way
4. The cold-start window doubled, and it now expires for every viewer. In param.rs, ColdStartMaxPostAgeSecs went from 86400 to 172800, one day to two. The larger change sits beside it in author_cold_start.rs: cold_start_freshness_eligible, which returned true for anyone outside the treatment arm and so exempted them from the deadline entirely, was deleted, and apply_cold_start now measures every candidate against params.max_post_age whatever arm the viewer is in. The test asserting the old exemption, control_ignores_max_post_age, was deleted rather than updated, which is the clearest signal the exemption was the thing being removed. The treatment arm also stopped insisting the candidate arrive by the mixture-of-experts retrieval path: cold_start_corpus_eligible lost its is_phoenix_moe clause, so any post by an author in the treatment corpus is eligible. PhoenixColdStartMaxResults is new at 0. Takeaway a post from a small account now has two days in the cold-start lane instead of one, and the lane shuts on the second day for every reader rather than staying open for most of them
5. Twelve new counters measure how repetitive a feed is, and nothing scores them. composition.rs is new: Composition::from_counts returns how many distinct keys a slate holds, the largest single share, a Herfindahl concentration index and a normalized entropy. tweet_type_metrics.rs gained twelve predicates built on it. Nine are about authors, among them unique_author_lte_5, single_author_gte_50_pct, author_repeat_gte_3_in_slate and author_not_engaged_by_viewer. Three are about semantic ids, the hierarchical topic codes now reachable on a candidate through semantic_id_prefix in candidate.rs: has_semantic_ids, sid_l1_repeat_in_slate and sid_l2_repeat_in_slate. response_diversity_stats_side_effect.rs is new and records the same shape at three points, the final order, the top 10, and a pre-heuristic ordering, on 5 percent of feed builds unless EnableResponseDiversityStatsExperimentBucket puts the viewer in a bucket. Every one of these is a metric or a side effect; no scorer and no filter reads any of them. Takeaway X is now counting how often one author and one topic repeat inside a single feed, and instrumenting a thing is usually what happens before it is optimized, so this is the one to watch rather than the one to act on
6. Under the Hood now shows when a post was withheld under a legal demand. The repo’s own README.md added a notable update dated 18 September: reports now say whether an account or its posts had visibility limited to comply with law, including which country a post was withheld in. A takedowns directory appeared in the component table to go with it, producing the takedown-reason list the visibility rules read, and merging a post’s own reasons with the ones attached to its author’s account. Takeaway if reach fell in one country and nothing about your posting changed, this is the first place in the open code where the answer is reportable rather than guessable
Also in this range, and not worth an item each: PhoenixRetrievalMOEInferenceClusterId moved from "Experiment2Memy04" to "Experiment3Memy04", a third deployment target for the retrieval model rather than a change in what it does, and the ranking side is still untouched. EnableAdsBrandSafetyVerdictV2 went from true to false, turning off an ads brand-safety verdict that was on last week. A Brazil 2026 election filter arrived at 143 lines. The visibility-filtering service dropped its entire bundled twemcache client, around 3245 lines across eleven files, in favor of a composite tweet-entity hydrator; that is plumbing, and it changes no ranking. On the labeling side, the reply-spam scorer was adjusted and a new Grok flow was added for spam that hides adult media inside video.
The standing numbers
For anyone new here, read from param.rs at this commit: a share via copy link 20.0, a reply between two accounts that follow each other, on top of the reply weight 15.0, a reply 5.0, a quote 5.0, a like 0.5, each extra post of yours in one viewer’s feed build 0.5, and that decay stops here 0.25. From config.rs, MAX_POST_AGE is 48 hours, so nothing older than that enters the For You feed at all.
Caveats
These are defaults in open-source code, overridable per experiment at runtime. The ranking model itself (Phoenix) is trained, not rule-based; the weights blend its per-viewer predictions and don’t describe what it learned. Reading weights as tactics is inherently lossy. I’d rather you know that than not.
Sources: param.rs, config.rs, fav_holdout_filter.rs, inventory_holdout_filter.rs, ranking_scorer.rs, author_cold_start.rs, composition.rs, tweet_type_metrics.rs, response_diversity_stats_side_effect.rs, candidate.rs, README.md, all at 8b25829.