starl3xx.fun / algo / 002

Issue № 002 (commit 902a06f)

No weight moved. Retrieval switched to long dwell, the cold-start boost narrowed, and nine visibility-filtering rule files became three without changing what they do

X Algo Watch

Every Sunday I diff the open-source X ranking code, xai-org/x-algorithm, against last week’s mirror and write down what actually moved. This issue covers the previous mirror at bc8e5f0 (28 August) through 902a06f, dated 4 September. Ten files net, 2,039 to 2,049.

Start with what did not happen: not one scoring weight changed. Every number on the ladder is where it was a week ago. This was a plumbing week, and the interesting parts are one default, one narrowed eligibility check, and a rewrite that was careful to change nothing.

Everything below is a default in the code, not a claim about production: the README describes defaults as production values kept in sync by a cron job, and describes experiments running on top of them, so what any given person gets is the default plus whatever arm they are in. Every number is copied from the diff; where I’m inferring what a change means for posting, I say so.

If you change one thing this week

  • Keep writing for the pause. dwell now steers retrieval as well as ranking, and retrieval decides whether a post enters the pool at all, so a post that holds someone is favored at the earlier and harsher stage.
  • Don’t pile replies into one thread. A new reply-spam classifier reads up to eleven posts of a chain at once and returns a set rather than a verdict, and it throws the flag away when only one post is named, so it is built to catch a pattern and not a stray reply.
  • If you’re under 1k, earn it in-feed on day one. The cold-start boost narrowed again for a slice of viewers, on top of the top-k that went to 2 last week.

Everything below is the evidence for those three, plus what else moved.

What moved

1. Retrieval now runs on long dwell too. In param.rs, PhoenixRetrievalAggregationType moved from DENSE_WITH_SHORT_DWELL to DENSE_WITH_LONG_DWELL. Its ranking-side twin, PhoenixAggregationType, made that same move last week and is untouched here, so the two halves of the pipeline now share one definition of attention. This is the more consequential of the pair. Ranking decides where a post lands once it is in the pool; retrieval decides whether it gets into the pool at all. A post that holds someone is now favored at the earlier, harsher stage. Takeaway keep earning the pause; dwell now decides whether you get into the pool, not only where you land in it

2. The cold-start boost narrowed, for some viewers. In author_cold_start.rs, a candidate in the treatment viewer arm must now also satisfy is_phoenix_moe, meaning its served_type is ForYouPhoenixRetrievalMoe. Belonging to the treatment author corpus used to be enough. The control and holdout arms are untouched, and no default in that file changed: most of the rest of the diff is a refactor that lifts thirteen scattered params.get calls into one ColdStartParams struct. One real change hides in that refactor, and it is about logging rather than ranking: author-corpus bucketing now runs before the early return on EnableViewerColdStart, so both arms record the same experiment impressions whether or not the feature is on. Two new tests, both_arms_log_the_same_codivert_impressions and both_viewer_arms_log_the_same_author_impressions, pin that. For accounts under the 1,000-follower cap, the exploration budget is now narrower again for a slice of viewers, on top of the cold-start top-k that narrowed to 2 last week. Takeaway if you’re under 1k, what you earn in-feed on the first day is what you get; traffic from outside X still buys you nothing

3. The visibility rules were rewritten, and the rewrite changes nothing. Nine modules in rules (nsfw_age_gating.rs, nsfw_interstitial.rs, nullcast_rule.rs, socialgraph_rules.rs, tes_rules.rs, tweet_flag_rules.rs, tweet_label_drops.rs, user_label_drops.rs, user_rules.rs) were collapsed into three: author_rules.rs, tweet_rules.rs and a new rule_spec.rs. The policies are now static RuleSpec tables instead of boxed trait objects. I want to be precise about how I know it changes nothing, because this is a big diff to wave through: about 5,750 changed lines inside rules alone, and roughly 7,600 across visibility-filtering once config.rs, the safety-label codec, the hydration layer and the whole twemcache tree are counted. There is a test in registry.rs called wired_rule_order_matches_pre_migration_sequence, and it pins a hardcoded ordered list of rule names, which means it proves the new wiring matches what someone wrote down rather than what the old code did. So I checked the old wiring by hand: base_home_rules() resolved to 28 rules in an order, the out-of-network list to 26, and the new tables reproduce both, name for name, in sequence. The golden-corpus fixtures gained new Allow cases and changed no existing verdict. Sometimes the report is that a very large diff is a refactor. Takeaway nothing to act on, and that is the finding: a diff this size can be a refactor, and this one is

4. One thing in that rewrite did become configurable. A new params.rs turns the NSFW age-gating country list, previously a hardcoded NSFW_GATING_COUNTRIES constant, into a hot-reloadable rust_vf_nsfw_gating_countries parameter, with the same sixteen-country default: ar au br ca de es fr gb id it kr mx nl ph pt th. The file also carries a drift counter that compares the live value against the old Scala config. The default did not move, so nothing changed today, but a list that used to require a deploy to change can now change between two refreshes of your feed. Takeaway nothing to act on, but a list that used to need a deploy can now change between two refreshes, so watch it

5. A new reply-spam classifier reads the whole thread. In reply_spam, a new classifier, classifier_multi_step_reply_spam.py, hands a reply chain to a language model and gets back a set of post IDs it judges spam, not a verdict on one reply. MAX_THREAD_SIZE is 10, and that caps the ancestors: _thread_posts takes the first five and last five, then appends the reply being scored, so the model sees up to eleven posts. Three filters run on that set afterwards, and they are the interesting part. _drop_singleton throws away the whole flag if only one post was named. _drop_other_author_posts keeps only posts by the author of the reply being scored. _drop_if_newest_reply_not_flagged throws away the flag if the reply that triggered the check is not in the set. Its output goes to reply ranking, via TaskWriteMultiStepReplySpamReplyRanking, not to the For You feed. This is built to catch a pattern, not a post. Takeaway one off-topic reply is explicitly not what this fires on; several from you in one thread is

Five smaller things in the same commit. The biggest single file change in home-mixer is data, not logic: brazil_2026_election_filter.rs had its account list refreshed again, and its asserted length moved from 2315 to 2554. AdsBlenderType moved from partition_organic_low_risk to multi_risk, which is ads placement and sits outside organic ranking. Two new parameters ship inert: EnablePhoenixScoreStatsExperimentBucket defaults to false and PhoenixExperimentOverrides to an empty string. candidate.rs and phoenix_scorer.rs gained a backbone_scores field carried alongside the existing scores, and as_score_info was renamed as_score_info_no_prediction_scores and now sends an empty prediction_scores back rather than the candidate’s own. And the slate-context struct gained two more fields nothing reads yet, exact_k and exact_gap, joining recon_cos_milli, recon_count_above and recon_gap_above from last week. Takeaway that is the same replacement flagged in issue 001, four weeks of fields deep and still switched off; still the thing to watch

The standing numbers

Nothing on the ladder moved this week, so it reads exactly as it did in issue 001. From param.rs, each one against a like:

  • share via copy link: 20.0, forty times a like
  • reply between two accounts that follow each other: 5.0 plus a +15.0 boost, also forty times
  • reply or quote: 5.0, ten times
  • like: 0.5, the unit

Each extra post of yours in one viewer’s feed build is multiplied by 0.5, compounding to a floor of 0.25. Nothing older than 48 hours enters the For You feed. A reply or a repost from an account a viewer doesn’t follow is dropped outright, so a reply earns its weight from your existing followers rather than from new reach; a quote carries the same weight and does travel out of network.

Caveats

These are defaults in open-source code, and an experiment arm can override any of them for any given viewer. The ranking model itself (Phoenix) is trained, not rule-based; the weights blend its per-viewer predictions and don’t describe what it learned. X publishes squashed drops rather than its real history, so a commit here is a range, and I can read what the code says a good deal more reliably than I can read why it says it. Reading weights as tactics is inherently lossy. I’d rather you know that than not.

Sources: param.rs, author_cold_start.rs, candidate.rs, registry.rs, params.rs and the reply_spam directory, all at 902a06f.