Skip to main content
Editorial reference only. Independent editorial knowledge base on digital journalism. No accreditation, no qualification, no award of any kind, and no guarantee of employment or publication.
Newsroom Horizon

Theme 03 · Distribution

Recommender systems and what readers are shown

A recommender is an editorial instrument with no editor in the room. It expresses a publisher’s priorities in code, and those priorities are usually invisible until somebody measures what the system actually promotes.

Scope note. This theme page describes practice and public rules. It is general editorial information, not legal advice, and it confers no qualification of any kind.
01

1. What the system is optimising

Every recommender maximises something measurable. The choice of objective is the whole editorial argument, compressed into one line of configuration. A system tuned for click-through will favour headlines that create curiosity gaps. A system tuned for completed reading time will favour longer pieces that hold attention, including some that few people open. A system tuned for return visits within seven days will favour series, follow-ups and habitual formats. None of these objectives is dishonest, and none is neutral. The mistake is adopting a default objective from a vendor and then treating the resulting front page as an editorial judgement rather than an arithmetic one.

02

2. Explicit and implicit signals

Explicit signals are things a reader states: a followed topic, a saved article, a chosen newsletter, a dismissed recommendation. Implicit signals are things a reader merely does: what they opened, how far they scrolled, how long the page stayed in view, what they returned to. Implicit signals are abundant and cheap, which is why most systems lean on them, but they are also poor evidence of preference. Opening an article can mean interest, mistrust, obligation or a mis-tap. Reading to the end can mean absorption or confusion. Systems built almost entirely on implicit signals tend to learn what is hard to resist rather than what is worth reading.

03

3. Candidate generation, ranking and re-ranking

Most production systems work in three stages. Candidate generation narrows the whole archive to a manageable set, usually by recency, section, topic similarity or collaborative patterns. Ranking scores those candidates against the objective. Re-ranking then applies constraints: do not show the same story twice, cap the number of items from one section, insert an editorial slot at position three, suppress items subject to a legal hold. The re-ranking layer is where most editorial policy can actually be implemented, because it is expressed as rules rather than as learned weights, and because it can be read, reviewed and changed by people who do not build models.

04

4. Cold start, new work and the recency trap

A newly published item has no engagement history, so a purely behavioural system has nothing to score it on. The usual fixes are content-based similarity, an exploration budget that shows a proportion of new items to a random slice of readers, and hard editorial slots that bypass ranking altogether. Without them, a system quietly converges on established formats and known topics, and original reporting that does not resemble anything already popular is buried by its own novelty. The same dynamic affects new sections, new reporters and coverage of places that have historically received little attention.

05

5. Feedback loops and popularity bias

Recommenders learn from behaviour that they themselves shaped. An item promoted into a prominent slot accumulates engagement, which raises its score, which prolongs its prominence. The system then reports high engagement as evidence that its choice was correct. Breaking that circularity requires deliberate measurement: compare engagement against position, hold out a control group that receives chronological or editorially ordered results, and track how much of total attention the top few items absorb. If the top five items take a very large share of all reading, the system is concentrating attention rather than distributing it, whatever the aggregate numbers look like.

06

6. Editorial override and curation slots

Automated ranking and human curation are not alternatives; the workable arrangement is a hybrid with clear boundaries. Reserve fixed positions for editorial choice, and define them in advance rather than negotiating them daily. Give the desk a pin, a suppress and a legal-hold action, each logged with a reason and a person. Set floors as well: at least one item from public-interest coverage in the top band, at least one item outside the reader’s dominant topic. Floors are more durable than exhortation, because they survive busy days and staff changes.

07

7. Measuring diversity of exposure

Diversity is measurable if it is defined first. Useful measures include the number of distinct sections a reader is exposed to in a session, the share of readers who see at least one public-interest item per week, the concentration of impressions across the catalogue, and the overlap between the sets of items shown to two randomly chosen readers. Each captures something different, and each can be gamed in isolation. Reported together, they describe whether the system is producing many small private front pages or a shared one with personal edges, which is a question a newsroom should be able to answer with numbers.

08

8. Transparency about ranking

The Digital Services Act requires certain online platforms to set out, in their terms and conditions, the main parameters used in their recommender systems and the options available to change them, with very large platforms subject to further duties including a non-profiling option. Whatever a publisher’s own legal position, the practice is worth adopting: a short, plain page explaining what the site’s ranking uses, what it deliberately ignores, where editorial choice overrides it, and how a reader can reset or switch to a chronological view. Read the current text of the instrument before making claims about its scope.

Comparison

Signals and what they actually tell you

Common ranking signals, their reliability and their characteristic distortion
SignalTypeWhat it evidencesCharacteristic distortion
Click-throughImplicitHeadline pull at the moment of displayRewards curiosity gaps over substance
Scroll depthImplicitPage traversal, not comprehensionLong pages score well even when abandoned mid-way
Attention timeImplicitTime the item held the viewportConfuses difficulty with interest
Return within 7 daysImplicitHabit formationFavours serialised and routine formats
Followed topicExplicitStated interestAges badly; readers rarely revisit their settings
Saved for laterExplicitPerceived valueSparse; many readers never use it
DismissalExplicitStated rejectionUnder-collected, so negative preference is weakly learned

No single signal is sufficient. The useful question is not which signal is best but which distortion the current mix is producing, and whether anyone is measuring it.

Desk checklist

Auditing your own recommender

  • The optimisation objective is written down in plain language and dated.
  • Someone outside the engineering team can explain what the system rewards.
  • A holdout group receives chronological or editorially ordered results for comparison.
  • Impression concentration across the catalogue is measured, not assumed.
  • New items receive an exploration budget rather than competing on history alone.
  • Editorial pin, suppress and legal-hold actions exist and are logged with a reason.
  • A reader-facing page explains the main ranking inputs in plain language.
  • Readers can switch to a non-personalised view without creating an account.
  • The audit is repeated after every objective change, not only annually.

An audit that only runs when the system is redesigned will miss drift, which is how most ranking problems arrive.

Questions readers ask

Is algorithmic ranking incompatible with editorial values?

No, but it encodes them whether or not anyone intends it to. The objective function, the re-ranking rules and the editorial slots are all value choices, and they should be made deliberately and reviewed.

Should a news site offer a chronological feed?

Many publishers do, as a simple way of letting readers see the full output without ranking. It is also a useful internal control, because it provides a baseline against which the ranked feed can be compared.

Does the Digital Services Act apply to a small news website?

Scope depends on the service and its role, and the heaviest duties fall on very large platforms and search engines. Read the current text and take qualified advice rather than relying on a summary.