A&A INSIGHTS
Existing articles, missing search results: learning from Spotify’s query reformulations
For specialist publishers and course owners: use Spotify’s historical search work to examine failed queries, reformulations and destinations. Separate vocabulary gaps from missing content, retain title search, and set clear criteria for a limited semantic-search trial.
日本語で読む
THE STARTING POINT
If staff keep sending readers links to articles that already exist, examine search failures before commissioning more content. Drawing on Spotify’s 2021 and 2022 firsthand publications, this article proposes reviewing the initial query, its reformulation and the destination together. Preserve searches for known titles and evaluate problem-based searches separately. Applying this approach to a small publisher is an A&A proposal, not a delivered client project or evidence of increased sales.
Separate missing content from mismatched vocabulary
Hypothetical example
Consider a hypothetical specialist publisher. A reader types “no reply after sending an estimate,” but the existing article is titled “Designing post-meeting follow-up.” After failing to find it, the reader asks support and receives that URL. Another reader searches for “estimates for overseas customers,” which the archive does not cover. Both searches fail, but one calls for examining the route to existing material and the other for examining a content gap.
A&A perspective
A&A proposes not turning every zero-result term directly into an assignment for a new article. Group the initial wording, the rewritten query and the article eventually opened. A route that reaches existing material becomes a candidate for improving descriptions or retrieval. Opening a page does not establish resolution, however. The reader may have changed goals. An editor should compare the question with the actual article before retaining the sequence as a candidate for the same task.
What Spotify’s reformulations reveal
From the sources
Spotify’s March 17, 2022 engineering article describes pairing the query before a successful reformulation with the episode reached. It also retained existing retrieval and added semantic retrieval as another candidate source, recognizing the value of exact term matching.
From the sources
A separate research article dated August 30, 2021 reports that podcast searchers spent more time entering and rewriting queries than music searchers. It records differences in behavior across search targets.
A&A perspective
What a small publisher can borrow is a way of observing the target and the search behavior separately, rather than a large training infrastructure. Here, A&A proposes separate evaluation groups for known-title searches and questions expressed as a problem. That is our proposed distinction, not an equivalence with Spotify’s music-versus-podcast categories. The two publications describe separate historical research and design work; neither establishes Spotify’s current architecture.
Check a reformulation before treating it as a correct answer
A&A perspective
Keep the initial query, reformulation, destination article ID, the passage that answers it and any unresolved conditions in the review record. A similar title is insufficient if the article concerns a different audience or period. Separate arrivals at outdated material and cases where a reader wanting a free explanation receives only a paid-course introduction. Treating clicks as correct answers would encourage the system to recommend material that was found but could not be used.
Hypothetical example
In the hypothetical example, an editor checks the change from “no reply” to “post-meeting follow-up” against the existing article. Keep it as a candidate if the body explains what the next contact should clarify; reject it if it merely explains taking minutes. If search history is unavailable, prepare evaluation questions from actual inquiries after removing personal information, or let editors write hypothetical questions. Label invented questions as such and store them separately from observed reader searches.
Retain title search and choose where semantic retrieval belongs
A&A perspective
A&A proposes comparing candidates in one category before replacing search across the site. For inputs containing a known title, course name or article number, check that current destinations remain reachable. For problem-language inputs, test adding articles that answer the question despite using different words. Merge overlapping candidates by article ID, and check that an explicitly named title is not displaced by another article that is only broadly similar.
A&A perspective
If reformulations cluster around a few fixed expressions, aliases or better descriptions may be worth testing first. If readers use many different expressions that an editor can consistently connect to existing articles, there is a concrete case for comparing semantic retrieval. More candidates alone cannot fix absent articles, a mismatch in intended access or outdated content. The owner should be able to choose which kind of mismatch warrants spending before approving a search investment.
Evaluate what happens after the result is found
A&A perspective
Before the trial, record the share of evaluation questions that reach an article containing an answer, separately for title and problem searches. Compare current retrieval with the proposed addition on the same questions, with an editor checking the answer passage. Include questions held apart from those used to develop the change. In a live trial, also examine waiting time, repeated searches for the same task and requests for staff to send a URL. Define retention scope and duration, and keep inputs containing personal information out of the training worksheet.
A&A perspective
Continue only if problem searches improve, missed known-title results and waiting time stay within agreed limits, and review and maintenance remain manageable. Include embedding creation, index updates, usage fees and checks of wrong candidates in the comparison. Track paid-course applications as a separate event after search rather than equating more views with more sales. Fewer inquiries may also mean readers gave up, so interpret that count alongside the suitability of the destinations.
Choose a search repair or a new article
A&A perspective
At the next editorial meeting, sort failed searches into three groups. Questions that can be linked to existing material go to retrieval or description improvements. Questions with no answer go to a decision about whether a new article is worthwhile. Questions with an answer that does not fit the reader’s circumstances go to a review of audience and access descriptions. This gives search and editorial staff distinct work instead of routing every failure back to writers.
A&A perspective
For a discussion with A&A, prepare pairs of questions and existing articles with personal information removed before comparing product names. They help distinguish retrieval work from content preparation. The related article “What to reuse when one topic becomes a blog, social post and hiring story” addresses adapting material for other channels. This article concerns the earlier decision: whether readers can reach knowledge that already exists.
Separate missing content from a mismatch in vocabulary. The clue is not just the failed query, but the reformulation and destination that follow it. Keeping known-title retrieval intact while comparing problem-based searches in one category makes it possible to decide between adding semantic retrieval and repairing headings or aliases. Start by reconnecting existing articles with readers’ questions.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Introducing Natural Language Search for Podcast Episodes
Spotify Engineering · 2022-03-17
Accessed 2026-09-20 - Neural Instant Search for Music and Podcasts
Spotify Research · 2021-08-30
Accessed 2026-09-20
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-20