The Most Important Word in an E-commerce Query

search
query understanding
Before you rank products, you have to understand the query. Decomposing a shopper’s query into its constituent parts so retrieval and ranking can act on each one, instead of on a bag of words.
Author

Srinivasan Seshadri

Published

July 14, 2026

This post is about our work at Zettata during 2013–2016. Zettata built e-commerce search as a service, and the conviction we started with — the founding premise, not something we discovered along the way — was that before you rank products, you have to understand the query. This post is about what that meant in practice: decomposing a shopper’s query into its constituent parts, so that retrieval and ranking could act on each part appropriately, instead of treating the query as an undifferentiated bag of words.

Instincts from web search, and where they broke

I came to e-commerce search from web search, carrying instincts that had served me well there. For the query-specific signals — setting aside graph-based signals like popularity and topical authority — TF-IDF combined with proximity scoring worked well enough in that setting, because the web corpus was so large that some page usually contained all the query’s words, close together. Search only needs ten good results, and at web scale you could generally count on finding them.

E-commerce search removes that cover. Even a large retailer has on the order of ten million products, not billions. Two things follow. First, very often the words of the query simply do not occur together in any single product. The engine is then forced to drop a word — and bag-of-words scoring has no idea which word is safe to drop. Second, when the engine gets the product wrong, the results are not just weaker, they are plainly wrong, in a way a shopper notices immediately.

IDF, the workhorse of term weighting, actively misleads in this setting. Consider the query mauve dress. “Mauve” is the rarer word, so IDF weights it more heavily than “dress.” But mauve is the negotiable part of that query and dress is not. A purple dress is a perfectly reasonable result if no mauve dress is in stock. A mauve scarf — which matches the “important” word — is useless. Or consider iphone case: a bag-of-words engine, weighting by rarity, lets a strong match on “iphone” carry a screen protector or a charging cable to the top of the results, even though neither one matches “case” at all. The word that actually names what the shopper wants gets outvoted by the word that happens to be rarer. The rarity of a word and its importance to the query are simply different axes, and in a small corpus the gap between them is fatal.

So the question every failed query forced on us was: when you must relax, what do you relax first? Relax “mauve” long before you’d ever touch “dress.” Answering that question requires knowing which word plays which role — which is to say, it requires decomposing the query.

The product signal

The single most important discovery was also the simplest. Shoppers, it turns out, almost always tell you what they want in plain terms. As best I recall from looking across our merchants’ query logs, well over ninety percent of queries explicitly named the product itself: blue jeans, iphone cases, washing machine detergent. We called this the product signal. It is a fact about how people phrase shopping queries, not something a system gets for free — finding it reliably, word by word, inside a query took real machinery, which I get to below — but once you know to look for it, it is almost never the word IDF would have picked. Once you know the product word, everything else in the query organizes itself around it: descriptive modifiers, brand, color, size, audience, price sensitivity.

Our goal, stated plainly, was to extract as much structure from the query as possible and match it against structured attributes of the catalog — to turn a fuzzy keyword search, as far as we could, into something closer to a structured lookup. A query like nike shoes for women wants to become: product is shoes, brand is Nike, gender is women. Every part of the query that stops being fuzzy text and becomes a structured predicate is a part that can be scored deliberately, on its own terms, instead of being swept into the same undifferentiated weighting as everything else.

How we found the product word

The machinery, honestly described, was simple — this was 2013, and we used the simple natural-language techniques of that time.

The raw material came from the catalogs themselves. Any one retailer’s catalog was too small to learn from reliably, so we crawled the product corpora of the leading retailers of the day to give ourselves a wider training set. Against that corpus, part-of-speech extraction gave us candidate product nouns — most product words are common nouns, though brand names like “iPhone” are a notable exception. We then validated candidates against the product taxonomy we maintained: a genuine product name clusters within a node of the product tree. Modifiers, meanwhile, mostly fell out naturally as adjectives.

The interesting cases were the words that could play either role. “iPhone” is a product in its own right in the query iphone 5, and a modifier in the query iphone case. Our mechanism let every word first check whether it could be a product and whether it could play the role of a modifier, and then simple positional rules arbitrated. A noun that comes before another known product noun is more likely to be the modifier: since “case” can be a product, iphone case is most plausibly iPhone describing case. In iphone 5, “5” cannot plausibly be a product, so “iPhone” takes the head role.

Noun-noun compounds needed one more piece of care. iphone case is best treated as a compound product — a unit — because within reason a Samsung case is not a good result for that query. This is worth pausing on, because it qualifies the mauve-dress lesson. Not every non-head word is relaxable. “Mauve” in mauve dress can be relaxed; “iphone” in iphone case is constitutive of what is being asked for. Part of understanding a query is knowing which of its modifiers are preferences and which are requirements, and we did not always get that boundary right.

The rest of the stack

The product signal was the most granular signal, but not the only one. We also computed a category signal, which mattered in two situations: when the product word spanned categories (a cabinet might be a filing cabinet or a storage cabinet, and the other words in the query often predispose it one way), and when there was no product word at all. We had a query classifier assigning queries to their most plausible categories from early on. Beyond product and category came the most plausible brands, colors, age groups (infant, toddler, teens, adults), gender, and so on.

Queries with no product signal were a small minority, and the glaring example was gifts — christmas gifts, gifts for dad. As best I recall we special-cased gifts, because it was by far the most popular product-less query.

One more distinction shaped how much decomposition a query needed: what I think of as soft products versus hard products. Apparel is soft — a purple dress for a mauve-dress query is a fine outcome, and the modifiers are genuinely fungible. Hardware sits at the hard end: 1/4 inch spiral phillips head screw is essentially a database row wearing a query’s clothes, and there the same protection that guards the product word has to guard the attributes too — a 1/2 inch screw is not a good result for a 1/4 inch query, no matter how well it scores on everything else. We understood this distinction, but I do not think we calibrated it, category by category, as carefully as it deserved.

What retrieval did with the parts

Decomposition only matters if retrieval and ranking act on it. In practice, we retrieved a wide set of candidates and then scored down to a final ten, folding in signals like inventory availability along the way.

The decomposed query drove that scoring. A product-signal match earned a far larger boost than a modifier match did — a hit on “dress” mattered more than a hit on “mauve” — though both were real boosts, and a full match on everything scored highest of all. That gradient, not a hard rule, is what let the system relax the least important part of a query first when the full query returned nothing or nearly nothing, which was the routine condition in a small corpus. This is what structure bought that undifferentiated term weighting could not. A system scoring every word alike, or by rarity, has no way to tell “dress” and “mauve” apart; ours did, because the decomposition told it which was which.

I don’t have a number I can quote for how much this mattered — no controlled measurement I can point to today. What I have is the anecdotal kind of evidence: a standing set of queries that would consistently trip up the home-grown search most retailers had built on top of Solr or Elasticsearch, while ours handled them without trouble. That was enough, at the time, to convince us the approach was right. It is worth naming as conviction earned by watching queries fail elsewhere, not as something I measured and can still cite.

Behavior data, and doing without it

I do not want to leave the impression that we thought linguistics could replace behavioral data. For a large retailer, click and purchase logs are gold: something like eighty percent of the business rides on the top million queries, and those map almost directly to their best-selling products. But we routinely had to prove ourselves without that gold — some merchants did not have enough of it, others withheld it as a test — and cold start made behavioral data insufficient even where it existed: a new product has no clicks yet, at any retailer, so a system leaning only on behavior fails precisely on the merchandise a retailer most wants to move.

So the corpus-linguistics-and-taxonomy machinery was the backbone, and behavioral data was the refinement layered on top when we had it. That ordering — structure first, behavior second — was forced on us by circumstance as much as chosen, but I came to believe it was the right ordering anyway.

What we never got to

The English-like queries were the clearest unfinished business. By 2016 we could see the query stream slowly drifting toward natural language, and our decomposition machinery, built for two-to-five-word noun phrases, was not the right tool for what to buy for my sister’s birthday. We knew it, and we had not solved it.

That is where our work on the query side stood when the Zettata chapter closed. The catalog side — extracting structured attributes from product data itself, which is the other half of turning search into structured lookup — was its own body of work, and deserves its own post.