Grouping a keyword list so each group becomes one page. The proper way needs ranking data; the approximation you can do in a browser is good enough to plan from.
Keyword clustering is grouping a keyword list so that each group becomes exactly one page. It is the step between a keyword tool's output — several hundred phrases — and a content plan, and it is the step that decides whether you end up with two posts competing for the same query.
Done properly it uses ranking data. Done approximately it uses the words themselves, and the approximation is good enough to plan from as long as you know where it is weak.
Why clusters matter more than keywords
Two posts targeting the same query is the most common self-inflicted SEO problem there is. Neither ranks as well as one would have, links split between them, and search engines pick whichever they prefer regardless of which you meant.
The defence is a rule rather than a judgement: one cluster, one URL. If two phrases belong to the same cluster, they belong on the same page — not on two pages that link to each other.
That also means a cluster has a job. It has a pillar, which is the page, and it has a set of phrases that page has to answer somewhere in its body.
The proper keyword clustering method, and why you probably cannot run it
Serious cluster tools group by SERP overlap: they run every phrase through a search engine and put two phrases together when the results returned for both are substantially the same pages. If Google thinks two queries deserve the same answer, they are the same page.
It is the right method because it uses the only opinion that matters — and it is the same logic behind Google's own advice against creating pages that duplicate each other's purpose. It is also expensive — one search per phrase, at scale, against an API that charges for it — which is why the tools that do it are subscriptions rather than free.
The approximation: shared words, weighted by rarity
Without ranking data, the honest fallback is to group on the words the phrases share. The naive version of this fails immediately and predictably.
Take an Instagram keyword dump. Nearly every phrase contains "instagram". Group on shared words and you get one cluster called instagram, containing everything.
The fix is to weight each word by how rare it is in the list — a measure called inverse document frequency. "Instagram" appears in every line and carries almost no information; "carousel" appears in six and carries a lot. Two phrases sharing "carousel" are far more likely to be the same page than two sharing "instagram".
Word
Appears in
Weight
instagram
48 of 50 phrases
Almost none
hashtag
12 of 50
Moderate
carousel
6 of 50
High
lamination
1 of 50
Highest, and useless alone
That last row is the counterweight. A word appearing once cannot group anything, so rarity is valuable up to the point where it is unique.
Intent splits a topic into two pages
This is the rule that catches people who have got the grouping right.
"What is a baker's percentage" and "baker's percentage calculator" share every content word. They are the same subject and they are two different pages, because the people typing them want different things — one wants to understand, the other wants a tool.
A guide ranking for a "calculator" query loses to anything that simply is a calculator. Serving both from one page usually serves neither.
Looking for something specific — login, pricing, download, docs, usually somebody's brand.
What to do with the singletons
A phrase that matches nothing should stay on its own rather than being pushed into the nearest cluster.
This feels untidy and it is the correct behaviour. A wrong grouping is worse than no grouping, because a wrong grouping is the one that ships — it becomes a post that half-answers two things. An ungrouped phrase is a decision you have not made yet, which is a much cheaper state to be in.
Some of them are genuinely their own page. Some are noise from the keyword tool. Both are easier to judge when they are listed as singletons than when they are buried in a cluster they do not belong to.
Picking the pillar
The pillar is the page the cluster becomes, and it should be a phrase from your list rather than a head term you invented.
Invented head terms read well in a plan and get no traffic, because nobody typed them. Where you have search volumes, the highest-volume phrase in the cluster is usually the right pillar; where you do not, the shortest phrase carrying the cluster's defining words is a reasonable stand-in.
Then the rest of the cluster becomes the outline. Each phrase in it is a question the page has to answer somewhere, which is a considerably better brief than a word count. Keyword density covers what to do with the phrase once you are writing, which is less than most people think.
Where clustering meets what you can actually make
A cluster list is a demand-side plan. It says nothing about whether you can make twenty videos about baker's percentages.
The other half is the supply side — the recurring themes you can sustain — and the two should agree. Content pillars covers that end, and the useful check is whether every pillar has at least one cluster behind it and every large cluster has a pillar that can feed it.
For YouTube specifically, the same clustering applies to a different corpus. YouTube keyword research covers where the phrases come from before any of this starts.
What is keyword clustering?
Grouping a keyword list so each group becomes one page. It is how you avoid writing two posts that compete for the same query.
How do you cluster keywords without SERP data?
Group on shared content words, weighted so that a rare shared word counts far more than a common one. It is an approximation of SERP clustering and it is good enough to plan from.
Why are two similar keywords in different clusters?
Usually different search intent. "What is X" and "X generator" share words and need different pages.
What is the pillar of a keyword cluster?
The page the cluster becomes. Pick a real phrase from your list — the highest-volume one — rather than inventing a head term nobody searches.
How many keywords should be in a cluster?
However many genuinely belong on one page. A cluster of thirty is usually a sign the grouping is too loose and hides sub-topics.
Paste your list into the keyword cluster splitter — a volume column pasted alongside is used to pick each pillar.
A search snippet is assembled, not published. Google picks the title, often rewrites it, chooses the description from the page, and cuts both to fit a fixed column.
Copying hashtags out of a screenshot by hand is how most people do this. Here is how to pull them off a public post from its URL, and what the limits are.
A YouTube timestamp link starts the video at an exact second. Here is the URL format, the chapter rules, and how to build both without counting seconds by hand.