Keyword Clustering
A self-hosted Keyword Insights. SERP overlap, plus the economics per cluster.
Two searches, one set of results.
That's the test. I type two searches into Google. If it hands back mostly the same pages, Google's already treating them as one question, and one question deserves one page rather than two.
Keyword Clustering runs that test across a whole list of search terms. It pulls the results for each one, looks at the top seven and puts a pair together when three of those seven match. Three's my default, and the form takes anything from one to seven. The groups fall out of that.
Wording can't tell you this. Two phrases can read almost the same and send Google somewhere completely different. The results page is what Google actually did. That's what I group on.
A subscription tool already does the grouping, and it does it well. I wanted it running on my own machine, without a monthly bill. I also wanted the part nobody sells: what a group is worth.
Every group comes out with a price to win and a monthly value. Then a payback figure. How long before the page earns back what it cost to make, and underneath all of that, the question I built this for: is a group worth writing at all when I could just buy those same clicks on ads?
The grouping rule settled early and stopped moving. The money side didn't. It moved nineteen times in a week. Eleven on one day. I'd worked the formula out longhand on paper first, before any of it was code. Paper doesn't tell you when the output looks wrong. I shipped a version, ran a few real cases through it and changed it again the same day.
Some of it I put in and took straight back out. Development cost, for instance. I added columns for it, then dropped them the same day, because the people who'd use this are publishing content and dev cost was noise to them.
Then I bolted on a proper finance model, with a discount rate and a horizon in months. That lasted a day. The next day I pulled the discounting out and put plain payback back.
What survived the week is the ranking, and the ranking is the part I use.
It weights money I'm already spending on ads at double. That's confirmed dollars I'd be pulling back, and the organic figure sitting beside it in the same row is only a guess about a page I haven't written yet, so the two don't get equal billing.
There's a second sort as well. It asks how far a dollar goes. What those clicks are worth on ads, weighted by the odds of reaching the top three, set against what the page costs me to make.
Every figure leans on an estimate. So the tool marks the ones it guessed. It reads intent off the results page where it can. Failing that, the wording. When neither of those works it stamps the term weak, and that label carries all the way into the money math.
When there's no score for how hard a term is to win, the tool works one out from the results and flags that it did. A term it couldn't look up gets written down as a failure. "Nobody searches for this" and "the lookup broke" shouldn't look the same on my screen.
The money numbers were built to feed Competitor Analysis. They're done and they work. Nothing over there reads them yet.
Let me know where I've got the math wrong.