Auditing Your Stack List Against Real Postings

The tech stack list problem gives you a filter based on what you can defend under questioning. That filter is necessary, and it is only half the job: it tells you what to keep of what you already wrote. It says nothing about whether the surviving list resembles what the roles you want are asking for.

Most engineers’ skills sections are a record of employment history. Whatever the last three employers happened to run is on the page, in roughly the order it entered your life. That is a perfectly honest document and it is not the same document as the market you are applying into.

The gap between the two is measurable. If you can write a script you can measure it in an evening, and if you can’t, you can measure it with twenty-five browser tabs and a pen.

Decide what you are measuring first

Not “in-demand technologies.” Not an industry report. The frequency of terms in postings you would actually apply to.

Define that population before you collect anything:

  • Title band — “backend engineer” and “platform engineer” are different markets with different vocabularies, and averaging them produces a list that describes nobody.
  • Seniority — a corpus that mixes graduate and staff postings measures the mean of two distributions.
  • Region, or remote — stack fashions are regional. So are the terms used for the same thing.
  • Company size or sector, if you have a preference. A fintech corpus and a developer-tooling corpus disagree loudly.

The narrower the definition, the more useful the result and the fewer postings you will find. That trade is the whole design of the experiment — decide it deliberately rather than letting a search box decide it for you.

Collecting the corpus

You need the full text of each posting, its employer, and its date. Save the text, not a link — you will want to re-run the extraction after you fix your term list, and half the links will be dead by then.

The mechanical part is dull and worth automating if you can. Job-search APIs hand back postings as records rather than as pages: Serply, for instance, publishes a documented Google Jobs endpoint returning postings with the employer, listed perks, and remote flags as fields. That turns the corpus into JSON you can iterate over instead of HTML you have to read, which matters mostly because it makes the count repeatable — you can change one filter and run it again.

If you are doing it by hand, twenty-five postings read properly beats two hundred skimmed. How many is enough? A defensible stopping rule: keep adding postings until another ten don’t change the top of your ranking. If the order is still shuffling, you are looking at noise.

Extracting terms, and the trap in it

The obvious approach is to check each posting for the terms already on your resume. Do that — it is the cheapest half of the answer — but notice what it can’t do. A vocabulary made of your own terms cannot discover the term you are missing. That is the tautology at the centre of this exercise and it defeats most people’s first attempt.

So run two passes. The counting pass uses a fixed vocabulary. The discovery pass is manual: read ten postings end to end and write down every proper noun, product name, and capitalised term you did not have on your list. Add them to the vocabulary and count again.

Normalise before you count

This is where a naive script quietly lies to you.

  • Aliases. Kubernetes and K8s. PostgreSQL and Postgres. GCP and Google Cloud Platform. Node and Node.js. Collapse them, or you will get two half-sized counts and conclude that neither term matters.
  • Word boundaries. Substring matching on Go hits “Django”, “MongoDB”, “going”. Matching R hits everything. Anchor on word boundaries and keep case sensitivity for short tokens.
  • Umbrella phrases. “CI/CD”, “cloud”, and “microservices” appear constantly and name a category rather than a technology. Count them if you like, but don’t act on them.
  • The same vacancy listed more than once. Syndicated duplicates inflate whatever the duplicated posting contains. De-duplicate on employer plus title before counting.

Reading the output

Cross the frequency with what is currently on your resume. Four cases, and only two of them require work.

Frequent in the corpus, on your resume, buried at the end of a line. Promote it. This is the cheapest change available and it is usually several terms, not one.

Frequent in the corpus, absent from your resume. Two very different situations that look identical from the outside. Either you have the experience and never wrote it down — extremely common, because the thing had a local name in your head and you called it “the deploy stuff” for three years — or you genuinely don’t have it. Sort these before you do anything else; the first kind is an hour of editing and the second is a project.

Rare in the corpus, on your resume, and it is your best work. Keep it. A term appearing in a handful of postings is not worthless — it is how you end up in the specific role rather than the generic one, which is the argument in a specialism the reader wasn’t looking for. Rarity is only a reason to cut when the thing is also not yours.

Rare in the corpus, on your resume, and it is filler. Gone. You already knew.

There is a fifth case worth naming: terms that appear in nearly every posting. Version control, a cloud, “Agile”. A term at very high frequency carries about as much discriminating information as a term at zero, because everyone applying has it on the page. The useful signal lives in the middle of the distribution.

What the count does not license

It does not authorise a claim. A frequency table tells you that a term is expected; it does not put the experience in your hands. Adding a term because it scored well, with nothing behind it, fails the same interview question as before — how did you use this? — and now it fails it in the first five minutes. The count changes your ordering and your emphasis, and it tells you what would be worth going and learning. It never writes a line for you.

It does not measure depth. “Kubernetes appears in most of your target postings” and “these teams want someone who has operated a cluster” are different findings, and the count only produces the first. What each posting means by the term is a separate reading problem — what an engineering job posting tells you about the team covers pulling that out of the prose.

It is not about keyword-matching software. Whether a term is machine-extractable from your file is a different question with a different answer. This exercise is about whether you are describing the right work to a human being at all.

And it goes stale. You measured a region and a title band during a particular few weeks. Re-run it when your search stops working, not continuously.

The part that actually changes

What you get out of this is not a list of technologies to bolt on. It is the end of arguing from memory about a population you never looked at. Most disagreements engineers have with their own skills sections — is this too long, should I drop the second cloud, does anyone still care about this — are arguments about frequencies nobody has counted.

Counting them takes an evening, and afterwards the reordering is obvious. Then go back to naming the problem rather than the tool in the bullets underneath, because a well-ordered list of nouns is still only a list of nouns.