Three more data sites I built
I built three more sites that publish counts and nothing else.
The first one, Podcastquery, loads the public index of podcast feeds and reports how many there are.
I assumed the headline number would be the number of podcasts, and that most of them would be at least occasionally alive.
There are 4,711,475 feeds, and 437,693 of them published an episode in the last 90 days.
892,068 feeds have exactly one episode and never got a second.
So the nearly-five-million figure people quote is mostly a count of sign-ups on hosting platforms.
I hit the same problem on ProfessionLens, which holds state licence rolls for doctors, agents, engineers and a few other licensed jobs.
Florida lists 972,132 licensed insurance producers, and 304,219 of them hold no active carrier appointment, so no insurer has appointed them to sell its policies.
I had been reading licence counts as headcounts of people doing the work, and in that one state the licence roll runs about a third above the appointed one.
JournalistLabs breaks in a different way.
It holds 127,602 editorial staff at 8,800 outlets, and every one of those rows is a job title the person wrote themselves.
My count for the New York Times comes to 128% of the 1,700 journalists the paper has published, and my BBC count comes to 72% of its 5,500 figure, which I read as former staff still listing the Times and BBC people never filling in a profile.
The ratio inside a single newsroom still held up, and it was not what I expected: the Times lists 1,144 editors against 632 reporters, while the BBC lists it the other way round.
After three of these I now put the denominator, the source and the pull date under every number on all three sites, and I say which rows I dropped.
That is mostly so I can tell later whether a number moved or whether I just counted differently.
If you spot one that looks wrong, tell me and I will go back to the file.
PS: If you're interested in following my journey, sign up below:
