Every article measured six ways — pageviews, length, references, edits, age and language coverage — then ranked by whatever seemed most interesting, including several things nobody asked for.
Total revisions since the page was created. These are the pages humanity cannot stop fiddling with.
Raw article size in kilobytes of wikitext. Several of these are longer than novellas.
Number of other-language Wikipedias that carry an article on this subject.
Last 7 days of pageviews versus the 7 days before. Something happened.
The steepest attention collapses. Last week's obsession, this week's footnote.
Straight pageviews over the last 7 days.
The largest number of views any of these articles pulled in on one calendar day.
Peak day divided by average day. A tall thin needle of attention and then nothing.
Coefficient of variation of daily views. Emotionally unstable articles.
Lowest day-to-day variation. People read these at exactly the same rate forever.
Weekend traffic as a percentage of weekday traffic. Above 100 = a leisure subject.
The inverse: articles whose traffic dies on Saturday. Homework, jobs, and news.
Old articles that are still read heavily today — age in years weighted by current traffic.
Earliest creation date. Some of these predate the Wikipedia logo.
Most recently created pages that are already pulling serious traffic.
Edits per year averaged over the article's entire lifetime.
How many different human beings have touched this page.
Edits divided by editors. A high ratio means the same people are going back and forth.
Lowest bytes-of-article per lifetime edit. Enormous effort, very little page.
Count of <ref> tags in the wikitext.
References per 1,000 words of prose. Nothing goes unsourced here.
Words of prose per reference. Long articles running on vibes.
Count of citation-needed and dubious-claim maintenance tags.
Share of references that are re-used pointers to an earlier citation.
Weekly views divided by reference count. Enormous audience, thin evidence base.
Estimated readable words after stripping templates, refs and markup.
File and image inclusions in the wikitext.
Total headings at every level. Articles with a table of contents like a phone book.
Wikitable count. Congratulations, you have written a spreadsheet.
Every {{...}} in the source. The machinery behind the page.
Blue-link density: how tangled this page is into the rest of Wikipedia.
Links pointing off Wikipedia entirely.
Number of categories the page files itself under.
Bullet and numbered list items per page. Prose was optional, apparently.
Words per section — the articles that most need a subheading.
Languages per 10 KB of English article. The world cares more than en-wiki does.
Bytes of English article per language edition. Vast here, invisible everywhere else.
Heavily read in English but barely exists in any other language.
Weekly views divided by article size. Short pages doing heroic traffic numbers.
Lowest weekly views per lifetime edit — enormous editorial effort, minimal audience.
Character count of the article title alone.
The terse end of the naming spectrum.
Composite score across references, sections, images and language coverage.
Protected pages ranked by traffic. Too popular, or too contentious, to leave open.
Ten more rankings unlock once this file has a second measurement of the same articles — fastest growing, being cut down, edit storms, attracting new editors, newly translated, getting better sourced, sustained risers and sustained decline.
Each collector run stamps a dated snapshot onto every
article it measures. Run it again in a week:
python3 wiki_stats.py --merge --limit 500 --discover 50
| Article | Views 7d | Δ week | Size KB | Words | Refs | Edits | Editors | Langs | Created | Images | Sections | Score | Data age | Runs |
|---|