Hamis ID > Cikkek > Where Our China Name Data Comes From

Ez a cikk még nincs lefordítva erre: Magyar — az eredetit olvasod ezen a nyelven: English. Elérhető még:Deutsch, English, Українська

Where Our China Name Data Comes From

Ask which surname is the most common in China and you will get a confident answer that the data does not support. 王 and 李 differ by 0.06 percentage points — about 800,000 people out of 1.3 billion, well inside the noise of a household registration snapshot. "The most common surname in China" is a question without a meaningful answer, and the first thing an honest dataset has to do is stop pretending otherwise.

This page documents the source behind our Chinese surname weights, what we could and could not measure, and the large hole in the middle of the published record that we deliberately did not fill with invention.

The registry we used

FieldValue
Source全国户籍人口姓氏统计 — national household-registration surname analysis
Publisher公安部治安管理局 — Public Security Administration Bureau, Ministry of Public Security
Reference dateApril 2007
Depth publishedTop 100 surnames with population shares
Below rank 100Nothing. No published table exists.
BasisHousehold registration (户籍), not census sampling

Published anchors we rely on:

  • 王 — 92,881,000 people, 7.25%
  • 李 — 92,074,000 people, 7.19%
  • 张 — 87,502,000 people, 6.83%
  • Top 10: 王, 李, 张, 刘, 陈, 杨, 黄, 赵, 吴, 周 — each over 20 million; 42.9% combined
  • Top 100: 84.77% of the population

What we measured vs what we estimated

Measured. Ranks 1–100, and by extension the shape of the head. Weights are a linear rescaling of the published shares against 王:

weight = round(255 × share / 7.25)

There is an unusual detail here worth being open about. When we audited the file's existing head against the 2007 table, we expected to find fitted values and found registry values instead: 23 of the top entries deviated by under 2% (only 黄 was off by 4%). The plateau at the top — 王 255, 李 252, 张 240 — was not constructed to look flat. It fell out of the real shares. The 2026 recalibration rescaled ranks 1–157; it did not re-derive them.

Estimated — and we want to be blunt about it. Ranks 158–460 are all sitting on weight 1, and that is wrong by a factor of about 50. See below.

Pitfalls specific to China

The top 5 is a plateau, not a ranking

王 7.25%, 李 7.19%, 张 6.83%. The gap between first and second is 0.06 pp. Household registration is not a census; it lags migration, deregistration and birth reporting by margins comfortably larger than 0.06 pp.

So: 王, 李 and 张 form a plateau. Any dataset, article or infographic that presents "#1 王" as a fact is over-reading the source. Our weights preserve the plateau because the plateau is what the measurement says. Adjacent ranks being nearly equal is not an error to be smoothed out — it is the finding.

The same caution applies further down. 赵, 吴 and 周 all land on weight 71 in our file, which means our data cannot order them. The registry can, barely — 赵 2.06%, 周 2.02%, 吴 2.00%. On an integer scale those are one number.

公安部 and 袁义达 count differently — and neither is wrong

This is the trap that produces two contradictory "official" Chinese surname rankings.

The 2007 MPS analysis counts 萧 and 肖 separately, likewise 戴/代 and 傅/付. The demographer 袁义达's widely cited rankings merge them, on the grounds that the split is an artefact of 20th-century simplification and clerical substitution rather than a real division of lineages.

Both are defensible. They produce different rank orders, and lists derived from them cannot be spliced together. Ours follows the MPS convention because our weights come from the MPS table, and mixing conventions inside one dataset is how you end up with a surname counted once and a half.

Readers of our Taiwan page will recognise the shape of this: Taiwan's registry counts 温 and 溫 separately because it records written forms. Same underlying phenomenon, different institutional answer, and in both cases the source's convention wins over our sense of what "should" be one surname.

There is no published table for ranks 100–300 — so we did not invent one

Here is the honest hole in the middle of our Chinese data.

The MPS publishes the top 100 (84.77%). 袁义达's 2006 work publishes a slice — "129 surnames at ≥0.1% account for 87%" — which is an aggregate, not a per-rank table. The only systematic source for ranks 100–300 is the book 《中国姓氏·三百大姓》 (袁义达 / 邱家儒, 2007), which we did not have.

So ranks 158–460 in our corpus all carry weight 1. That is factually wrong and we know exactly how wrong. Sitting on that shared floor are 欧阳 with roughly a million bearers and 壤驷 / 漆雕 / 澹台 with a few thousand each. The real ratio is on the order of 50:1; our data says 1:1. Worse, those entries are ordered by their sequence in the 百家姓 primer, not by frequency — that section of the file is a list, not a curve.

We could have interpolated a plausible-looking power curve across that range. It would have looked better and been fabricated. The gap is real, it is documented, and the fix is the book, not a formula.

Coverage vs the number in the file

Our corpus is 460 surnames, covering roughly 91% of the population. That immediately explains a number people try to "correct": the top-10 share inside our file is 42.3%, not 42.9%. Different denominators — 42.9% is a share of all Chinese people, 42.3% is a share of a file that does not contain every Chinese surname.

The honest in-file target is 42.9% ÷ 91% ≈ 47%, and we do not reach it. 42.3% is the ceiling, not a shortfall: the top 10 (1,496 weight units) and ranks 11–157 (1,742 units) are pinned to real registry shares, and the 303 tail entries cannot weigh less than 1 (=303 units). Any number above 42.3% requires understating ranks 11–157 — i.e. lying about them to make an aggregate look right.

The check that caught us

One diagnostic is worth publishing because it is portable. If a weight of 1 corresponds to share(top1) / 255 of the population, then summing all weights and multiplying tells you how much population the file claims to describe:

claimed coverage = sum(weights) × share(top1) / max(weight)

Before the 2026 recalibration our Chinese file claimed 114% of China. That is arithmetically impossible, and it proved the tail was over-weighted without needing any judgement call about which entries looked wrong. After recalibration it claims 100.7%.

Top 10

RankSurnameWeightBearersShare of population
125592,881,0007.25%
225292,074,0007.19%
324087,502,0006.83%
4189> 20,000,000~5.37%
5158> 20,000,000~4.49%
6107> 20,000,000~3.04%
782> 20,000,000~2.33%
871> 20,000,0002.06%
971> 20,000,0002.00%
1071> 20,000,0002.02%

What is measured and what is arithmetic. Ranks 1–3 are published counts and shares. Ranks 4–7 show shares reconstructed from the stored weight — the file keeps an integer weight, not the original share, so these are recoverable only to within rounding. Ranks 8–10 show the published shares, which is why they are not in weight order: the file cannot separate 2.00% from 2.06%. The MPS states only that each of the top 10 exceeds 20 million bearers; per-surname counts below rank 3 are not in the published summary we used.

Note the reconstruction runs slightly low: our weights round-trip to a top-10 of ~42.5% against the published 42.9%, and a top-100 of 84.0% against the published 84.77%. That ~0.8 pp shortfall is integer rounding accumulating across 100 entries, in a consistent direction.

Known limitations

  • Ranks 158–460 are flat and unordered. The most serious defect in this locale. ~50:1 of real variation compressed to 1:1, sequenced by a classical primer rather than by frequency. Fixable only by acquiring 《中国姓氏·三百大姓》.
  • The data is from 2007. Nineteen years old. The MPS has published newer reports (《二〇二〇年全国姓名报告》, 《二〇二一年全国姓名报告》) but they do not carry an equivalent full top-100 share table, so we did not switch. For scale: the 2020 report gives the top 5 as 30.8% of registered population; our 2007-calibrated head implies 31.1%. That is a temporal stability signal from the same agency — evidence the distribution has barely moved — and explicitly not independent confirmation that our numbers are right.
  • The top 10 cannot be ordered by our data below rank 3. Nor should it be, really. See the plateau section.
  • Coverage is ~91%. The remaining ~9% of the population holds surnames absent from our file.
  • Household registration is not a census. 户籍 lags actual residence. The direction of the error is known; the size is not.
  • Simplified/variant merging follows the MPS convention. If you need 袁义达's convention, ours is the wrong file.
  • The top-100 consistency check is not verification. 83.4% in-file × 100.7% claimed coverage = 84.0% against the published 84.77% is a self-consistency check against the same source the weights came from. It shows the curve is arithmetically coherent past the head. It cannot show it is true.

Data as of 2026-07-17

Registry reference date: April 2007. Dataset recalibrated: 17 July 2026 (top-10-in-file 37.3% → 42.3%; claimed coverage 114% → 100.7%). Corpus: 460 surnames covering ~91% of the population. Male and female files are byte-identical — Chinese surnames do not inflect for gender.

Sources

← Cikkek