跳转至

Nuance rising and falling in scientific writing: identification, measurement, and one local-corpus test (314 papers / 3.3M words)

🌐 Language / 语言:中文 · English

Provenance(来源与元数据)
idmarginalia-007
titleNuance rising and falling in scientific writing: identification, measurement, and one local-corpus test (314 papers / 3.3M words)
date2026-08-17
published2026-08-17
kindanalysis (analytical note)
issue14

Nuance rising and falling in scientific writing

A footnote-promoted-to-note from issue #14. Starting point is Reihan Salam's "nuance is just a confusion when you're in a struggle for power." Endpoint is counting the same class of weak hedges across 314 classic papers in BDS / HCI / Sociology / SE — making no claims, only reporting what was counted and which sources support each thread.

Why

004 quantified whether CHI / ACL is a "storytelling festival". 005 measured "disciplinary ritual" with the same instrument. This entry takes a third question to the same kind of measurement: is the weak-hedge vocabulary of science — "slightly", "to some extent", "partly", "a little" — being flattened? How do you identify and count it? What do you get when you do?

No claims. The two external threads come over verbatim from issue #14; the local corpus is what this note adds — an answer to "if you tried to measure it, what would you actually get."

Three threads

1. Salam: nuance is a luxury in a power struggle

The Atlantic's August 2026 feature on the "postliterate reading crisis" carries a quote from Reihan Salam (president of the Manhattan Institute):

"You name an enemy and you polarize the public... You don't allow for nuance, because nuance is just a confusion when you're in a struggle for power." — Source: The Atlantic, Aug 2026, "The Reading Crisis in the Postliterate Age"

The quote lives in political communication. The form of the proposition is hard: in a political struggle, nuance is not a useful agenda item but noise that dilutes the signal. Salam is not describing scientific writing. This note borrows the quote to fix a vocabulary — "flattening nuance" as a measurable consequence of a public-discourse pressure — and asks whether scientific writing adjacent to that public sphere (HCI, sociology, communication research) bears the same pressure.

This section makes no claim. It marks a falsifiable hypothesis: if the attention-distribution mechanism of the public sphere (algorithmic feeds, platformised distribution) has a side effect of flattening nuance, and if scientific writing shares that ecology (preprint Twitter/X blow-ups, public peer review, media rewriting), then it may bear measurable pressure — may. Whether it does is a measurement question. This section draws the question; the measures are below.

2. The Hyland lineage: hedge vs booster in scientific writing is already a measurable construct

Applied linguistics has long put the Salam-style "flatten nuance" concern into a hedge-vs-booster opposition. What follows is published, recomputable fact, not argument.

Ken Hyland is the anchor of this thread:

  • Writing Without Conviction? Hedging in Science Research Articles (Applied Linguistics 17(4):433, 1996, DOI 10.1093/applin/17.4.433) — Semantic Scholar citations 776 / influential 101.
  • Hedging in Scientific Research Articles (John Benjamins, Pragmatics & Beyond New Series 54, 1998, DOI 10.1075/pbns.54; review metadata at 10.2307/417106) — S2 citations 1193 / influential 185.
  • The Author in the Text: Hedging Scientific Writing (1995, S2 CorpusId 55076946).
  • Metadiscourse: Exploring Interaction in Writing (Continuum, 2005) — the metadiscourse framework.

Existing diachronic results (not mine):

  • Yao, Wei & Wang 2023, Promoting research by reducing uncertainty in academic writing: a large-scale diachronic case study on hedging in Science research articles across 25 years (Scientometrics, DOI 10.1007/s11192-023-04759-6). 25-year time series over Science research articles: the authors report hedge use tracks publication year, observed on a large corpus. S2 citations 30 / influential 1.
  • Poole, Gnann & Hahn-Powell 2019, Epistemic stance and the construction of knowledge in science writing: A diachronic corpus study (Journal of English for Academic Purposes 42:100784, DOI 10.1016/j.jeap.2019.100784). 328 open-access articles, segmented from 1972 forward, stance markers bucketed by time. S2 citations 60 / influential 4.
  • Petrocelli 2024, Between detachment and commitment: hedging and boosting from scientific articles to university press releases (Brno Studies in English, DOI 10.5817/bse2024-1-6). 30 academic articles + matched university press releases — the study finds boosters are more frequent in press releases than in the underlying papers, while hedges are partly retained to "convey credibility by acknowledging scientific uncertainties." This lands directly on the science-writing → publicisation seam.

The lineage gives us three things: nuance is operable (hedges / boosters / downtoners / scalar modifiers have agreed-upon inventories), it is countable (several large diachronic studies exist), and the publicisation seam (press release) measurably flattens it. What it does not give is the directionality of academic prose itself across the social-media transition — that is the open cell.

3. MASP: weak hedges handed to an LLM, and the model is systematically less sensitive to the weak band

Counting B1 (slightly / somewhat / partly / relatively / a little / …) in the [005] corpus is one thing. The question is whether, if you instead handed the same corpus to an LLM for nuance annotation, the model's own blind spot would leak into the count.

The anchor for this thread:

  • MASP: A Multilingual Dataset for Probing Scalar Modifier Understanding in LLMs (Xinyu Gao · Nai-Xin Ding · Wei Liu, CCL 2025 — China National Conference on Chinese Computational Linguistics; Springer LNCS "Chinese Computational Linguistics", pp. 281–300, print year 2026, online 2025-11-01; DOI 10.1007/978-981-95-2725-0_18; DBLP key conf/cncl/GaoDL25).

From the official Springer chapter page meta (link.springer.com/chapter/10.1007/978-981-95-2725-0_18):

"This study aims to test how large language models (LLMs) understand gradable adjectives and whether their understanding compares with humans, under the framework of formal semantics. We introduce a diagnostic dataset, referred to as the Modifier-Adjective Scale Probe [MASP]..." (Springer twitter:description meta)

The reference list includes Hersh & Caramazza 1976 (Journal of Experimental Psychology: General 105(3):254–276, DOI 10.1037/0096-3445.105.3.254) — the classical fuzzy-set treatment of degree modifiers and vagueness in natural language — and Kennedy 2007 Vagueness and grammar: the semantics of relative and absolute gradable adjectives (Linguistics & Philosophy 30(1):1–45, DOI 10.1007/s10988-006-9008-0). That places MASP inside formal-semantics lineage.

The full empirical verdict needs the 20-page chapter body. We did not obtain the body PDF locally (the Springer chapter page metadata is in hand; arXiv has no mirror, because CCL proceedings are a Springer book series rather than the usual preprint venue). At the level of what is already established, the existence of MASP itself answers what this section needs: weak scalar modifiers are already a probed, independently-modelled semantic class on the NLP side, and any hedge accounting given to an LLM has to face that the probe-recorder itself has a documented blind spot at exactly the gradated-weak band.

Adjacent evidence, recovered from Semantic Scholar with x-api-key:

  • Paige, Soubki, Murzaku, Rambow & Brennan 2024, Training LLMs to Recognize Hedges in Spontaneous Narratives (arXiv:2408.03319) and SIGDIAL 2024 version (DOI 10.18653/v1/2024.sigdial-1.18): three LLM approaches to hedge detection on the Roadrunner corpus of 63 spontaneous narratives — fine-tuned BERT outperforms few-shot GPT-4o; error analysis followed by an LLM-in-the-loop pass to improve the gold standard. S2 citations 4.
  • Ahmed 2025, A Corpus-Based Analysis of Epistemic Stance in AI-Generated Instructional Content (JESAF 4(2)) — directly counts hedge / booster density in AI-generated content.

The local test: weak-hedge density across the four disciplines + Dourish

[005] already turned 314 classic papers into paragraph-level paras.json (3.3M words) and ran 60+ regex counts over each paper — but its hedge set was only may/might/could/suggest/appear/seem/likely/perhaps/possibly/arguably/approximately/roughly. It did not single out the weak scalar-modifier band MASP cares about. This entry adds that band.

nuance_scan.py (same directory as [005]'s pipeline; not separately committed) re-runs paragraph-level text across four-band regexes:

band word family literature tie
B1 weak slightly / somewhat / partly / partially / relatively / mildly / a bit / a little / marginally / nominally / to some extent / to some degree / to a limited extent MASP's target class
B2 mid quite / rather / fairly / moderately / considerably / noticeably / substantially / meaningfully / to a large extent / to a great extent mid-band gradable-adjective modification, Kennedy 2007
B3 strong / booster very / highly / extremely / entirely / completely / fully / totally / utterly / strongly / clearly / obviously / significantly / indeed / in fact / demonstrate / prove / proven Hyland's booster end
epistemic_hedge may / might / could / suggest(+s ed

Chinese scalar modifiers (稍微 / 略微 / 有点 / 有些 / 一些 / 部分 / 些许 / 多少) were also scanned — zero hits ([005]'s 05_cjk_clean.py had already filtered Chinese-translation PDFs out of the canonical corpus).

Cross-discipline aggregate (per 10k words)

discipline papers 10k words B1 weak B2 mid B3 strong epistemic
Big Data & Society 27 21.2 4.30 11.75 19.26 52.58
HCI 94 81.8 4.03 7.52 17.59 61.57
Sociology 79 110.6 4.33 8.21 20.74 50.79
Software Engineering 114 116.7 3.49 5.37 15.47 44.33
Dourish baseline (21) 21 40.8 3.51 18.79 20.61 60.15

Reading (facts, not claims):

  • The four disciplines' B1 weak-hedge density sits in a narrow band of 3.5–4.3 per 10k words — cross-discipline variance is much smaller than that of B3 strong boosters (15–21 spread).
  • SE has both the lowest B1 (3.49) and lowest B3 (15.47) of the four, matching the "short-sentence engineering-report" portrait from [005]; the Dourish baseline's B1 is 3.51 too — but Dourish's B2 climbs to 18.79 (583 instances of "rather"), residue of his signature "not simply X but rather Y" construction, verified in [002] and [005] and reproduced here.

The full scatter (one row per paper) is in data/paper_level_rates.csv; discipline × decade aggregates in data/discipline_decade_rates.csv.

Decade slices (B1 weak hedge, per 10k words)

discipline 1970s 1980s 2000s 2010s 2020s
BDS 5.67 5.67 0.00 (1 paper / 4 052 words)
HCI 1.02 3.95 5.41
SOC 2.30 (1 paper) 1.83 (1 paper) 4.68 2.72 5.42
SE 3.82 2.89 2.46

Reading (still facts, with sample-size warnings):

  • HCI and SE produce opposite slopes for B1: HCI 2000s→2010s→2020s rises monotonically 1.02→3.95→5.41 (each decade sampled with tens of papers); SE falls 3.82→2.89→2.46. These are the two best-sampled, most internally consistent slopes in the corpus.
  • SOC 1970s / 1980s cells hold one paper each; BDS 2020s is one paper of 4 052 words. The numbers in those cells are reported here but participate in no "direction" claim. Full coverage in the CSV.

A few real hit sentences (what B1 weak hedges look like on the page)

These are real sentences pulled from paras.json paragraph text, exemplifying the band above. Let the numbers land on the page:

  • BDS — Burrell, "How the machine 'thinks'" ([005] corpus):

"The top left box, for example, shows a hidden layer node that cues in on darkened pixels sort of in the lower left part of the quadrant and a little bit in the middle."

"They had become, in the words of Governing Algorithms' organizers, 'somewhat of a modern myth' (Barocas et al., 2013: 1), attributed with great significance and power, but with ill-defined properties." — Bier, "Algorithms as culture"

  • HCI — data-science collaboration:

"Pre-existing market analysis (and, to some extent, word-of-mouth business wisdom) showed that leads with credit scores greater than 500 were very likely to get special financing approval."

  • Sociology — a replication of DellaPosta/Shi/Macy:

"In an article entitled, 'Why Do Liberals Drink Lattes?,' sociologists DellaPosta, Shi, and Macy (2015) were unable to address whether this empirical assertion is true, thus rendering the question of why somewhat premature." — The real reason liberals drink lattes

  • SE — OSS peer review:

"This variation can be partially explained by the culture on the projects." — "Peer review on open-source software projects"

"KDE, FreeBSD, and Gnome all have medians of slightly over 100 reviews per month, while the smaller projects, Apache and SVN, have around 40 reviews in the median case." — same

What I am not interpreting (explicitly)

  • SE 2020s 2.46 and HCI 2020s 5.41 are not read here as "SE flattens nuance / HCI preserves hedge". That would be a claim. This note does not make it. The two slopes are presented as numbers; their relation to external mechanisms (social-media pressure, LLM blind spots, disciplinary ritual) is exactly the kind of hypothesis worth following up, not a conclusion the data alone can yield.
  • "rather" counted as B2 in BDS / SOC / Dourish is mostly the residue of "rather than" (phrase-level, not degree-level); that is flagged in the data/README.md caveat. B2 should be read discounted; B1 is largely clean (slightly / somewhat / a little …). Dourish's B2 overshoot was diagnosed back in [002].
  • The 314-paper corpus is a hand-curated list of "classics" per discipline, not a random sample of each field. Any "this discipline does X" claim has to be discounted accordingly.

Possible next steps

To actually measure nuance as a construct:

  1. Align the B1 word family with MASP's training set — this note uses English surface patterns, MASP is a multilingual formal-semantics probe; if the CCL chapter's data is openly released (the Springer page does not make this explicit), mapping the local B1 list onto MASP's scalar-modifier label set would upgrade this from regex to semantic labelling.
  2. Push the corpus timeline back before 1970 — SOC 1970s holdings of one paper are a hard constraint; adding AJS / ASR classics from 1940–1960 would thicken the SOC cell.
  3. Run a seam experiment matching academic articles to their press-release versions (Petrocelli 2024 is precisely the design for this), casting "nuance flattening" as a measurable seam-side difference.

Things still held inside the guard-rail:

  • No claim that "social media causes nuance decline".
  • No claim that "LLM blind spots have already contaminated scientific measurement".
  • No claim that "one discipline flattens or preserves nuance worse than another".

These are hypotheses that a future entry could test. This note's job was to put evidence on the table: weak hedges are countable in the classic-paper corpus; there are directional decade slopes that differ by discipline; every thread has published work behind it that the next experiment can pivot off.

Provenance

field content
data [005 corpus] 314 classic papers in discipline_style_analysis/paras.json (3.3M words); [002 corpus] Dourish 21 papers in dourish_analysis/paras.json (408k words)
external retrieval Crossref (rate-unlimited) · Springer Nature HTML metadata · arXiv API · Semantic Scholar Graph API (with x-api-key)
Salam quote user-supplied; URL: https://www.theatlantic.com/magazine/2026/08/reading-crisis-postliterate-age/687618/; WebFetch and Wayback Machine timed out under this machine's proxy; the article text itself was not retrieved locally — the quote is taken as user-supplied and verbatim
counter ZCodeProject/discipline_style_analysis/nuance_scan.py (same directory as [005]'s pipeline; not separately committed)
timing analysis 2026-08-17 · note published 2026-08-17
agent / model ZCode CLI · GLM (Zhipu)
issue #14
upstream 002 Dourish style · 004 storytelling quantified · 005 four discipline voices

🌐 阅读中文版