# Coding rubric — sentiment & framing of Swedish 2026 election posters

Version 1.2 · 2026-09-05 (v1.0 variables unchanged; v2 and v3 addenda below)

This rubric turns the qualitative read in `poster-analysis.md` into variables that can be
applied consistently across the corpus and compared across parties. It is designed so that a
second coder — human or model — given the same image and text would produce the same labels.

## Why "sentiment" is defined the way it is here

A campaign poster is, by construction, positive about its own party. Classic polarity
(positive/negative) therefore carries almost no information. The variation that matters is in
**which emotion is solicited**, **what moral language is used**, **who is addressed**, **who
(if anyone) is cast as the antagonist**, and **how much of a causal story is told**. The rubric
codes those, plus the visual cues that carry affect (colour temperature, faces, type scale).

## Unit of analysis

One **unique design**. Alternate resolutions and file formats of the same artwork collapse to one
unit (e.g. S `01-ekonomin.jpg` and `01-ekonomin-1920px.jpg` = one unit). A shared slogan rendered
as visibly different artworks (different person, pose, or layout) counts as **separate units**, but
they carry the same `design_family` tag so family-level and unit-level counts can both be reported.

## Evidence class (carried from the catalogue)

- **A** — clean official asset; text and layout analysed confidently.
- **B** — official design seen in a press photograph; perspective/obstruction limits detail.
- **C** — press-collage only; identification, not fine-grained visual claims.

Every coded row records its evidence class. Findings are reported A-only and all-classes so the
weaker material never silently drives a conclusion.

---

## Variables

### 0. Identification
| Field | Values |
|---|---|
| `id` | short code, e.g. `S-01`, `M-WC-07` |
| `party` | S, V, L, C, M, KD, MP, SD, MED, SjvP, EnadRöst |
| `scope` | national / regional / local |
| `evidence` | A / B / C |
| `design_family` | free text tag shared by slogan-siblings (e.g. `vi-far-saker-gjorda`) |
| `slogan_sv` | verbatim headline (square brackets = uncertain) |
| `slogan_en` | plain gloss |

### 1. Dominant affect (`affect_primary`, optional `affect_secondary`)
The single emotion the design most works to produce in the viewer. Choose from:

| Code | Affect | Recognisable by |
|---|---|---|
| `gravity` | Seriousness / sober authority | formal portrait, muted palette, "på allvar", restraint |
| `hope` | Optimism / forward lift | bright light, "vinner", "framåt", smiling candid |
| `pride` | National / group pride | flag-coded colour, "Sverige", collective "vi" |
| `fear` | Threat / alarm | crime, violence, danger, security nouns |
| `anger` | Indignation / grievance | naming a culprit, injustice, "orättvisa", punishment |
| `reassurance` | Protection / safety restored | "trygg", "skydda", a promise to shield |
| `solidarity` | Recognition / being-on-your-side | everyday scene, "på din sida", listening posture |
| `warmth` | Belonging / community calm | shared places, children + elders, greenery |
| `resolve` | Managerial competence / "we deliver" | "vi får saker gjorda", results ledger, busy leaders |
| `nostalgia` | Loss & restoration of a past | "igen", "återfå", implied golden condition |

Rule: pick `affect_primary` first from text+image gestalt; add `affect_secondary` only if a
clearly distinct second register is present (e.g. SD: `nostalgia` primary, `reassurance` secondary).

### 2. Temporal stance toward the system (`stance`)
Captures the maintenance/restoration/reform/transform axis from the analysis.

| Code | Meaning |
|---|---|
| `restore` | return to a lost/past condition |
| `maintain` | keep present institutions, run them more seriously |
| `reform` | adjust discrete settings incrementally |
| `transform` | change structural ownership/decision arrangements |

### 3. Emotional intensity (`intensity`, 1–5)
1 = calm/measured · 2 = mild · 3 = moderate · 4 = charged · 5 = high-arousal/alarm.
Anchors: S text poster = 1; L pastel policy card = 2; V billionaire-tax joke = 3; M crime card = 4;
SD policy receipt / "utvisas" language = 5.

### 4. Constructed subject — who is addressed (`subject`)
From the five political subjects in the analysis, extended:

`adult-nation` · `squeezed-household` · `autonomous-individual` · `local-community` ·
`bounded-nation` · `crime-victim` · `senior` · `family-parent` · `taxpayer`

### 5. Antagonist / out-group (`antagonist`)
`none` or one of: `criminals` · `immigrants-asylum` · `islamists` · `political-opponents` ·
`profiteers-billionaires` · `waste-bureaucracy` · `developers-highrise`.
Record `antagonist_explicit` = yes/no (named on the poster vs. merely implied).

### 6. Moral foundations (`mf_primary`, `mf_all`)
Haidt's foundations; code the primary and list all salient:
`care` · `fairness` · `loyalty` · `authority` · `sanctity` · `liberty`.
Guide: "trygghet/skydda" → care or authority; "frihet" → liberty; "rättvist/klyftor" → fairness;
"Sverige/nation/igen" → loyalty; "straff/ansvar/kontroll" → authority; "moské/renhet" → sanctity.

### 7. Causal completeness (`causal_score`, 0–4)
Entman frame elements present on the design. +1 each for: names a **problem**; names a **cause**;
assigns **blame/agent**; specifies a **remedy/mechanism**. Slogan-only promise with no diagnosis = 0–1;
V's care-vs-profit or MP's inequality→safety = 2–3; SD policy receipt = 4.

### 8. Visual affect cues
| Field | Values |
|---|---|
| `color_temp` | warm / neutral / cool / dark |
| `subject_of_image` | leader / citizen(s) / place-environment / text-only / object |
| `face_prominence` | none / small / medium / dominant |
| `type_scale` | measured / large / oversized |

### 9. Rhetorical device (`device`, may list two)
`second-person` · `imperative` · `pun-wordplay` · `joke` · `list-receipt` · `slogan-only` ·
`testimonial` · `comparison-vs-opponent`.

---

## Coding procedure

1. Collapse files to unique designs; assign `id` and `design_family`.
2. Read `slogan_sv` from the cleanest available asset (or the catalogue transcription for B/C).
3. Code text-driven fields (affect, stance, subject, antagonist, moral foundations, causal, device).
4. Code visual fields from the image.
5. Record evidence class and any uncertainty.

## Reliability plan (stated, partially executed)

- Full corpus coded once here (single coder = this model, drawing on the Codex catalogue as a
  second independent description where it exists — that overlap is used as a lightweight
  cross-check, not as true double-coding).
- **To make any published number citable**, a ~30-design subsample should be independently
  double-coded (second human or a different model with only the rubric + image), and
  Krippendorff's α reported per variable. α ≥ 0.67 tentative, ≥ 0.80 firm. Categorical variables
  (affect, subject, moral foundation) will be the hardest and should be reported with their α.

## Sampling caveat (must travel with every result)

The archive is **uneven** by construction (see `poster-sources.md`): S/V/L are complete national
sets; C national exists as print masters but the richest C material is a single local Sundbyberg
series; SD national designs are photo-crops (class B); M is over-represented by 29 web cards that
are 16:9 web format, not street posters; KD/MP/M national appear partly only in a press collage
(class C); MED is a single hyper-local Danderyd suite. **Per-party affect distributions are shaped
by what was collectable, not only by campaign emphasis.** Counts are reported per design *and*
flagged by evidence class; cross-party claims lean on class-A material and on within-party pattern,
not raw frequency.


---

# v2 addendum — machinery coding

Added 2026-09-04. Does **not** alter any v1.0 variable or label; it adds two columns to
`coded-data.csv`.

## Why

Every number this project publishes has a codebook and, since the reliability pass, a measured
α — except the one on the front of the website: *no party has a plan for who owns the machines
that will administer Sweden.* That claim was rhetorical. This addendum makes it a coded,
reproducible result, and it borrows the codebook from the other side of the argument.

The codebook is the twelve functional layers of **The Agentic State** vision paper (Ilves,
Kilian, Parazzoli, Peixoto, Velsberg; launched at the Tallinn Digital Summit, 2025-10-09 —
see `costs-sources.md` §9). Using an opponent's framework rather than one written here removes the
obvious objection that the categories were built to produce the finding.

## Variables

### `infra_ref` — 0/1 · **mechanical, headline-bearing**

Does the design refer, in visible text or image, to **any** technical capacity of the state:
software, a model, data systems, compute, networks, surveillance cameras, drones, automation?

This is a presence test on transcribed content — the same family as `antagonist`, which measured
α = 0.82, the most reliable variable in the v1.0 rubric. Published claims should rest on this
variable and no other.

Coding rule: code `1` for any named technical artefact or system, however mundane, including
security hardware. Code `0` for metaphorical or purely institutional language ("bygga Sverige
starkt", "infrastruktur" meaning roads and rail).

### `as_layer` — layers `1`–`12`, `;`-separated, or `none` · **interpretive, illustrative only**

Through which layer(s) of the Agentic State framework would the state have to deliver this
poster's promise? Assigned using the paper's own worked examples — a promise about healthcare
waiting times is coded L1/L2 because the paper's Italy case study is agents managing CT-scan
waiting lists.

**This variable is single-coder and has not been double-coded.** Expect disagreement comparable
to `stance` (α = 0.51), and for the same reason: reasonable coders draw the boundary between
"a political good" and "an administrative delivery mechanism" in different places. Quote the
aggregate shape, never a single design's layer.

Layers: 1 Public Service Design/UX · 2 Government Workflows · 3 Policy- and Rule-Making ·
4 Regulatory Compliance and Supervision · 5 Crisis Response · 6 Public Procurement ·
**7 Agent Governance · 8 Data and Privacy · 9 Tech Stack · 10 Cyber Security and Resilience ·
11 Public Finance · 12 People, Culture and Leadership**. Layers 1–6 are the paper's
*implementation* layers; 7–12 its *enablement* layers.

## Result

- `infra_ref = 1` for **2 of 80** designs: M-WC-06 (*Fördubblat kameramål*) and M-WC-08
  (*Ett svenskt drönarvärn*). Both are security-register hardware.
- **0 of 80** name a model, a datacentre, a cloud contract, compute, or the ownership of any
  of them.
- Layer coverage: **47 design-assignments to implementation layers (1–6), 4 to enablement
  layers (7–12)**. L7 Agent Governance, L9 Tech Stack and L12 People and Culture are empty.

The paper's own rule — *"any deployment of agents within an application layer requires careful
alignment across all enablement layers"* — is what gives that asymmetry its force: the campaigns
promise outputs from layers 1–6 while saying nothing about the layers that would have to carry
them.

## What would make this citable

Same standard as v1.0: a second, blind coder on a ~30-design subsample, with α reported per
variable. `infra_ref` should score very high (it is nearly mechanical); `as_layer` should be
expected to score poorly and be presented accordingly.


---

# v3 addendum — machinery coding for long-form platform documents

Added 2026-09-05. Extends the v2 machinery coding from posters to the eight 2026
*valmanifest* archived in `valmanifest/` (provenance: `valmanifest-sources.md`).
Applied by `build_manifest_dataset.py`; output in `manifest-coded.csv` and
`manifest-aggregations.txt`. Does not alter v1.0 or v2.

## Why the unit of analysis has to change

A poster is one utterance, so one design = one unit. A manifesto is 2 000–27 500 words, so
"one document = one unit" throws away everything and "one proposal = one unit" is neither
well-defined across eight differently-structured documents nor reproducible.

v3 therefore uses **two units at two levels of trust**:

- **The document**, for frequency. Mechanical, reproducible by re-running the script.
- **The passage**, for framing. A passage is a contiguous run of machinery language:
  consecutive lexicon hits less than **250 characters** apart are merged into one. That
  threshold is a rubric parameter, chosen because the passage count is *stable* there —
  45 passages at a 90-character gap, 32 at 250, 31 at 400. It sits on the plateau, not on
  a slope. Corpus total: **39 passages.**

## Variables

### Stage 1 — mechanical

- `hits` — lexicon matches per document.
- `hits_per_1k` — **the comparable figure.** The corpus varies fourteenfold in length
  (C 27 541 words, V 1 955); raw counts are a length measurement before they are a
  political one. Never quote a raw cross-party count.

### Stage 2 — interpretive, single-coder, **not** double-coded

- `frame` — how the passage frames the machinery. Twelve values, grounded in the corpus and
  listed in the script: `competitiveness`, `capacity`, `dependency`, `labour`,
  `admin-efficiency`, `regulation`, `rights`, `risk-social`, `risk-security`, `risk-energy`,
  `inclusion`, `culture`, plus `toc` for a table-of-contents artefact (excluded).
- `as_layer` — the Agentic State's twelve layers, exactly as in v2.
- `own_claim` — **0/1: does the passage assert public ownership or control of models,
  weights, compute or data infrastructure?** This is the variable the whole exercise exists
  to test, and it is the one to state plainly.

Expect `stance`-grade disagreement on `frame` and `as_layer` (v1.0 measured α = 0.51 for
`stance`). Quote the shape, not a single passage's code.

## What the first lexicon missed — read this before trusting any lexicon result

The first pass used a **core** lexicon only: unambiguous machinery words (`AI`, `algoritm`,
`språkmodell`, `datacenter`, `beräkningskraft`, `molntjänst`, `öppen källkod`,
`öppna vikter`, `automatiser*`, `digitaliser*`). It deliberately excluded high-frequency
ambiguous words, and an audit of those exclusions confirmed the exclusions were right:
`data` is mostly waiting-list statistics, `digital` is mostly *digital vård*, `teknik` is
mostly military or educational, `plattform` collides with *valplattform*, `upphandling` is
mostly procurement of welfare services.

**But the audit also found a false negative, and it was the most important passage in the
corpus.** Moderaterna write:

> "Minska Sveriges beroenden av utländsk teknik i kritiska system genom stärkta incitament
> och strategisk upphandling för svensk och europeisk teknik inom **säkerhetskänslig IT,
> moln och drift.**"

The core lexicon missed it because M write `moln`, not `molntjänst`. A second lexicon tier
(`INFRA`) was therefore added for infrastructure- and dependency-language: `moln`,
`halvledare`, `beroende* av utländsk`, `utländsk teknik`, `säkerhetskänslig*`,
`kritiska system`, `digital infrastruktur`, `datadelning`, and the `cyber*` family. That
raised the corpus from 32 to 39 passages and changed the substantive finding.

**The lesson is the general one: a presence test is only as good as its lexicon, and its
lexicon is a judgement call wearing mechanical clothing.** Any future extension of this
rubric must audit its own exclusions before publishing a zero.

## Results

### Stage 1

| party | words | hits | passages | per 1 000 words |
|---|---|---|---|---|
| S | 7 087 | 14 | 9 | **1.98** |
| C | 27 541 | 48 | 16 | 1.74 |
| L | 9 478 | 12 | 7 | 1.27 |
| M | 16 811 | 14 | 6 | 0.83 |
| KD | 4 646 | 1 | 1 | 0.22 |
| **SD** | 3 057 | **0** | **0** | **0.00** |
| **V** | 1 955 | **0** | **0** | **0.00** |
| **MP** | 3 222 | **0** | **0** | **0.00** |
| total | 73 797 | 89 | 39 | 1.21 |

**SD, V and MP contain no machinery language at all** — and that holds under the extended
lexicon, so it is not an artefact of word choice. Note that normalising reverses the raw
ranking: S is the densest, not C.

Absent from all eight manifestos, in ~73 800 words: **`språkmodell`, `öppen källkod`,
`öppna vikter`, `leverantörsberoende`, `digital suveränitet`.**

### Stage 2

Frames: `labour` 7, `risk-social` 7, `competitiveness` 5, `risk-security` 5,
`admin-efficiency` 4, `capacity` 3, `dependency` 2, and one each of `regulation`, `rights`,
`risk-energy`, `inclusion`, `culture`.

Layers: implementation (1–6) **13**, enablement (7–12) **18**, no layer 18. L10 Cyber
Security is the most-touched layer (7), then L9 Tech Stack (5). **L11 Public Finance is
empty.**

**`own_claim` = 0 of 38 coded passages, across eight manifestos and 73 797 words.**

## What this does and does not license

- **Licensed:** "Across the eight 2026 manifestos — 73 797 words — no passage claims public
  ownership or control of models, weights, compute or data infrastructure." Mechanical
  variable plus an explicit, auditable definition.
- **Licensed:** "SD, V and MP contain no machinery language at all."
- **Licensed:** the absent-terms list.
- **Not licensed:** treating this as confirmation of the poster finding. It is a
  *correction* of it — see below.
- **Not licensed:** any single passage's `frame` or `as_layer` as fact.

## How this changes the poster finding

The poster coding (v2) returned **47 implementation-layer assignments against 4 enablement**,
and that asymmetry carried the argument: the campaigns promise outputs from layers 1–6 while
saying nothing about the layers that would have to carry them.

**At manifesto length that asymmetry disappears — 13 implementation against 18 enablement.**
The parties are not ignorant of the enabling layers. M proposes a national cyber shield for
municipalities and strategic procurement to reduce dependence on foreign technology in
security-sensitive IT, cloud and operations. L wants Europe to build its own capacity in
semiconductors, cloud and AI. C proposes a national AI centre gathering research, compute
and industry. These are real infrastructure politics, and the poster corpus could not see
them because a poster has no room for them.

So the honest finding is narrower and harder than "they never mention the machinery":

> They mention it. They propose to buy it, to defend it, to attract it and to be skilled at
> using it. Across 73 797 words, not one passage proposes to **own** it.

The two passages that name technological dependency as a problem — M-04 and L-04 — both sit
in a **defence** chapter, and their instruments are **procurement** and **EU capacity**.
That is the Agentic State pattern exactly (`costs-sources.md` §9): the constitutional
question, filed under upphandling.

## What would make this citable

Same standard as v1.0 and v2: a second, blind coder over the 39 passages, α reported per
variable. `own_claim` should score very high — it has a sharp definition — and `frame`
should be expected to score poorly and be reported as such.
