> **Public edition.** Generated from the study's internal protocol by `qa/public-edition.mjs`. The names of websites the paper anonymises, and the internal Agency-NN and Unnamed-NN numbers, are replaced by [agency] or [unnamed website]; nothing else is changed.

> **Public edition note.** This is the protocol as frozen on 26 September 2026. The deviation log at the end records every later change. As run: 576 API trials and 48 Google captures over two windows (26 to 27 September and 4 October 2026). The paper and its downloads are at https://studio.aliqgroup.com/en/research/ai-answer-citations-malaysia-2026/.

# Pre-registered study protocol (framework §3; playbook 06 §6.6)

**Status:** DRAFT v0.1 for Jinn's approval, 2026-09-26. **Frozen** when approved: from then on prompts are added, never rewritten; any deviation is logged in §14 with date and reason and reported in the paper.
**Companion files:** `panel/prompts.json` (the panel, `frozen: false` until approval), `panel/run.mjs` (the instrument, built after approval), `analysis/codebook.md` (derived from §7 at coding time).

## 1. Research questions → analyses

| RQ | Question (from `brief.md`) | Pre-registered analysis |
|---|---|---|
| RQ0 | What sources does the answer rest on? | A1 source-type share of citations, per engine, per window, pooled |
| RQa | How often does the engine search before answering; does it differ by intent? | A5 search rate = trials with a recorded web search ÷ completed trials, by engine × intent |
| RQb | Provider-owned vs third-party share | A2 share of citations coded `agency-owned-service`, `agency-owned-listicle` or `freelancer-site` vs all other codes |
| RQc | Which third-party platforms recur | A4 top named platforms by citation count and by number of distinct prompts on which they appear (agencies anonymised) |
| RQd | Malaysian or foreign | A3 publisher-country share (MY / foreign / unknown) |
| RQe | Stability a week later | A6 W1→W2 overlap: Jaccard of cited domain sets per prompt × engine; share of W1 domains that reappear in W2 |
| RQf | Homepage or category page | A7 among agency-owned citations, page-type share (homepage / service or industry page / article or listicle / other) |
| extra | Concentration | A8 number of distinct domains that account for 50% and 80% of all citations, per engine |

Exploratory (reported separately, labelled exploratory): the Bahasa Malaysia arm; the Google AI Overview / AI Mode sub-panel; source type × intent; the GEO category compared with the other six.

**Working hypotheses (may fail):** H1 third-party pages > provider-owned pages in citation share on both engines. H2 the two engines cite different source-type mixes. H3 fewer than half of W1 cited domains reappear in W2 for the same prompt. The paper reports the measured shares whichever way they fall.

**Statistics:** descriptive shares with their denominators; Jaccard overlap; simple agreement and Cohen's kappa for coding reliability. No significance tests and no inferential claims: n is small and the prompts are designed, not sampled. Language: "in this audit", "possible contributing factors", never "causes".

## 2. Panel

- 42 English prompts: 7 categories (website, SEO, social media management, content production, paid advertising, branding, AI-answer visibility) × 3 intents (provider selection, informational or evaluation, cost) × 2 phrasings. The two phrasings are a paraphrase pair (playbook §6.6 item 3), one generic-Malaysia, one local with a business type.
- 6 Bahasa Malaysia prompts, exploratory (open item 4), wording checked by a native reader before freezing.
- Evidence tag per prompt: `E-GSC` (a real query shape from Search Console, 90 days to 2026-08-18), `E-CALL` (a question prospects asked on recorded calls, April to September 2026, paraphrased), `H` (editorial hypothesis). Draft count: 8 E-GSC, 16 E-CALL, 24 H.
- Rules: no brand names of any kind; no wording that leads toward any provider; no prompt reuses transcript wording; local prompts name Kuala Lumpur, Selangor or Petaling Jaya only.
- The GEO category is included like any other category (open item 3). D5 applies to **reporting**: Aliq's own citations across the whole panel are disclosed as one aggregate line and never per prompt or per category.

## 3. Surfaces, settings, models (pinned)

| Lane | Access | Model (pinned) | Settings | State recorded per row |
|---|---|---|---|---|
| Gemini | Generative Language API `generateContent` | `gemini-3.8-flash` (smoke-tested 2026-09-26; `modelVersion` recorded per row) | tool `google_search`; API default sampling; no location control exists, so local prompts are tagged "location: none" | model, modelVersion, `webSearchQueries`, grounding chunk count, resolved URLs with status |
| ChatGPT | OpenAI Responses API | `gpt-5.5-2026-04-23` (dated snapshot; smoke-tested 2026-09-26 with search on, KL location honoured) | tool `web_search` with `user_location` = Kuala Lumpur, Kuala Lumpur, MY (approximate); default sampling | model, whether a `web_search_call` occurred, `url_citation` annotations (raw URL kept, `utm_source=openai` stripped in a normalised field) |
| Google AI Overview + AI Mode | real Chrome through the connected extension | n/a (consumer surface) | `google.com/search?q=…&gl=my&hl=en` and `…&udm=50`; 12-prompt sub-panel (the `-P-a` and `-I-a` prompts of the six non-GEO categories); 1 draw per window; state = Jinn's signed-in Chrome on Windows, personalised, recorded as such | page text, whether an AI block rendered, links inside the AI block, wall or absence recorded as a row |
| Not run | Perplexity, Claude, Microsoft Copilot, Meta AI, Chinese assistants | | no API key or property; stated in the paper as not run, never as absent | |

Both API lanes are the vendors' retrieval stacks under a clean, un-personalised API state. They are repeatable but **not byte-identical to the consumer apps**; the paper says so in its method and limitations.

## 4. Draws, windows, schedule

- 3 draws per prompt per engine per window. Spacing ≥1.5 s (OpenAI) and ≥4 s (Gemini). Order: prompts in file order, draws consecutive.
- Two windows, **W1** and **W2**, at least 7 and at most 10 days apart, each completed within 48 hours, Asia/Kuala_Lumpur. Same instrument (sha256 recorded), same prompts file (sha256 recorded), same models.
- Trials: (42 + 6) × 2 engines × 3 draws × 2 windows = **576 API trials**, plus 24 Google captures. Cost estimate under USD 20 (OpenAI ≈ 9k tokens per trial at the smoke-test rate; Gemini flash negligible). Abort threshold: USD 30 spent, checked after W1.
- Pre-flight before each window: both keys present, a one-call balance check on OpenAI, disk space, the output file for that window does not exist (the instrument refuses to overwrite, lesson G1 from the 2026-09-21 audit).

## 5. Trial and citation definitions (framework §15 counting rules)

- **Trial** = one prompt × one engine × one draw × one window. States: `completed`, `error` (HTTP or API error, recorded with the message), `no-answer`. Errors and no-answers are kept and reported, never discarded or silently retried; a retry is a new row flagged `retryOf`.
- **Citation** = one URL the engine returned in its **structured citation field**: OpenAI `url_citation` annotations; Gemini `groundingChunks[].web.uri` resolved to the final URL. URLs that appear only in the answer text are recorded separately as `textUrls` and are not citations for A1–A8.
- **Unit for share metrics** = one (trial, cited URL) pair; **unit for domain metrics** (A4, A6, A8) = one (trial, registrable domain) pair, deduplicated within the trial. Denominators are printed with every table.
- **Domain** = the registrable domain (public suffix + 1; `www.` dropped). Subdomains are kept in the URL field.
- **Resolution:** Gemini redirect links are resolved at run time (`redirect: manual`, follow up to 5 hops, 10 s timeout); `uri`, `url`, `resolveStatus` and `resolvedAt` are stored. Unresolved links are coded `other/unresolvable` and counted in the denominator.

## 6. Instrument (`panel/run.mjs`, adapted from `docs/seo/audit-2026-09-02/geo-panel/run-llm.mjs`)

Keep: 3 draws, KL `user_location`, run-time Gemini link resolution, JSONL append, `searched` flag, model recorded.
Change: prompts file and output directory as arguments; window id, `startedAt`/`finishedAt` and `modelVersion` per row; keep **every** `output_text` part; store raw and normalised citation URLs; a `meta` header row with the sha256 of the script and the prompts file; refuse to start if the window's output file exists, if either key is missing, or if the OpenAI one-call check fails; exit non-zero if every call in a lane fails; `retryOf` for explicit re-runs. Unit tests with `fetch` mocked (pattern: `bin/__tests__/geo-visibility.test.js`), including the refusal paths; the same pinned prompts file is a fixture.

## 7. Source taxonomy and coding rules (pre-registered)

Each (trial, URL) row gets one **source type**, one **publisher country**, one **page type**, and a **basis** note (what was looked at: resolved page title/H1/about page, or domain only when the page was blocked).

| Code | Definition | Examples | Decision rule |
|---|---|---|---|
| `agency-owned-service` | A page on the site of a business whose primary offer is marketing, web, creative or advertising services, other than a listicle | an agency's SEO service page, homepage, case study | site's own about/services page shows it sells such services |
| `agency-owned-listicle` | A "top N" or comparison article published on an agency's own site | "10 best F&B agencies in Malaysia" on an agency blog | same site test as above AND the page ranks or lists multiple providers |
| `freelancer-site` | A personal site of an individual selling such services | a freelance designer's portfolio | individual, not a company |
| `independent-listicle` | A "top N", comparison or review article on a site that does not sell those services | a blog or magazine ranking agencies | publisher fails the agency test |
| `directory-review` | A directory or review platform profile or listing | Clutch, GoodFirms, Sortlist, DesignRush, TechBehemoths, Listing.my | the URL is a profile/listing or the platform's own category page |
| `press-release` | A newswire release or its syndicated copies | EIN Presswire, syndicated copies on TV-station or law-review sites | page is labelled press release or carries a newswire byline |
| `social-profile` | A profile or post on a social network | Facebook page, Instagram, TikTok, X, YouTube | host is a social network |
| `company-profile` | A company or employer profile on a job or business-registry site | Hiredly, LinkedIn company page, Indeed, Crunchbase, companieshouse.my | host is a job board, professional network or registry |
| `marketplace` | A freelance or services marketplace | Upwork, Fiverr, Kaodim | host is a marketplace |
| `news-media` | An editorial news or industry-media article | The Star, Marketing-Interactive, Marketech APAC | editorial byline; a labelled sponsored article is coded `news-media` with `sponsored: true` |
| `government-association` | Government, statutory body, industry association, university | MDEC, SME Corp, 4As (aaaa.org.my), MCMC | `.gov.my`, `.edu.my`, or a recognised association |
| `google-property` | A Google-owned page | Google Business Profile / Maps, support.google.com, business.google.com | host is google.com or a Google product domain |
| `vendor-docs` | Documentation or blog of a platform or software vendor that does not sell agency services | Meta Business Help, Google Ads Help, Semrush blog, HubSpot blog | vendor of tools, not of services |
| `other-unresolvable` | Dead link, blocked, no content, or none of the above | 404, 403 to our fetch, empty page | record the HTTP status |

**Precedence when two codes fit:** the host decides first (social, directory, marketplace, government, google, company-profile, press wire), then the publisher test (agency vs not), then the page shape (listicle vs service page). Aliq's own pages are coded by the same rules.

**Publisher country:** `MY` if the registrable domain ends in `.my`, or the site states a Malaysian address or Malaysian legal entity on its contact or about page, or the platform page is a Malaysian entity's profile; `foreign` if it states a non-Malaysian base; `unknown` otherwise. For platforms (Clutch, Facebook), the country is that of the profiled entity where the profile makes it clear, else `platform` (reported as its own bucket).

**Page type:** `homepage` (empty path or `/` after locale prefix), `service-or-industry-page` (title or H1 names a service or an industry), `article-listicle` (blog post, guide, ranking), `profile-listing`, `other`.

**Agency for anonymisation:** any `agency-owned-*` or `freelancer-site` domain becomes `[agency]`, `[agency]` … numbered by first appearance in W1 Gemini then W1 ChatGPT then W2. The mapping lives only in `analysis/agency-map.internal.json` (gitignored). All other domains are named.

## 8. Coding process and reliability

1. The controller codes every row from the resolved page (fetch title, H1 and about/contact page; 10 s timeout; basis recorded). Blocked pages: coded from domain and platform knowledge, basis = `domain-only`.
2. A second reader (the named reviewer or Jinn; **from 2026-09-27 the reviewer is Ricky Wong**, see README W3) re-codes a random 20% sample of rows (seeded RNG, seed recorded in `analysis/reliability.md`), blind to the first coding, on source type and publisher country, using `analysis/codebook.md` only.
3. Report simple agreement and Cohen's kappa for each field. **Pre-stated rule:** if simple agreement on source type is below 80%, the rules are revised, the revision logged, and the full set re-coded by both readers on a fresh sample.
4. Disagreements are resolved by discussion and each resolution is logged. An AI assistant may draft codes for review but is never the sole coder of record for the sample.

## 9. Google sub-panel (exploratory)

12 prompts (the `-P-a` and `-I-a` prompts of the six non-GEO categories) × 1 draw per window, through the connected Chrome extension: navigate to the search URL with `gl=my&hl=en`, then the same with `udm=50`; wait for the AI block; capture page text and the links inside the AI block; record `rendered: yes/no`, `walled: yes/no`. Personalised state (Jinn's signed-in Chrome) is recorded and the sub-panel is reported as exploratory with that caveat. If the extension is unavailable on a run day, the sub-panel is "not run" for that window.

## 10. What this design cannot establish

Causation (why a source is cited); the behaviour of the consumer Gemini and ChatGPT apps; behaviour on Malay or Chinese prompts beyond the six-prompt BM arm; representativeness for all Malaysian SME questions (the panel is designed, not sampled); the quality of any provider; any change over time longer than the gap between the two windows; anything about Aliq's visibility (not a finding of this paper).

## 11. Privacy, ethics, disclosure

No client-confidential text, name or price is sent to any API. Competing agencies are anonymised in every public artefact. Aliq's commercial interest (it sells AI-visibility services) is disclosed in the paper. No system is instructed to favour Aliq; no prompt names Aliq.

## 12. Data package (public)

`public/research/ai-answer-citations-malaysia/trials.csv` with columns: `window, engine, model, modelVersion, promptId, arm, category, intent, phrasing, evidenceTag, draw, status, searched, citationIndex, domainAnon, urlAnon (path kept for platforms, replaced by "[agency]" for anonymised agencies), sourceType, publisherCountry, pageType, sponsored, resolveStatus`. Answer text is **withheld** (it names competitors); the codebook and this protocol are published beside the CSV. Internal copies keep everything. **Amended again 2026-10-05 (D5, Jinn):** `domainAnon` carries no per-website pseudonym: every agency is `[agency]` and every unnamed website `[unnamed]`, so neither ALIQ Studio's rows nor any competitor's can be matched or tracked across rows. **Amended 2026-10-05 (§14):** `urlAnon` withholds the path of every page about one business and of any page whose path carries an anonymised name; the websites in `analysis/unnamed.json` are unnamed.

## 13. Pre-registered deliverable tables

T1 trials by engine × window × status · T2 search rate by engine × intent · T3 source-type share by engine (pooled windows) with counts · T4 provider-owned vs third-party by engine · T5 publisher country by engine · T6 named platforms (top 10 by citations; distinct prompts) · T7 W1→W2 overlap per engine (median Jaccard; share of domains reappearing) · T8 page type among agency-owned citations · T9 concentration (domains for 50% / 80%) · T10 exploratory: BM arm and Google sub-panel · T11 coding reliability. All generated by `analysis/build.mjs`; the paper's numbers are never typed by hand.

## 14. Deviations log

| Date | Deviation | Reason | Reported in paper? |
|---|---|---|---|
| 2026-09-26 | W1 split across four files: `runs-W1.jsonl` (OpenAI WEB-P-a d1–3 and WEB-P-b d1–2, then stopped by the controller), `runs-W1-openai-resume1.jsonl` (WEB-P-b d3 via the new `--first-draw` flag), `runs-W1-openai-rest.jsonl` (the other 46 prompts), `runs-W1-gemini.jsonl` (all 48). Engines now run in parallel rather than one after the other. No trial was discarded or repeated; the instrument gained `--first-draw` (tested, mutant killed) so its sha256 differs between the first file and the rest. | The first run was launched inside a harness call that can be stopped after ten minutes while a window takes about 100; the controller stopped it cleanly after 5 trials and relaunched both lanes as detached processes. | Yes, one line in the method section |
| 2026-09-27 | W1 resumed after a 25-hour pause: stopped 2026-09-26 ~04:10 KL, resumed 2026-09-27 05:38 KL, still inside the 48-hour window. Four more files: `runs-W1-openai-resume2.jsonl` (SEO-P-a d3), `runs-W1-openai-rest2.jsonl` (the other 41 prompts), `runs-W1-gemini-resume1.jsonl` (SEO-P-a d2–3) and `runs-W1-gemini-rest.jsonl` (the other 41). The instrument and prompts sha256 match the files from 2026-09-26 (`72e3ad0e…`, `24675786…`). Before launch, the recorded trials were checked against the resume lists: 39 recorded, 249 missing, 0 duplicates. As a result, W1 spans about 27 hours of clock time with a 25-hour gap, and prompt SEO-P-a's draws straddle the gap (OpenAI d1–2 before, d3 after; Gemini d1 before, d2–3 after). One failed launch ran in the wrong directory with no env file and made no API call. | Laptop restart (Jinn). | Yes, in the method section's timing line |
| 2026-09-27 | Google sub-panel W1: 24 of 24 captures (12 prompts × AI Overview and AI Mode). WEB-P-a was captured 2026-09-26 ~03:30 KL and the other 11 on 2026-09-27 05:45–06:00 KL. Method notes follow. **(a)** On the standard results page, the AI block on all 11 prompts captured 09-27 is headed "AI Mode reply for …", not "AI Overview". It is recorded as `google-aio` as planned. **(b)** The Chrome extension's privacy filter blocks any result that carries a query string, so query strings were cut from cited links. This affects 4 recorded links: 3 × `youtube.com/watch`, which lost the video id, and 1 × `ipa.com.my/coursedetails.aspx`. Domain-level coding is unaffected. **(c)** Container rule. WEB-I-a to CON-I-a used v1: the ancestor of the AI heading just below the results container. On ADS-P-a, v1 over-reached into the organic results and a private Google Ads account panel. That capture was discarded, never recorded, and the page was reloaded. ADS-P-a onward used v2: the largest ancestor whose text equals the AI block's. Every v1 text was checked to start with the answer, not with a panel. **(d)** Two standard pages were loaded twice: WEB-I-a, while the extractor was being built (the first load showed 5 links, of which 3 are shared with the recorded load), and ADS-P-a, after the v1 failure. In both cases the recorded capture is the later load. **(e)** SOC-P-a's AI block included Google's location widget, which names the browser's home area. That line was removed from the stored text as personal data. It shows the sub-panel is location-personalised. **(f)** BRD-I-a in AI Mode answered with no sources, confirmed on screen. It is recorded as completed with 0 links. **(g)** The WEB-I-a AI Mode text excerpt is about 140 characters, not 600, because the tool truncated the first read. | Tooling constraints found during the capture. | (a), (b), (e) and (f): yes, in the sub-panel caveat. The rest: internal. |
| 2026-09-27 | Method limitation, not a deviation, found by Lane A and confirmed by the controller at source (the OpenAI web-search guide, saved page A10): `url_citation` annotations "show only the most relevant references", while a separate `sources` field (`include: ["web_search_call.action.sources"]`) lists every URL consulted. The instrument records `url_citation` only, as pre-registered in §5, and stays unchanged for W2 so both windows match. The paper says ChatGPT citations here are the inline references a user sees, a subset of the pages consulted. | Found by source verification. | Yes, in limitations |
| 2026-09-27 | W1 OpenAI: 75 trials (25 prompts × 3 draws, CON-C-b onward) failed with `insufficient_quota` at 05:57 KL. After Jinn topped up, they were re-run 11:04–11:38 KL as `retryOf` rows (`runs-W1-openai-retry1.jsonl`, 75/75 completed). The analysis uses the latest attempt per trial (`resolveTrials`: a retry may only follow a failed attempt). The failed first attempts are kept and reported in T1. W1 therefore closed at 2026-09-27 11:38 KL, 32 hours after it opened. | Prepaid API balance ran out mid-window; a one-call pre-flight check does not prove the balance covers the window. | Yes, in T1 and one line in the method section |
| 2026-09-27 | Clarifications fixed before W2. **(a)** Retries (§5) apply to *errors* only. A *no-answer* is a result, reported and never re-run, since re-running until an answer appears would be cherry-picking. A retry must start at the first failed draw and may not re-run a completed draw (`panel/window-status.mjs`, tested). **(b)** Every W2 Google sub-panel capture uses the final W1 extractor ("rule v2", `panel/google-extract.mjs`). **(c)** The reliability sample's seed is fixed in advance in `analysis/reliability.md` (20261004), before W2 or any draw. The blind sheet groups pages by website to reduce the reviewer's effort; the sample itself is unchanged (random 20% of rows, §8). | Decided before W2 so both windows run under stated rules. | (a) and (c): yes, one line in the method section. (b): internal. |
| 2026-10-04 | Google sub-panel W2: 24 of 24 captures, pages loaded 11:39–11:53 KL (recorded 11:43–11:53), one load per page, no reloads; 0 walled, 0 location widgets, 0 Ads panels, 0 query strings stripped. **(a) Extractor changed for the standard results page (`AIO_W2` in `panel/google-extract.mjs`).** Google now also shows a visible "AI Overview" label (a `div` with `role=heading`, 11 characters) above the AI block. The W1 `AIO` matches that label first. Its container never reaches the 100-character stability floor, so the script ran past the extension's 45 s limit on the first W2 capture (WEB-P-a) and stored nothing. The hidden "AI Mode reply for …" heading that every W1 capture from 09-27 climbed from is still there (recorded as `heading` on all 12 W2 captures). `AIO_W2` changes one line: it prefers that heading and falls back to `AIO`'s original match. The container rule, waits and filters are unchanged. It was run on the same WEB-P-a page load (not reloaded) and on all 11 others. Driven in headless Chromium by `qa/check-aio-w2.mjs`: on a W1-shaped fixture both extractors give an identical capture; on the 10-04 shape `AIO` takes the label (0 links, never stable) and `AIO_W2` gives the W1 capture; 4/4 PASS, and removing the preference turns the last check RED (mutant killed). `AI_MODE` is unchanged. **(b)** The AI Mode text excerpt now begins with the query: the page echoes it twice, and the extractor strips only the first. This affects the stored 600-character excerpt, not the links. **(c)** The Chrome window was hidden (`document.visibilityState` = `hidden`, checked at the end; screenshots timed out), since Jinn had left. W1 was captured with Jinn present. Links still rendered (3–15 per AI Mode answer elsewhere). BRD-P-a in AI Mode gave 0 external links (its answer shows a map of local businesses, whose links are Google hosts and are filtered). This could not be confirmed on screen as W1's BRD-I-a was. On BRD-I-a's AI Mode page, the whole document held the same 1 external link as the main region, so nothing sat outside the extractor's scope. Recorded as completed with 0 links, per the W1 rule. | Google page change found during the capture; Jinn away, so the window could not be on screen. | (a) and (c): yes, in the sub-panel caveat. (b): internal. |
| 2026-10-04 | W2 API panel run as planned in one launch (`panel/run-window.sh W2`, 11:38:33 KL, inside the window from 11:38 on 4 October to 03:21 on 6 October). Two files: `runs-W2-openai.jsonl` (144/144 completed, 11:38–12:40 KL) and `runs-W2-gemini.jsonl` (144/144 completed, 11:38–12:25 KL). 0 errors, so no retries. The instrument and prompts sha256 match W1 (`72e3ad0e…`, `24675786…`, checked by the launcher). Models as pinned: `gpt-5.5-2026-04-23` and `gemini-3.8-flash`. Not a deviation; recorded so the method section's timing line has a source. W1 to W2 gap: 7 days from W1's close, 8 days from its start. Coding: 119 new domains and 366 new URLs coded by the controller on 4 October under the codebook, with one resolved case added before any row was coded under it (R10: link shorteners, file hosts and web archives are coded by the publisher they lead to). Every W1 code was checked unchanged afterwards (373 domains, 802 URLs). | Planned run. | Timing and gap: yes, in the method section. R10: in the published codebook. |
| 2026-10-04 | Analysis correction, not a design change. §5 says unresolved Gemini links are coded `other/unresolvable`, but `analysis/tables.mjs` had never applied that rule, because W1 had no such link. In W2 one Gemini link (GEO-P-b draw 1) failed to resolve at run time (`error:fetch failed`) and stayed on Google's redirect host, so it was counted as a Google page. `citationUnits` now applies §5: a link still on `vertexaisearch.cloud.google.com` is `other-unresolvable` ("dead, blocked or unknown"), country `unknown`, and is published as `[unresolved redirect]`. Test-first (RED shown), mutant killed. Effect: Gemini Google-property 1 → 0 of 1,513 and other-unresolvable 27 → 28. The other Gemini links with a failed fetch (8 fetch errors, 4 timeouts, 67 × 403, 4 × 404) had already left the redirect service for the publisher's own address, so their publisher is known and they keep their domain codes. The link was not in the reliability sample, and the fix does not change the number or order of units. | Found while checking T6 for the draft. | Internal (the corrected tables are what the paper reports). |
| 2026-10-05 | Disagreements from the reliability check (§8.4) resolved and applied to every row. **(a) Rules clarified after seeing the reliability data** (codebook R11–R13, Jinn 2026-10-04, agreed by the reviewer): a platform page profiling one business takes that business's country, a platform company's own pages take the company's country, and `foreign` needs a stated base. R11 was already in the codebook but could not apply, because country was coded per domain; `urls.csv` gained a per-URL `country` with its basis, carried by `prepare-coding.mjs`. **(b)** The reviewer's 39 decisions applied, except two that an existing rule decides ([agency] by R7; readkong.com by R10, with scribd.com coded the same way), each decided by Jinn. **(c)** The R13 re-check read the about or contact pages of 58 domains (the `unknown` ones the reviewer did not decide, plus `foreign` resting on general knowledge, plus fb.com), and the platform pages were coded page by page: 103 per-URL countries. Evidence: `analysis/recheck/`. **(d) Analysis correction:** T11 had compared the reviewer's codes with the live sheets, which after resolution would have reported agreement with the resolution; it now reads the codes frozen at the draw (`key.internal.csv`) and refuses to run without them. T11 is unchanged (92.6% / 0.87; 88.9% / 0.57). Every change is in `analysis/reliability/applied-2026-10-05.csv` (366 rows). **Effect:** T5 moves most (ChatGPT `platform` 17.0% → 2.5%, `foreign` 2.3% → 14.5%; Gemini `unknown` 4.5% → 1.1%); T3 agency lists Gemini 14.5% → 13.0%; T4 provider share Gemini 87.1% → 87.2%. Claim F08 reworded, method claim M19 added. |
| 2026-10-05 | R13 narrowed by Jinn (codebook R14, verbatim: "known website like meta tatler investing.com can be listed as foreign"): a well-known company or publisher based outside Malaysia is `foreign` by general knowledge. Applied to meta.com, fb.com, tatlerasia.com, investing.com, arxiv.org, statcounter.com and Facebook's 7 own pages (ledger step D). T11 unchanged; T5 ChatGPT `foreign` 14.5% → 15.5%, Gemini 5.4% → 5.8%. M19 reworded. |
| 2026-10-05 | **Second-reader round** (one Opus review agent, read-only, approved by Jinn; report `lanes/report-second-reader-2026-10-05.md`; every finding reproduced by the controller before it was acted on). **(a) Exploratory T10e added after seeing the data:** provider-owned share and source types by engine × kind of question (`tables.json` `T10.byEngineIntent`), because the pooled T4 shares answer the "who to hire" question with all three kinds of question mixed (ChatGPT 84.3% on provider-selection questions, 18.8% on how-it-works ones). The pooled T4 stays the pre-registered headline. **(b) Vendors unnamed (codebook R2 kept):** four vendors cited through their own pages selling marketing, web or advertising services (`analysis/unnamed.json`) are left out of T6 and pseudonymised `Vendor-NN` in the package; no code changes, T6's tenth row is now handld.org. **(c) §12 amended:** `urlAnon` withholds the path of every page about one business (social-profile, company-profile, directory-review, marketplace, press-release) and any other path that contains an anonymised name; this over-withholds where an agency's domain is a generic phrase (for example a website-design or video-production domain), which costs URL detail, never anonymity. **(d) Release gate:** `qa/render-draft.mjs --release` now also fails on an `unverified` claim ID and on a package whose ALIQ pseudonym has the disclosed row count (D5); the build still only warns, so it can run before the D5 decision. | The reader's P1-1 (a competitor named as "not a marketing provider"), P2-2/P2-3 (pooled figures read as one engine or one kind of question), b-2 (agency names in platform profile URLs in the public CSV), P2-9 and P2-11. Jinn ruled (a), (b) and (c) on 2026-10-05, each the recommended option. | (a) yes: labelled exploratory beside every use; (b) yes: one sentence under Table 4; (c) in the package README at release; (d) no |
| 2026-10-05 | Draft corrections from the same round, not design changes. **H3 was unreported:** §1's third working hypothesis (fewer than half of W1 domains reappear in W2) failed on both engines (55.8%, 54.2%) and the draft called H1 "the" working hypothesis; all three are now reported (H1 failed, H2 held, H3 failed). Also corrected: "82 of 84" are answers, not questions; the two sponsored or agency-supplied pages are both agency rankings (the draft said one), and the wording no longer points at the supplying agency; I01 and I04 are worded as "may"; the Bahasa Malaysia share uses all provider-owned codes; two AI Mode answers "had no links to outside websites" rather than "cited no sources"; the disclosure line covers the Google sub-panel (7 of 412 links). Claim statuses F04, F06 and F10 moved back to qualified. | Second reader's P1-2, P1-3, P1-4 and P2s, all reproduced by the controller. | yes (the corrected text) |
| 2026-10-05 | **Public URL rule hardened after an automated security review of the commit** ("anonymisation bypass / parser differential" in `publicUrl()`; a second finding was not delivered with details). Driven before fixing: the first rule checked the parsed path and query but returned the raw string, so fragments, percent-encoded names, subdomains and userinfo passed, and an unparseable URL crashed the build. None was live in the data, but a full read of all 355 non-provider cited pages found pages ABOUT one anonymised agency that no name match catches (a business-news profile, a trade-media appointment story, a Singapore government directory profile, a magazine profile of a founder). **Now:** only `google-property`, `vendor-docs` and `government-association` pages keep a full address, rebuilt from the checked parts (decoded; no fragment or userinfo); every other page shows `https://<registrable domain>/[path withheld]`; two government-coded profiles are listed by hand (`unnamed.json` `withholdPaths`); unparseable addresses are withheld. The unnamed list was widened from four vendors to 12 websites that sell these services but are coded otherwise (including the three sites read as agencies in the reliability step, pending Jinn's ruling on their codes); label `Unnamed-NN`. 12 mutants killed. No figure moved; agency numbers unchanged. | Security review; controller's own sweep. Within Jinn's 2026-10-05 "fix now" ruling on the package; the widening of the unnamed class goes past the case he ruled on and is reported to him. | Table 4 note (count generated); package README at release |
| 2026-10-05 | **The reviewer checked the unnamed list** (`analysis/reliability/unnamed-review.html`; answers `reliability/unnamed-ricky-2026-10-05.csv`, filed as downloaded). Rule agreed; 10 websites kept unnamed (now 7, see next); digimax.com.my and renthestudio.com named; **[agency], [agency] and [agency] recoded from `other-unresolvable` to `agency`** (blocked at first coding, read as agencies later; their cited pages are site roots, page type `homepage`); [unnamed website] kept (the reviewer could not open it either). 14 checked edits appended to `reliability/applied-2026-10-05.csv` as step E. **Effect:** provider-owned Gemini 87.2% → 87.4%, ChatGPT 59.7% → 59.9%; who-to-hire 85.9% → 86.1% and 84.3% → 84.8%; BM ChatGPT 61.5% → 62.4%; T8 homepage 4.1% → 4.3% and 18.2% → 18.5%. T11 unchanged (scored against the frozen first codes). Agency numbers re-sorted by first appearance (§7): 258 moved; ALIQ stays [agency]; nothing had been published. | A coding correction decided after seeing the data, by the human reviewer, for three sites the first coding could not read. | yes, through the generated figures; this row is published with the protocol |
| 2026-10-05 | **The last two domain-grain questions closed (Jinn, applying the controller's recommendation).** (a) **semrush.com stays `directory-review`** although two of its cited pages (its homepage and "AI Visibility Toolkit Pricing") are its own product pages, which would be `vendor-docs` at page grain. Both are third-party either way, so no headline figure depends on it; the directory share it inflates is ChatGPT's 4.2% (about 4.0% without them). Source type stays one code per website, apart from the existing per-page list and press-release flags. (b) **[unnamed website] keeps `marketplace` but joins `unnamed.json`**: its one cited page sells link building itself, so it falls under the unnamed rule the reviewer agreed. No figure moved; the unnamed count is now 8. | Open items from the reliability step (`analysis/reliability.md`). | (a) no: a known limit of domain-grain coding, recorded here; (b) through the Table 4 note's generated count |
| 2026-10-05 | **D5 decided (Jinn, at the release checklist): every agency is `[agency]` in the public package** (unnamed websites `[unnamed]`), instead of `Agency-NN` pseudonyms. Measured first: under the pseudonyms ALIQ Studio was the only agency with exactly the disclosed 75 rows, and even a rounded disclosure ("under 3%") would leave it one of 3 to 5 candidates. The disclosure line stays exact. Cost: readers cannot recompute the per-website overlap (T7) and concentration (T9) for agency rows from the CSV; both tables are published. `render-draft --release` now refuses any per-website pseudonym. No figure moved. Also: the studio site gets an author page for the byline (Jinn). | D5 (2026-09-19 rule; deferred 2026-09-27 to release). | yes: Supporting material describes the dataset this way |
| 2026-10-05 | **Wording of RQ0 in the paper (Jinn, after the E1 extraction test):** "what sources does the answer rest on?" becomes "what sources does the answer cite?", and the findings section becomes "What the answers cited". Shares in the paper and its tables print to one decimal (81.0%, not 81%), and kappa and Jaccard values to two. | The extraction test read "rest on" as reliance, which the audit does not measure: it measures citations. Mixed precision (81% beside 91.1%) read as rounding. | yes: wording and display only; no measure, value or analysis changes |
| 2026-10-06 | **Prompt phrasings, found after publication (E2 extraction test, controller check of `panel/prompts.json`).** §2 describes the two phrasings as "one generic-Malaysia, one local with a business type". The frozen file (v1.0, 26 September 2026) does not follow that for every pair: 13 of the 42 English prompts name a place (Kuala Lumpur, Selangor or Petaling Jaya), 10 of them phrasing "b" and 3 phrasing "a", and 11 "b" prompts ask about Malaysia as a whole. The prompts were run as frozen; nothing was re-run. | The design text was written before the prompt file and was not reconciled with it; the paper's method sentence (claim M17) was taken from the design. | Yes: the method section now describes the file as it is (13 local, 29 Malaysia-wide); no reported figure splits by phrasing. |

## 15. Approval

- Jinn: ☑ approved with all three open items included (GEO category, BM arm, Google sub-panel) · date: 2026-09-26 (planning session, AskUserQuestion record) · `panel/prompts.json` set to version 1.0, frozen: true.
- prompts.json sha256 at freeze: `2467578671c86e993875ed997ef9dc1f3cc222a0a5c5ed64a294f2806c6da98b`
