02 Chapter 2 · 智 AI

Skipping the IME? 跳过输入法?

Can you skip the input method and just type pinyin to an AI agent? Forty-eight real calls say: for routine work, yes — but homophones collapse, and one word kept in Hanzi brings the meaning back. The first AI paper, accepted at UKAIRS 2026 in Edinburgh.

Fig. 2 — Toneless pinyin as a walk that can be absorbed by the wrong word.

Fig. 2 shi shi: six words, one keystring Honesty tag: measured

跳过输入法?

  1. 跳过
  2. 挑过
  3. 调过
  4. 条过not a word

to skip the input method

AcceptedPoster

Skipping the IME? Keyboard-Friendly Multilingual Prompting for LLM Agents

Xiaoxiao Zhouyi · Yanzhen Li

With
Yanzhen Li
Track
Emerging Research · Poster
Venue
UK AI Research Symposium (UKAIRS) 2026, Edinburgh · 24–25 Nov 2026
Poster
#181 · Lunch & Poster Session 3 · 25 Nov 2026 · 13:15–14:15
Accepted
11 Aug 2026

Fig. 2.1 shi shi measured

  1. 01Typing, no IME
  2. 02What the agent hears
  3. 03The mixed prompt
  4. 04The score

01 / 04 · Typing, no IME

You type the sound.

7 keystrokes, no candidate list, no tones. Fast — and a little lossy.

02 / 04 · What the agent hears

The agent hears 6 words.

Tones would tell them apart. Toneless pinyin collapses fact, implementation, real time and the rest into one string.

03 / 04 · The mixed prompt

Keep the one that matters in Hanzi.

Scaffolding in pinyin, the critical term in native script. Ambiguity gone, speed kept.

04 / 04 · The score

One third, or all of it.

Homophone stress tasks — Codex (n = 15) and a Claude Code check (n = 9): pinyin-only prompts averaged 0.333; native and mixed prompts scored 1.00.

Pinyin only
0.333
Hanzi anchored
1.00
  1. 事实 shìshí fact
  2. 实施 shíshī implementation
  3. 实时 shíshí real time
  4. 时事 shíshì current affairs
  5. 史诗 shǐshī epic
  6. 失事 shīshì accident

Toneless pinyin reaches the agent as one string. Each candidate inks in by first-passage percolation: ink enters where the brush lands and every pixel darkens at its first-passage time (arrival fields baked at build).

Fig. 2.2 Homophone star illustrative

Type or pick a keystring

  1. shi shi
  2. shu
  3. mai

0.13P(you meant 事实) · Uniform guess among the homophones (no context).

Homophone star · shi shi事实shìshífact 13%YOU MEANT实施shíshīimplementation 13%实时shíshíreal-time 13%时事shíshìcurrent affairs 13%史诗shǐshīepic 13%逝世shìshìpassing away 13%失事shīshìaccident 13%誓师shìshīrally 13%shi shi
  • 事实 shìshí fact
  • 实施 shíshī implementation
  • 实时 shíshí real-time
  • 时事 shíshì current affairs
  • 史诗 shǐshī epic
  • 逝世 shìshì passing away
  • 失事 shīshì accident
  • 誓师 shìshī rally
Fig. 2.2 Toneless pinyin as a random walk on a star: every homophone is an absorbing target, and the toned spelling under each is the information the keystrokes deleted. Illustrative model (uniform prior plus a context drift), not measured model probabilities. Typing the critical term in Hanzi is a shortcut straight to the right target.

Table 2.1 What the agents actually did measured

ProbeComparisonQualityn
Arithmetic objectiveCodexall six input conditionsall 1.006
Cycling JSONCodexnon-initial modes vs initials0.3331.006
Abbreviation reliabilityCodexnative vs initials in two modes0.333 / 0.4441.0012
Homophone stressCodexnative / mixed vs pinyin-only modes0.3331.0015
Homophone checkClaude Codenative / mixed vs clean pinyin0.3331.009
48 valid calls · zh-CN · text-only · deterministic graders48

What came back

  1. shi shi jian yan expected: 事实检验 · 实施检验 fact check · implementation check

    • Codex · Native script事实检验 · 实施检验✓
    • Codex · Clean pinyinshi shi jian yan · shi shi jian yan✗
    • Codex · No spaces事实检验 · 事实检验✗
    • Codex · Mixed事实检验 · 实施检验✓
    • Claude Code · Clean pinyin实时检验 · 实时检验✗
  2. shu expected: 书 · 树 · 输 · 熟 book · tree · lose/input · ripe

    • Codex · Native script书 · 树 · 输 · 熟✓
    • Codex · Clean pinyin书 · 树 · 数 · 输✗
    • Codex · No spacesshu · shu · shu · shu✗
    • Codex · Mixed书 · 树 · 输 · 熟✓

Skipping the IME may be viable for routine scaffolding, but exact names, domain terms and homophones need native-script anchoring.

The emerging design rule · Emerging Research · Poster

Input styles the scaffold already defines

  1. Chinese pinyinloses tone and the final character choice
  2. Japanese romajiloses kanji choice
  3. Vietnamese keyboard inputloses diacritics
  4. Hinglishloses native-script spelling
  5. Arabiziloses script-specific spelling
  6. Russian transliterationloses native-script names
  7. Greeklishloses native-script names
  8. Korean keyboard-friendly inputloses native-script names

Your turn · the collapse

Type toneless pinyin

A tiny dictionary shows how many words share your keystrokes. An agent reading raw pinyin has to guess among them.

Try

  1. shi shi
  2. shou da
  3. zhou yi xiao xiao

shou da

  1. 首达first passage
  2. 手打typed by hand

2 words share these keystrokes.

Toneless, my two research lines are one word: 首达 shǒudá, first passage — and 手打 shǒudǎ, typing by hand.

Fig. 2 Toneless pinyin types itself; candidate words ink in; proof marks keep 事实 and strike the rest; the score moves from 0.333 to 1.00 when the key term stays in Hanzi.

Exhibits: papers

Skipping the IME? Keyboard-Friendly Multilingual Prompting for LLM Agents

Xiaoxiao Zhouyi, Yanzhen Li

UKAIRS 2026 · Emerging Research (poster), Edinburgh ·

Accepted · poster #181, Edinburgh, 25 Nov 2026

The record

Chinese speakers type pinyin and let an input method choose the characters. What if you skip that last step and send the pinyin straight to an AI agent? Faster — but 事实 (fact) and 实施 (implementation) are both ‘shi shi’. This chapter is the study that measured where meaning collapses, and the rule it arrived at.

ime-eval drives real coding agents with paired prompts — native script against keyboard-friendly variants — under strict rules: dry-run by default, every real call authorised, results written once. It moved in phases, each audited by an independent agent before any spend, and it kept its negative results: asking the model to restore the Chinese first made things worse.

The finding is small and sharp: routine instructions survive pinyin, but homophones do not — unless the critical word stays in Hanzi, which restores full correctness. The paper was accepted at UKAIRS 2026 in Edinburgh. Around it sits a way of working: models propose, other agents attack, computation decides.

An AI-native way of doing mathematics

Across the first-passage line a repeatable loop took shape: AI models propose derivations and code, separate agents attack the claims, numerical computation acts as the arbiter, and exact algebra is machine-checked in Lean 4. The documented cases include proposed formulas that failed numerical checks and were discarded, a repository-grounded audit that caught errors an outside model could not see, and multi-round reviews — up to about sixty agents in a single audit — before each submission. The principle in one line: AI can generate more scientific possibilities; the research skill is deciding which ones deserve belief.

Ongoing

Agents in one audit
≈ 60
Lean-checked theorems (fold paper)
46

A research line that bends toward AI

Alongside first-passage research on random walks, the PhD work broadened during 2026 toward AI: evaluating LLM agents, human–AI interaction and multilingual input, and using multi-model AI workflows as research instruments — adversarial audits whose disagreements are settled by computation. The habit carries straight over: measure the hidden route, then refuse to believe the number until something independent confirms it. The UKAIRS 2026 acceptance is its first peer-reviewed AI output.

Ongoing

GOSIM Paris 2026

Attended GOSIM Paris 2026, ‘The Agentic AI Convergence’, at Station F on 5–6 May. The takeaway, written up afterwards: open-source agent ecosystems are moving away from model showcases toward tool interoperability, model serving, edge deployment, production systems and trustworthy workflows — the same direction the evaluation work was heading.

Attended

ime-eval: a harness for keyboard-friendly prompting

The question was personal: can I just type pinyin at my agent? ime-eval turns it into an instrument. It drives real command-line coding agents rather than raw APIs, pairs every native-script prompt with keyboard-friendly variants — clean pinyin, no spaces, light typos, initials, and mixed prompts that keep key terms in native script — and defines profiles for eight input styles, from Japanese romaji and Vietnamese keyboard input to Hinglish, Arabizi, Russian transliteration, Greeklish and Korean. Hard rules from day one: every configuration generated automatically, dry-run by default, results written once, and sandboxes that keep answers from leaking to the agent. The harness is named here but not yet public.

Completed

Input styles
8
Pilot prompt variants
144

A relay of three models, refereed by a human

The study ran as a relay between AI systems. An orchestrating model wrote each phase brief; a coding agent implemented and ran it; a third, independent agent audited readiness before any real call was spent — for the homophone phase it returned ‘safe to run’. Every phase ended with a state-and-decision packet handed back to the orchestrator, and every step was committed to git. The human role was the architect's: choose the phases, authorise the spend, carry the baton.

Completed

Independent readiness audits
4

Four calls, four bugs, then a clean 1.0

Before any model call was spent, the harness was hardened: leakage-neutral work directories, an environment-gated switch for real runs with a hard cap on calls, preflight checks and deterministic scoring, each reviewed by an independent audit. The first real smoke test took four authorised calls and exposed four invocation bugs — relative schema paths, a blocked input stream and the strict rules of structured output. Once they were fixed, the pipeline answered correctly and scored 1.0. The test suite grew from 47 to 87 along the way.

Completed

Calls to first success
4
Bugs found
4

Catching a silent duplicate

A zero-cost audit of all 24 pilot tasks found that the ‘abbreviation’ condition produced prompts byte-for-byte identical to clean romanisation on three tasks — one Chinese, two Vietnamese — which would have silently wasted real calls. A deterministic, script-aware transform now renders syllable initials for Chinese, Japanese and Korean (so ‘shang dian li’ becomes ‘sdl’), and the variant lock records a digest of every rendered prompt, so that even a code-only change is caught by drift checks.

Completed

Six input shapes: a ceiling, then a cliff

An arithmetic task came back correct under all six input shapes — good evidence that the plumbing worked, but a saturated test. A constraint-following task discriminated: clean pinyin, pinyin without spaces, pinyin with typos and a mixed prompt with English technical terms all returned the required three-item JSON array, while pinyin initials recovered the topic but lost the structure and scored 0.333. Under a placeholder keystroke model the initials needed an estimated 0.39 of the native keystrokes: the effort saved was real, and so was the quality lost.

Completed

Initials on the formatting task
0.333 (native 1.00)

A negative result: restoring the Chinese first made it worse

A repetition study of twelve calls confirmed that the abbreviation regression is stable: direct answers to pinyin-initial prompts averaged 0.444 against 1.0 for native prompts. A ‘normalise, then answer’ wrapper did not rescue it. The mean fell to 0.333 because the model judged the initials too lossy and declined to answer in all three repetitions — turning confident wrong output into honest abstention. Native prompts were unaffected. It is a negative result, and it stayed in the paper.

Completed

Direct vs normalise-first
0.444 vs 0.333

The key finding: toneless pinyin erases homophones

Three adversarial tasks hinge on words that collapse in toneless pinyin: 事实 ‘fact’ and 实施 ‘implementation’ are both shi shi; 书, 树, 输 and 熟 are all shu. Native Hanzi scored 1.00, and so did the mixed workflow, where ordinary instruction is typed in pinyin but the critical terms stay in Hanzi. Clean pinyin, pinyin without spaces and pinyin with typos all averaged 0.333 — the agent either left the answer in pinyin or picked the wrong character. The one task that survived, 买 ‘buy’ against 卖 ‘sell’, did so because 入 and 出 disambiguate it. And the mixed prompt still saved effort: an estimated 0.815 of the native keystrokes, under a placeholder model.

Completed

Native Hanzi · mixed
1.00 · 1.00
Pinyin only (clean, no-space, typo)
0.333

A second agent, the same boundary

Nine real calls to Claude Code on the same three homophone tasks reproduced the pattern exactly: native and mixed prompts scored 1.00, clean pinyin averaged 0.333. The pinyin ‘shi shi jian yan’, for example, came back as 实时检验 (‘real-time check’) for both items, instead of 事实检验 (‘fact check’) and 实施检验 (‘implementation check’). Two earlier batches were excluded as infrastructure failures rather than counted. The check made the finding cross-agent — reported in the paper as a small check, not a model comparison.

Completed

Valid calls
9

When the tool changes underneath: pin the model

A prepared 48-call boundary study, with a new glossary condition, returned an error on every call. For reproducibility the harness ignored the command-line tool's user configuration, so it fell back to a built-in default model that the signed-in account could not use. The batch was reported as a failure rather than scored as data; the model was then pinned explicitly and a one-call smoke confirmed the fix. The audit trail surfaced a second lesson: the exact model behind the earlier successful phases had not been recorded. Product tools change underneath you — pin exact model IDs.

Completed

Submitted to UKAIRS 2026

‘Skipping the IME? Keyboard-Friendly Multilingual Prompting for LLM Agents’ went to the UK AI Research Symposium 2026 on 8 June: a two-page paper in the Emerging Research track, under the theme of synergistic human–AI collaboration, typeset so that the Chinese examples render inside an ACM template. It reports 48 valid text-only zh-CN calls — 39 to Codex and 9 to Claude Code — and proposes a design rule: skipping the IME may be fine for routine scaffolding, but exact names, domain terms and homophones need native-script anchoring. Authors: Xiaoxiao Zhouyi and Yanzhen Li, University of Bristol.

Valid calls
48 (39 + 9)
Pages
2

Scaling to eight scripts

The next study was built before a single call was spent: 44 tasks across eight languages (264 prompt variants), split into high-risk exact-script tasks and controls; a fixed matrix of frontier models; a statistics engine with paired bootstrap confidence intervals; and generators for every table and figure. An adversarial audit from eight native-fluency perspectives replaced four homophone pairs that did not truly collapse. The main run is waiting on a go/no-go decision, so there are no multilingual results yet — and none are claimed.

Planned

Tasks · languages · variants
44 · 8 · 264

Triangulated audit: three minds and a calculator

A reusable audit skill that attacks a research package from three angles over several rounds — a multi-agent workflow with access to the repository and tools, a strong external model, and a numerical and empirical arbiter — and settles every disagreement between models by computation rather than authority. The audit that motivated it showed why: the strongest external model got a formula and a threshold wrong that the other participants got right; only the union plus the arbiter was complete; and a mistaken refutation forced a certification that made the result stronger.

Shipped

What the next version will measure

What comes next: position the work against related research; match the title to the zh-CN evidence; soften the design rule to fit the small homophone samples; run a small typing study, so that both sides of the effort–quality trade-off are measured rather than modelled; build about thirty homophone tasks with graded ambiguity to trace a real performance curve; pin every model version; and look harder at why ‘normalise first’ hurts. The eight-language harness is ready for the broader claim.

Planned

Accepted: the first AI paper

On 11 August 2026 UKAIRS accepted the paper for a poster session in the Emerging Research track. The reviews found the question genuinely interesting, the artefact useful-looking and the inclusivity angle worth pursuing, and recommended leading with the ‘shi shi’ example and the three-layer analysis — encoding change, compression, tone loss. It is the first peer-reviewed AI paper, first-authored.

Accepted

Poster #181, Wednesday 25 November

The organisers assigned the poster to Wednesday 25 November 2026, 13:15–14:15, in Lunch & Poster Session 3, and sent the presentation guidance. The public poster list shows it as #181 under ‘F. Synergistic Human–AI Collaboration’, Emerging Research, Xiaoxiao Zhouyi, University of Bristol.

Scheduled

UKAIRS 2026, Edinburgh

The UK AI Research Symposium runs on 24–25 November 2026 at the John McIntyre Conference Centre in Edinburgh, and the poster is up on Wednesday 25 November. The plan, written into the paper: a side-by-side demo showing, for each task, the native prompt, its keyboard-friendly variant, the agent's response, the grade and the effort estimate — so that visitors can toggle from routine prompts to exact-term stress prompts and watch where meaning collapses. Chapter 02 of this site is a first draft of that demo.

Scheduled · 25 Nov 2026

Evidence ledger

Every public item in this chapter, with its date, status and link.
DateItemStatusLinks
An AI-native way of doing mathematicsOngoing
A research line that bends toward AIOngoing
GOSIM Paris 2026Attended
ime-eval: a harness for keyboard-friendly promptingCompleted
A relay of three models, refereed by a humanCompleted
Four calls, four bugs, then a clean 1.0Completed
Catching a silent duplicateCompleted
Six input shapes: a ceiling, then a cliffCompleted
A negative result: restoring the Chinese first made it worseCompleted
The key finding: toneless pinyin erases homophonesCompleted
A second agent, the same boundaryCompleted
When the tool changes underneath: pin the modelCompleted
Submitted to UKAIRS 2026
Scaling to eight scriptsPlanned
Triangulated audit: three minds and a calculatorShipped
What the next version will measurePlanned
Accepted: the first AI paperAccepted
Poster #181, Wednesday 25 NovemberScheduled
UKAIRS 2026, EdinburghScheduled · 25 Nov 2026