Skipping the IME? 跳过输入法?
Can you skip the input method and just type pinyin to an AI agent? Forty-eight real calls say: for routine work, yes — but homophones collapse, and one word kept in Hanzi brings the meaning back. The first AI paper, accepted at UKAIRS 2026 in Edinburgh.
Fig. 2 — Toneless pinyin as a walk that can be absorbed by the wrong word.
AcceptedPoster
Skipping the IME? Keyboard-Friendly Multilingual Prompting for LLM Agents
Xiaoxiao Zhouyi · Yanzhen Li
Fig. 2.1 shi shi measured
- 01Typing, no IME
- 02What the agent hears
- 03The mixed prompt
- 04The score
01 / 04 · Typing, no IME
You type the sound.
7 keystrokes, no candidate list, no tones. Fast — and a little lossy.
02 / 04 · What the agent hears
The agent hears 6 words.
Tones would tell them apart. Toneless pinyin collapses fact, implementation, real time and the rest into one string.
03 / 04 · The mixed prompt
Keep the one that matters in Hanzi.
Scaffolding in pinyin, the critical term in native script. Ambiguity gone, speed kept.
04 / 04 · The score
One third, or all of it.
Homophone stress tasks — Codex (n = 15) and a Claude Code check (n = 9): pinyin-only prompts averaged 0.333; native and mixed prompts scored 1.00.
- Pinyin only
- 0.333
- Hanzi anchored
- 1.00
- 事实 shìshí fact
- 实施 shíshī implementation
- 实时 shíshí real time
- 时事 shíshì current affairs
- 史诗 shǐshī epic
- 失事 shīshì accident
Toneless pinyin reaches the agent as one string. Each candidate inks in by first-passage percolation: ink enters where the brush lands and every pixel darkens at its first-passage time (arrival fields baked at build).
Fig. 2.2 Homophone star illustrative
Type or pick a keystring
- shi shi
- shu
- mai
0.13P(you meant 事实) · Uniform guess among the homophones (no context).
- 事实 shìshí fact
- 实施 shíshī implementation
- 实时 shíshí real-time
- 时事 shíshì current affairs
- 史诗 shǐshī epic
- 逝世 shìshì passing away
- 失事 shīshì accident
- 誓师 shìshī rally
Table 2.1 What the agents actually did measured
| Probe | Comparison | Quality | n |
|---|---|---|---|
| Arithmetic objectiveCodex | all six input conditions | all 1.00 | 6 |
| Cycling JSONCodex | non-initial modes vs initials | 0.3331.00 | 6 |
| Abbreviation reliabilityCodex | native vs initials in two modes | 0.333 / 0.4441.00 | 12 |
| Homophone stressCodex | native / mixed vs pinyin-only modes | 0.3331.00 | 15 |
| Homophone checkClaude Code | native / mixed vs clean pinyin | 0.3331.00 | 9 |
| 48 valid calls · zh-CN · text-only · deterministic graders | 48 | ||
What came back
shi shi jian yan expected: 事实检验 · 实施检验 fact check · implementation check
- Codex · Native script事实检验 · 实施检验✓
- Codex · Clean pinyinshi shi jian yan · shi shi jian yan✗
- Codex · No spaces事实检验 · 事实检验✗
- Codex · Mixed事实检验 · 实施检验✓
- Claude Code · Clean pinyin实时检验 · 实时检验✗
shu expected: 书 · 树 · 输 · 熟 book · tree · lose/input · ripe
- Codex · Native script书 · 树 · 输 · 熟✓
- Codex · Clean pinyin书 · 树 · 数 · 输✗
- Codex · No spacesshu · shu · shu · shu✗
- Codex · Mixed书 · 树 · 输 · 熟✓
Skipping the IME may be viable for routine scaffolding, but exact names, domain terms and homophones need native-script anchoring.
Input styles the scaffold already defines
- Chinese pinyinloses tone and the final character choice
- Japanese romajiloses kanji choice
- Vietnamese keyboard inputloses diacritics
- Hinglishloses native-script spelling
- Arabiziloses script-specific spelling
- Russian transliterationloses native-script names
- Greeklishloses native-script names
- Korean keyboard-friendly inputloses native-script names
Your turn · the collapse
Type toneless pinyin
A tiny dictionary shows how many words share your keystrokes. An agent reading raw pinyin has to guess among them.
Try
- shi shi
- shou da
- zhou yi xiao xiao
shou da
2 words share these keystrokes.
Toneless, my two research lines are one word: 首达 shǒudá, first passage — and 手打 shǒudǎ, typing by hand.
Exhibits: papers
Skipping the IME? Keyboard-Friendly Multilingual Prompting for LLM Agents
UKAIRS 2026 · Emerging Research (poster), Edinburgh ·
Accepted · poster #181, Edinburgh, 25 Nov 2026
The record
Chinese speakers type pinyin and let an input method choose the characters. What if you skip that last step and send the pinyin straight to an AI agent? Faster — but 事实 (fact) and 实施 (implementation) are both ‘shi shi’. This chapter is the study that measured where meaning collapses, and the rule it arrived at.
ime-eval drives real coding agents with paired prompts — native script against keyboard-friendly variants — under strict rules: dry-run by default, every real call authorised, results written once. It moved in phases, each audited by an independent agent before any spend, and it kept its negative results: asking the model to restore the Chinese first made things worse.
The finding is small and sharp: routine instructions survive pinyin, but homophones do not — unless the critical word stays in Hanzi, which restores full correctness. The paper was accepted at UKAIRS 2026 in Edinburgh. Around it sits a way of working: models propose, other agents attack, computation decides.
An AI-native way of doing mathematics
Across the first-passage line a repeatable loop took shape: AI models propose derivations and code, separate agents attack the claims, numerical computation acts as the arbiter, and exact algebra is machine-checked in Lean 4. The documented cases include proposed formulas that failed numerical checks and were discarded, a repository-grounded audit that caught errors an outside model could not see, and multi-round reviews — up to about sixty agents in a single audit — before each submission. The principle in one line: AI can generate more scientific possibilities; the research skill is deciding which ones deserve belief.
Ongoing
- Agents in one audit
- ≈ 60
- Lean-checked theorems (fold paper)
- 46
A research line that bends toward AI
Alongside first-passage research on random walks, the PhD work broadened during 2026 toward AI: evaluating LLM agents, human–AI interaction and multilingual input, and using multi-model AI workflows as research instruments — adversarial audits whose disagreements are settled by computation. The habit carries straight over: measure the hidden route, then refuse to believe the number until something independent confirms it. The UKAIRS 2026 acceptance is its first peer-reviewed AI output.
Ongoing
GOSIM Paris 2026
Attended GOSIM Paris 2026, ‘The Agentic AI Convergence’, at Station F on 5–6 May. The takeaway, written up afterwards: open-source agent ecosystems are moving away from model showcases toward tool interoperability, model serving, edge deployment, production systems and trustworthy workflows — the same direction the evaluation work was heading.
Attended
ime-eval: a harness for keyboard-friendly prompting
The question was personal: can I just type pinyin at my agent? ime-eval turns it into an instrument. It drives real command-line coding agents rather than raw APIs, pairs every native-script prompt with keyboard-friendly variants — clean pinyin, no spaces, light typos, initials, and mixed prompts that keep key terms in native script — and defines profiles for eight input styles, from Japanese romaji and Vietnamese keyboard input to Hinglish, Arabizi, Russian transliteration, Greeklish and Korean. Hard rules from day one: every configuration generated automatically, dry-run by default, results written once, and sandboxes that keep answers from leaking to the agent. The harness is named here but not yet public.
Completed
- Input styles
- 8
- Pilot prompt variants
- 144
A relay of three models, refereed by a human
The study ran as a relay between AI systems. An orchestrating model wrote each phase brief; a coding agent implemented and ran it; a third, independent agent audited readiness before any real call was spent — for the homophone phase it returned ‘safe to run’. Every phase ended with a state-and-decision packet handed back to the orchestrator, and every step was committed to git. The human role was the architect's: choose the phases, authorise the spend, carry the baton.
Completed
- Independent readiness audits
- 4
Four calls, four bugs, then a clean 1.0
Before any model call was spent, the harness was hardened: leakage-neutral work directories, an environment-gated switch for real runs with a hard cap on calls, preflight checks and deterministic scoring, each reviewed by an independent audit. The first real smoke test took four authorised calls and exposed four invocation bugs — relative schema paths, a blocked input stream and the strict rules of structured output. Once they were fixed, the pipeline answered correctly and scored 1.0. The test suite grew from 47 to 87 along the way.
Completed
- Calls to first success
- 4
- Bugs found
- 4
Catching a silent duplicate
A zero-cost audit of all 24 pilot tasks found that the ‘abbreviation’ condition produced prompts byte-for-byte identical to clean romanisation on three tasks — one Chinese, two Vietnamese — which would have silently wasted real calls. A deterministic, script-aware transform now renders syllable initials for Chinese, Japanese and Korean (so ‘shang dian li’ becomes ‘sdl’), and the variant lock records a digest of every rendered prompt, so that even a code-only change is caught by drift checks.
Completed
Six input shapes: a ceiling, then a cliff
An arithmetic task came back correct under all six input shapes — good evidence that the plumbing worked, but a saturated test. A constraint-following task discriminated: clean pinyin, pinyin without spaces, pinyin with typos and a mixed prompt with English technical terms all returned the required three-item JSON array, while pinyin initials recovered the topic but lost the structure and scored 0.333. Under a placeholder keystroke model the initials needed an estimated 0.39 of the native keystrokes: the effort saved was real, and so was the quality lost.
Completed
- Initials on the formatting task
- 0.333 (native 1.00)
A negative result: restoring the Chinese first made it worse
A repetition study of twelve calls confirmed that the abbreviation regression is stable: direct answers to pinyin-initial prompts averaged 0.444 against 1.0 for native prompts. A ‘normalise, then answer’ wrapper did not rescue it. The mean fell to 0.333 because the model judged the initials too lossy and declined to answer in all three repetitions — turning confident wrong output into honest abstention. Native prompts were unaffected. It is a negative result, and it stayed in the paper.
Completed
- Direct vs normalise-first
- 0.444 vs 0.333
The key finding: toneless pinyin erases homophones
Three adversarial tasks hinge on words that collapse in toneless pinyin: 事实 ‘fact’ and 实施 ‘implementation’ are both shi shi; 书, 树, 输 and 熟 are all shu. Native Hanzi scored 1.00, and so did the mixed workflow, where ordinary instruction is typed in pinyin but the critical terms stay in Hanzi. Clean pinyin, pinyin without spaces and pinyin with typos all averaged 0.333 — the agent either left the answer in pinyin or picked the wrong character. The one task that survived, 买 ‘buy’ against 卖 ‘sell’, did so because 入 and 出 disambiguate it. And the mixed prompt still saved effort: an estimated 0.815 of the native keystrokes, under a placeholder model.
Completed
- Native Hanzi · mixed
- 1.00 · 1.00
- Pinyin only (clean, no-space, typo)
- 0.333
A second agent, the same boundary
Nine real calls to Claude Code on the same three homophone tasks reproduced the pattern exactly: native and mixed prompts scored 1.00, clean pinyin averaged 0.333. The pinyin ‘shi shi jian yan’, for example, came back as 实时检验 (‘real-time check’) for both items, instead of 事实检验 (‘fact check’) and 实施检验 (‘implementation check’). Two earlier batches were excluded as infrastructure failures rather than counted. The check made the finding cross-agent — reported in the paper as a small check, not a model comparison.
Completed
- Valid calls
- 9
When the tool changes underneath: pin the model
A prepared 48-call boundary study, with a new glossary condition, returned an error on every call. For reproducibility the harness ignored the command-line tool's user configuration, so it fell back to a built-in default model that the signed-in account could not use. The batch was reported as a failure rather than scored as data; the model was then pinned explicitly and a one-call smoke confirmed the fix. The audit trail surfaced a second lesson: the exact model behind the earlier successful phases had not been recorded. Product tools change underneath you — pin exact model IDs.
Completed
Submitted to UKAIRS 2026
‘Skipping the IME? Keyboard-Friendly Multilingual Prompting for LLM Agents’ went to the UK AI Research Symposium 2026 on 8 June: a two-page paper in the Emerging Research track, under the theme of synergistic human–AI collaboration, typeset so that the Chinese examples render inside an ACM template. It reports 48 valid text-only zh-CN calls — 39 to Codex and 9 to Claude Code — and proposes a design rule: skipping the IME may be fine for routine scaffolding, but exact names, domain terms and homophones need native-script anchoring. Authors: Xiaoxiao Zhouyi and Yanzhen Li, University of Bristol.
Submitted
- Valid calls
- 48 (39 + 9)
- Pages
- 2
Scaling to eight scripts
The next study was built before a single call was spent: 44 tasks across eight languages (264 prompt variants), split into high-risk exact-script tasks and controls; a fixed matrix of frontier models; a statistics engine with paired bootstrap confidence intervals; and generators for every table and figure. An adversarial audit from eight native-fluency perspectives replaced four homophone pairs that did not truly collapse. The main run is waiting on a go/no-go decision, so there are no multilingual results yet — and none are claimed.
Planned
- Tasks · languages · variants
- 44 · 8 · 264
Triangulated audit: three minds and a calculator
A reusable audit skill that attacks a research package from three angles over several rounds — a multi-agent workflow with access to the repository and tools, a strong external model, and a numerical and empirical arbiter — and settles every disagreement between models by computation rather than authority. The audit that motivated it showed why: the strongest external model got a formula and a threshold wrong that the other participants got right; only the union plus the arbiter was complete; and a mistaken refutation forced a certification that made the result stronger.
Shipped
What the next version will measure
What comes next: position the work against related research; match the title to the zh-CN evidence; soften the design rule to fit the small homophone samples; run a small typing study, so that both sides of the effort–quality trade-off are measured rather than modelled; build about thirty homophone tasks with graded ambiguity to trace a real performance curve; pin every model version; and look harder at why ‘normalise first’ hurts. The eight-language harness is ready for the broader claim.
Planned
Accepted: the first AI paper
On 11 August 2026 UKAIRS accepted the paper for a poster session in the Emerging Research track. The reviews found the question genuinely interesting, the artefact useful-looking and the inclusivity angle worth pursuing, and recommended leading with the ‘shi shi’ example and the three-layer analysis — encoding change, compression, tone loss. It is the first peer-reviewed AI paper, first-authored.
Accepted
Poster #181, Wednesday 25 November
The organisers assigned the poster to Wednesday 25 November 2026, 13:15–14:15, in Lunch & Poster Session 3, and sent the presentation guidance. The public poster list shows it as #181 under ‘F. Synergistic Human–AI Collaboration’, Emerging Research, Xiaoxiao Zhouyi, University of Bristol.
Scheduled
UKAIRS 2026, Edinburgh
The UK AI Research Symposium runs on 24–25 November 2026 at the John McIntyre Conference Centre in Edinburgh, and the poster is up on Wednesday 25 November. The plan, written into the paper: a side-by-side demo showing, for each task, the native prompt, its keyboard-friendly variant, the agent's response, the grade and the effort estimate — so that visitors can toggle from routine prompts to exact-term stress prompts and watch where meaning collapses. Chapter 02 of this site is a first draft of that demo.
Scheduled · 25 Nov 2026