AI PRODUCT · OPEN SOURCE · SPEECH
I turn hard AI problems into tools people can use.
I’m Jayden, an AI product builder and independent researcher working across agent workflows, speech evaluation, and practical AI products.
01 / Selected work
Products, tools, and open research.
Each project starts with a concrete problem and ends with something others can inspect, use, or reproduce.
CN-NewsTTS Bench
A target-level benchmark for evaluating whether Chinese news TTS products pronounce difficult written forms correctly from raw input.
MeowFinder
A multilingual, scenario-based guide that helps people find hiding cats using familiar sounds and evidence-informed search steps.
02 / Research thread
From measuring pronunciation to questioning the measurement.
Two releases form one research arc: build a scalable benchmark, then investigate a blind spot in the automated evaluation itself.
Production problem
Written news forms change spoken meaning.
Scores, model names, units, abbreviations, and mixed strings are common—and easy for TTS systems to misread.
Open benchmark
Evaluate targets from raw input.
CN-NewsTTS Bench releases a fixed public test, multi-ASR transcripts, scoring code, initial results, and a leaderboard.
Evaluation audit
ASR can “correct” what listeners hear.
Human audits and cross-model controls reveal false negatives that automated round-trip checks can conceal.
03 / More tools
Small, focused systems.
Experiments and utilities that turn recurring friction into a repeatable workflow.

Claude-Codex Bridge
A local, CLI-first bridge for bidirectional Claude Code and Codex review loops.
View on GitHub04 / About
Product judgment, backed by working systems and inspectable evidence.
I work at the boundary between AI product design and technical investigation. My projects usually begin with an ambiguous real-world failure—then become a tool, benchmark, dataset, or public product that makes the problem easier to understand and act on.
I’m especially interested in evaluation, agent workflows, speech products, and the details that determine whether an AI system is genuinely useful outside a demo.