I am Jacob

What do you want to talk about?

Enter to send

Enter to send, Shift+Enter for newline

Research Interests

How I Think About Collaboration in the Agent Era

Current questions spanning multi-human steering, agent-driven engineering, and faster software evolution.

Multi-Human Steering for Agent Systems

The Question

How should teams coordinate around a shared agent system when planning, analysis, and execution all happen faster than any single person can fully review?

Why This Matters

One person should not carry the full cost of understanding, auditing, and steering a high-speed agent. I want collaboration models where people guide the system from different areas of expertise instead of chasing it alone.

What I'm Testing Now

I am studying interaction models where multiple people provide feedback, constraints, and reinforcement through a unified agent interface so teams can stay aligned without forcing everyone to inspect everything.

Interfaces for Agent-Driven Software Engineering

The Question

What new development environments, representations, and visualizations do we need when agents can read, modify, and generate code faster than humans can track?

Why This Matters

Traditional code editors assume humans inspect changes directly, while pure CLI workflows compress away too much context. Neither is enough when agent activity becomes dense, parallel, and continuous.

What I'm Testing Now

I am exploring workflow designs that make agent behavior observable, reviewable, and steerable so humans can keep judgment and control inside highly automated engineering loops.

Shortening the Path from User Need to Software Change

The Question

How can software evolve much faster in agent-mediated product teams without losing safety, controllability, or human understanding?

Why This Matters

As agents get better at turning intent into concrete changes, the distance between end users and software updates can shrink dramatically. We need new processes that preserve oversight while letting products adapt faster.

What I'm Testing Now

I am thinking about frameworks, governance mechanisms, and transition paths that help teams move into agent-driven work while redirecting human effort toward deeper reasoning, creativity, and cross-domain problem solving.

Publications

Research, in Papers

Recent work on model evaluation and agent-managed software, alongside research in rhythm games and speech technology.

View all publications
  1. 2026 0 citations Model EvaluationBenchmarks

    Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison

    Jhen-Ke Lin, Hong-Yun Lin

    arXiv:2608.30044 · Preprint

    Contribution

    Combines semantic density weighting, score equating, and task-relevant residual pooling; evaluates model comparisons on Artificial Analysis benchmarks and WildScores.

    Summary

    A framework for comparing models using public benchmark scores while accounting for repeated capabilities, benchmark difficulty, and the target task.

    Open publication
  2. 2026 0 citations LLM EvaluationGrading

    Grading Needs a Rubric, Not Intelligence

    Jhen-Ke Lin

    arXiv:2608.17938 · Preprint

    Contribution

    Separates one-time rubric extraction from repeated grading, and uses rubric ablations to examine the role of the official answer in grading reliability.

    Summary

    Studies when lower-cost language models can reliably grade open-ended exam answers using explicit rubrics.

    Open publication
  3. 2026 0 citations Rhythm GamesEvaluation

    ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation

    Jhen-Ke Lin

    arXiv:2607.12857 · Preprint

    Contribution

    Tests evaluation signals against controlled corruptions and reports a quality profile instead of collapsing distinct properties into one total score.

    Summary

    Evaluates generated rhythm-game charts through separate timing and chart-quality signals without requiring reproduction of an official note sequence.

    Open publication
  4. 2026 0 citations AI AgentsSoftware Engineering

    BUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases

    Jhen-Ke Lin

    arXiv:2605.06136 · Preprint

    Contribution

    Separates behavioral correctness from intent recovery, reporting accuracy, repeatability, implementation coverage, and inspection effort.

    Summary

    Measures whether downstream agents can recover intended behavior and design decisions from generated repositories, and how much inspection that requires.

    Open publication
  5. 2026 2 citations Speaking AssessmentMultimodal

    Session-Level Spoken Language Assessment with a Multimodal Foundation Model via Multi-Target Learning

    Hong-Yun Lin, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen

    ICASSP 2026

    Contribution

    Uses multi-target learning and a frozen Whisper speech prior for acoustic-aware calibration, evaluated on the Speak & Improve benchmark.

    Summary

    Assesses an entire spoken response session with a multimodal foundation model, jointly learning holistic and trait-level proficiency objectives.

    Open publication

Side Projects

Open Source and Side Projects

A mix of developer tools, ML experiments, and open-source utilities I have built over time. Some overlap with my research interests; many simply solve practical problems I care about.

Followers

Loading live metrics...

Public repositories

Loading live metrics...

Total Stars

Loading live metrics...

Most starred repository

Loading live metrics...

Svelte 613 stars

d1-manager

A Cloudflare D1 workbench for browsing schema, editing rows, running SQL, and turning natural-language questions into executable queries.

TypeScript 902 stars

LeetCode-Stats-Card

A Cloudflare-hosted card generator that turns LeetCode activity into customizable, embeddable visuals for GitHub profiles and personal sites.

TypeScript 43 stars

selflare

A CLI that compiles Cloudflare Workers into Cap'n Proto configs and ships them as minimal Docker deployments while preserving common platform bindings.

Rust 29 stars

gradio-rs

A Rust client and CLI for Gradio apps with API discovery, file uploads, queued jobs, and both blocking and async inference.

Rust 11 stars

rhythm-rs

A Rust rhythm-game workspace that separates engine timing, Taiko rules, chart parsing, and a playable game into reusable crates.

Updating live metrics...