Publications

Papers and preprints on model evaluation, agent-managed software, rhythm-game generation, and speech technology.

Google Scholar

Citation counts checked September 29, 2026.

2026

  1. Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison

    Jhen-Ke Lin, Hong-Yun Lin

    arXiv:2608.30044 · Preprint

    A framework for comparing models using public benchmark scores while accounting for repeated capabilities, benchmark difficulty, and the target task.

    Read paper 0 citations
  2. Grading Needs a Rubric, Not Intelligence

    Jhen-Ke Lin

    arXiv:2608.17938 · Preprint

    Studies when lower-cost language models can reliably grade open-ended exam answers using explicit rubrics.

    Read paper 0 citations
  3. ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation

    Jhen-Ke Lin

    arXiv:2607.12857 · Preprint

    Evaluates generated rhythm-game charts through separate timing and chart-quality signals without requiring reproduction of an official note sequence.

    Read paper 0 citations
  4. BUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases

    Jhen-Ke Lin

    arXiv:2605.06136 · Preprint

    Measures whether downstream agents can recover intended behavior and design decisions from generated repositories, and how much inspection that requires.

    Read paper 0 citations
  5. Session-Level Spoken Language Assessment with a Multimodal Foundation Model via Multi-Target Learning

    Hong-Yun Lin, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen

    ICASSP 2026

    Assesses an entire spoken response session with a multimodal foundation model, jointly learning holistic and trait-level proficiency objectives.

    Read paper 2 citations

2025

  1. The NTNU System at the S&I Challenge 2025 SLA Open Track

    Hong-Yun Lin, Tien-Hong Lo, Yu-Hsuan Fang, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen

    SLaTE 2025

    Combines wav2vec 2.0 and Phi-4 multimodal predictions through score fusion; the system placed second in the Speak & Improve 2025 SLA open track.

    Read paper 5 citations
  2. Advancing Automated Speaking Assessment Leveraging Multifaceted Relevance and Grammar Information

    Hao-Chien Lu, Jhen-Ke Lin, Hong-Yun Lin, Chung-Chun Wang, Berlin Chen

    SLaTE 2025

    Integrates question, image, exemplar, and response relevance with detailed grammar-error features to improve content, language-use, and overall speaking scores.

    Read paper 1 citation
  3. A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions

    Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin, Berlin Chen

    SLaTE 2025

    Addresses scarce labeled speaking-assessment data using proficiency-conditioned text generation and speaker-aware speech synthesis.

    Read paper 1 citation
  4. Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems

    Jhen-Ke Lin, Hao-Chien Lu, Chung-Chun Wang, Hong-Yun Lin, Berlin Chen

    SLaTE 2025

    Compares removed, generic, and acoustically precise hesitation annotations when fine-tuning Whisper for verbatim L2 speech transcription.

    Read paper 4 citations

2024

  1. Development of an English Oral Assessment System with the GEPT Dataset

    Hao-Chien Lu, Chung-Chun Wang, Jhen-Ke Lin, Berlin Chen

    O-COCOSDA 2024

    An oral assessment system for the GEPT Intermediate picture-description task that considers acoustic features, content, and language use.

    Read paper 1 citation

2023

  1. The NTNU ASR System for Formosa Speech Recognition Challenge 2023

    Hao-Chien Lu, Chung-Chun Wang, Jhen-Ke Lin, Tien-Hong Lo

    ROCLING 2023

    Describes the NTNU automatic speech recognition system submitted to the Formosa Speech Recognition Challenge 2023.

    Read paper 4 citations