Publications
Papers and preprints on model evaluation, agent-managed software, rhythm-game generation, and speech technology.
Google ScholarCitation counts checked September 29, 2026.
2026
Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison
Jhen-Ke Lin, Hong-Yun Lin
arXiv:2608.30044 · Preprint
A framework for comparing models using public benchmark scores while accounting for repeated capabilities, benchmark difficulty, and the target task.
Read paper 0 citationsGrading Needs a Rubric, Not Intelligence
Jhen-Ke Lin
arXiv:2608.17938 · Preprint
Studies when lower-cost language models can reliably grade open-ended exam answers using explicit rubrics.
Read paper 0 citationsChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
Jhen-Ke Lin
arXiv:2607.12857 · Preprint
Evaluates generated rhythm-game charts through separate timing and chart-quality signals without requiring reproduction of an official note sequence.
Read paper 0 citationsBUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases
Jhen-Ke Lin
arXiv:2605.06136 · Preprint
Measures whether downstream agents can recover intended behavior and design decisions from generated repositories, and how much inspection that requires.
Read paper 0 citationsSession-Level Spoken Language Assessment with a Multimodal Foundation Model via Multi-Target Learning
Hong-Yun Lin, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen
ICASSP 2026
Assesses an entire spoken response session with a multimodal foundation model, jointly learning holistic and trait-level proficiency objectives.
Read paper 2 citations
2025
The NTNU System at the S&I Challenge 2025 SLA Open Track
Hong-Yun Lin, Tien-Hong Lo, Yu-Hsuan Fang, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen
SLaTE 2025
Combines wav2vec 2.0 and Phi-4 multimodal predictions through score fusion; the system placed second in the Speak & Improve 2025 SLA open track.
Read paper 5 citationsAdvancing Automated Speaking Assessment Leveraging Multifaceted Relevance and Grammar Information
Hao-Chien Lu, Jhen-Ke Lin, Hong-Yun Lin, Chung-Chun Wang, Berlin Chen
SLaTE 2025
Integrates question, image, exemplar, and response relevance with detailed grammar-error features to improve content, language-use, and overall speaking scores.
Read paper 1 citationA Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions
Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin, Berlin Chen
SLaTE 2025
Addresses scarce labeled speaking-assessment data using proficiency-conditioned text generation and speaker-aware speech synthesis.
Read paper 1 citationAcoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems
Jhen-Ke Lin, Hao-Chien Lu, Chung-Chun Wang, Hong-Yun Lin, Berlin Chen
SLaTE 2025
Compares removed, generic, and acoustically precise hesitation annotations when fine-tuning Whisper for verbatim L2 speech transcription.
Read paper 4 citations
2024
Development of an English Oral Assessment System with the GEPT Dataset
Hao-Chien Lu, Chung-Chun Wang, Jhen-Ke Lin, Berlin Chen
O-COCOSDA 2024
An oral assessment system for the GEPT Intermediate picture-description task that considers acoustic features, content, and language use.
Read paper 1 citation
2023
The NTNU ASR System for Formosa Speech Recognition Challenge 2023
Hao-Chien Lu, Chung-Chun Wang, Jhen-Ke Lin, Tien-Hong Lo
ROCLING 2023
Describes the NTNU automatic speech recognition system submitted to the Formosa Speech Recognition Challenge 2023.
Read paper 4 citations