Hybrid LLM-Based Framework for Sentence-Level Evaluation of LLM Text Responses
MSc Thesis Defense by: Nehme Haidura
Date: Thursday, September 10, 2026
Time: 2:00pm to 3:00pm
Location: Essex Hall 122
Abstract:
The rapid growth of large language models (LLMs) has made it difficult to evaluate both model performance and the quality of generated text. Domain specific LLM (DSLLM) research varies widely in architecture, task, domain, and metric choice. This makes it difficult to determine how these factors influence reported performance, and to draw consistent conclusions across studies. At the same time, automatic text evaluation methods often capture only selected dimensions of response quality rather than assessing several complementary properties within a unified evaluation process. In this thesis, we conduct two studies to address these problems.
The first study systematically reviews around 60 DSLLM papers to examine how architecture, domain adaptation, task, and evaluation method jointly shape model performance. Drawing on findings across the reviewed studies, this study identifies distinct capability patterns across encoder-only, decoder-only, and encoder-decoder architectures, and shows that domain specialization alone does not guarantee superior performance. The study finds that evaluation practices across the field are fragmented and rarely account for evolving or outdated knowledge, which limits the field’s ability to draw consistent conclusions. These findings highlight broader limitations in current LLM evaluation and motivate the multidimensional generated-text evaluation investigated in the second study.
To address this gap, the second study introduces HLFEval, a hybrid LLM based framework for sentence-level evaluation of generated text that combines semantic relevance, fluency, safety, and factuality into a single score using learned metric weights, with semantic relevance assessed through both token-level linguistic analysis and LLM-generated similarity scores. Across 12 non-overlapping evaluation batches, HLFEval shows stable learning behaviour and consistent performance, and matches or outperforms established evaluation methods.
Together, these studies contribute a clearer understanding of what drives DSLLM performance and a more comprehensive, multidimensional framework for evaluating generated text, addressing two central gaps in how LLMs are currently assessed.
Keywords: Large Language Models (LLMs), Domain-Specific LLMs, Automatic Text Evaluation, Sentence-Level Evaluation, Evaluation Metrics
Thesis Committee:
Internal Reader: Dr. Jianguo Lu
External Reader: Dr. Esam Abdel-Raheem
Advisor: Dr. Ziad Kobti
Co-Supervisor: Dr. Hussein Assaf
Chair: Dr. Muhammad Asaduzzaman
