MSc Thesis Defense: A Systematic Vulnerability Assessment For LLM Guardrails by Ahmed Abdalmagid

Tuesday, October 13, 2026 - 12:00

A Systematic Vulnerability Assessment For LLM Guardrails

MSc Thesis Defense by: Ahmed Abdalmagid

Date: Tuesday, October 13th, 2026

Time:  12:00pm to 2:00pm

Location:  ED1121 (Education Bldg)

 

Abstract:

Large language models (LLMs) are increasingly deployed across a wide range of applications, including security and safety-critical domains where model failures or adversarial manipulation can have significant consequences. To mitigate these risks, researchers have proposed both internal post-training safety mechanisms and external guardrails to enforce policies and prevent unsafe or undesired model behavior. Existing studies have evaluated guardrail models against specific datasets and adversarial attacks and have identified individual vulnerabilities observed during training or inference. However, the literature lacks a systematically derived vulnerability taxonomy for LLM guardrails, limiting our ability to characterize their attack surface, compare guardrail mechanisms, and evaluate their security in a structured manner.

In this work, we study LLM guardrails as policy enforcement mechanisms and introduce a systematic vulnerability taxonomy derived through a three-stage mapping procedure that translates established vulnerability concepts from network security appliances into the LLM guardrail domain. We use the resulting taxonomy to guide the security evaluation of representative guardrail models under adversarial attacks. Our experimental results show that Granite Guardian 3.2-3B consistently outperforms the other evaluated LLM-based guardrails against adversarial attacks. We also find that a lightweight non-LLM classifier is competitive with, and in several cases outperforms, LLM-based guardrails, highlighting that increased model complexity does not necessarily translate into greater robustness.

These findings suggest studying LLM guardrail security through systematic vulnerability analysis rather than primarily through attack- or dataset-specific evaluations. The proposed taxonomy provides a foundation for more structured and comparable guardrail security assessments and can be extended as new vulnerability classes and attack mechanisms emerge.

 

Keywords: Guardrails, LLMs , Vulnerability assessment

 

Thesis Committee:
Internal Reader: Dr. Alioune Ngom
Internal Reader: Dr. Jianguo Lu
Advisor: Dr.  Sherif Saad
Chair: Dr.Ikjot Saini