MSc Thesis Proposal: Combating the Anisotropic Problem: Fine-tuning Sentence Embedding Using Jensen-Shannon Divergence by Shiqi Zhang

Monday, September 21, 2026 - 10:00

Combating the Anisotropic Problem: Fine-tuning Sentence Embedding Using Jensen-Shannon Divergence

MSc Thesis Proposal by: Shiqi Zhang

Date: Sept 21, 2026

Time: 10:00

Location: Essex Hall 122

Abstract:

Sentence embeddings are dense vector representations of sentences, and developing high-quality sentence embeddings plays a central role in both LLM research and downstream applications. Most generative AI systems rely on retrieval-augmented generation (RAG), which depends on sentence embeddings, and embeddings are the cornerstone of tasks such as semantic search, classification, and clustering. Embeddings from pretrained LLMs typically need fine-tuning using human-labeled similarity pairs. This is complicated by anisotropy: pretrained sentence representations occupy a narrow region of embedding space, so nearly all sentence pairs exhibit high cosine similarity (e.g., above 0.9), even pairs humans label as unrelated — a substantial mismatch between the similarity distribution embeddings naturally produce and the distribution implied by human labels. Loss functions such as Mean Squared Error (MSE) address this by regressing directly toward the labeled value, forcing disproportionately large, disruptive changes to the embedding geometry to reach targets far outside its natural range. Combating this mismatch, however, does not necessarily mean eliminating anisotropy itself. We propose two loss functions that instead correct for it directly: a mean-std-adjusted version of MSE, which rescales training targets to align with the pretrained model's own output distribution rather than an arbitrary external scale; and Batch JSD, which min-max normalizes predictions and labels within each batch and minimizes their Jensen-Shannon divergence as probability distributions, retaining relative-magnitude information that purely ordinal objectives do not encode. Experiments on BERT-base and RoBERTa-base across seven STS benchmarks show that our loss functions outperform the state of the art embedding methods, including CoSENT and AnglE. In addition to the STS evaluation, we also explore its impact on the geometric shape of the resulting embeddings.  

Keywords: [Insert 3-5 keywords] Sentence embeddings, fine-tuning , loss function, anisotropy, Jensen-Shannon Divergence

Thesis Committee:

Internal Reader: Dr. Dan Wu

External Reader: Dr. Abdul A. Hussein

Advisor: Dr. Jianguo Lu

Vector Institute Logo