Evaluating NLP Bias with Perturbation Sensitivity Analysis
Evaluating NLP Bias with Perturbation Sensitivity Analysis
Created using ChatSlide
This research investigates the implications of text bias in NLP models, highlighting how unintended social associations can lead to significant harms in moderation systems. Through Perturbation Sensitivity Analysis (PSA), we analyze the impact of replacing entity anchors on toxicity and sentiment classifications across various online comment genres. Our results reveal that sentiment is more sensitive to name changes, with label alterations affecting 10-40% of cases. The findings underscore...