The invention discloses an uncertainty guided few-sample harmful speech detection method (U-GIFT). According to the method, a pre-training
language model is finely adjusted based on a small number of labeled samples, and a semi-supervised self-training and uncertainty guiding strategy is combined. Monte Carlo Dropout is started in the reasoning stage, multiple times of random
forward propagation are carried out to obtain sample posterior distribution, prediction entropy and
information gain are calculated, pseudo-
label samples are sorted and screened, and only high-confidence samples are selected to be added into a
training set. And in order to reduce the influence of a pseudo labeling error, designing a stability weighting mechanism, giving a
sample weight according to a prediction variance, and constructing a joint
loss function, so that the model preferentially learns a stable sample to improve the
detection performance. According to the method, the semantic and attention mechanism of the pre-training model is utilized, the detection effect is remarkably improved under the conditions of few samples, imbalance, multiple languages and cross domains, models such as BERT, RoBERTa, XLM-R, LLaMA2 and DeepSeek-R1 are compatible, and the method is suitable for content auditing and
risk prevention and control.