The application discloses a multi-
modal hate speech detection method based on rational enhancement dynamic anchor points, relates to the technical field of multi-
modal content detection, and comprises the following steps: collecting multi-
modal social platform data streams, extracting text and visual modal segments, and constructing a sample set; generating event triples by
semantic role labeling and
parsing the text, and establishing a cross-modal entity alignment mapping table in combination with the position of the visual
face detection frame; extracting a text deep
semantic vector and a visual area
feature vector, and fusing to generate a multi-modal joint representation
tensor; calculating the attention weight between
modes by a dynamic
anchor point generator, determining a rational enhancement dynamic
anchor point set, and re-weighting the joint representation
tensor to obtain a corrected
feature matrix; and inputting the
feature matrix into a pre-trained classification decision forest to output a hate speech detection result. The method realizes accurate alignment of cross-modal entities, dynamically optimizes feature weights, and improves the multi-modal hate speech detection effect.