A discrimination detection method for multi-modal multi-language information
Patent Information
- Application Number
- CN202510137826.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-02-07
AI Technical Summary
[0005]此外,传统的单模态方法在应对多模态内容(如图像与文本的结合)时,往往因无法充分捕捉模态间的交互关系而表现出较低的检测准确性
(1)本申请通过采用ViT图像编码器和XLM-R文本编码器,显著提升了对图像和多语言文本信息的特征提取能力,能够更好地捕捉视觉和语言模态中的关键信息。交叉注意力机制的引入有效解决了图像与文本模态之间的细粒度交互信息难以捕捉的问题,增强了模态间的语义关联,使得融合特征更加精准和具有深度语义表达能力。动态记忆机制的结合进一步强化了模型对长距离依赖和复杂语境细节的捕捉能力,有效提高了对复杂歧视信号的识别准确性。
Smart Images

Figure CN120144794B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of cybersecurity and information content security, and in particular to a discrimination detection method for multimodal and multilingual information. Background Technology
[0002] With the rapid development of the internet and social media, the forms of information dissemination have become increasingly diversified, exhibiting a multimodal characteristic combining text, images, and videos. However, the convenience of this multimodal information dissemination is also accompanied by potential negative impacts, especially the covert spread of discriminatory content, which poses a significant threat to individual mental health and social harmony. Social media platforms (such as Twitter, Facebook, and Instagram) have become an important part of people's lives, but discriminatory speech spreads rapidly through these platforms, creating negative cross-linguistic and cross-cultural impacts. Detecting and identifying such content has become a significant challenge in the fields of cybersecurity and information management.
[0003] Currently, most discriminatory speech detection methods still primarily focus on the analysis of single-modal text data. Traditional methods include machine learning-based models such as Naive Bayes, Support Vector Machines, and Random Forests. These methods extract text features (such as bag-of-words models, TF-IDF, and n-grams) and use classifiers to identify discriminatory content. However, these methods typically rely on manual feature extraction and cannot effectively capture deep semantic relationships within the text. With the rise of deep learning technology, models based on Recurrent Neural Networks (RNNs) and Long Short-Term Memory Networks (LSTMs) have gradually replaced traditional algorithms. These models can better capture the contextual semantics and sequence dependencies of text, thus significantly improving the performance of discriminatory speech detection. Subsequently, large-scale pre-trained language models, represented by BERT, have further advanced discrimination detection. These models can learn richer semantic information through pre-training on large-scale corpora, thereby improving their ability to understand complex text content.
[0004] In recent years, multimodal discrimination detection has gradually become a research hotspot. Multimodal methods combine text and image information sources, extracting image features through convolutional neural networks (CNNs) and processing text information using LSTM or Transformer architectures, significantly improving the performance of discrimination detection. For example, some studies have successfully detected covert discriminatory content containing both images and text in social media content. However, these studies mainly focus on monolingual environments (especially English), lacking effective identification capabilities for discriminatory speech in multilingual and cross-cultural contexts. In practical applications, due to the diversity of language and cultural backgrounds, discriminatory content often exhibits different metaphorical, contextual, and symbolic expressions, posing higher demands on existing detection methods.
[0005] Furthermore, traditional unimodal methods often exhibit lower detection accuracy when dealing with multimodal content (such as combinations of images and text) because they fail to adequately capture the interactions between modalities. Although Transformer-based multimodal models have made some progress in recent years, existing methods still fall short in capturing fine-grained interaction information between modalities. Simultaneously, due to the multilingual nature of information on social media platforms, existing methods lack support for multilingual scenarios, particularly exhibiting weak ability to identify discriminatory remarks in non-English languages, thus limiting their application in a globalized environment. Summary of the Invention
[0006] This application aims to at least partially address one of the technical problems in the related art.
[0007] Therefore, the first objective of this application is to propose a discrimination detection method for multimodal and multilingual information.
[0008] The second objective of this application is to propose a discrimination detection device for multimodal and multilingual information.
[0009] To achieve the above objectives, the first aspect of this application proposes a discrimination detection method for multimodal and multilingual information, comprising: The original image and text data are preprocessed. The preprocessing process includes word segmentation and stop word removal of the text data, and resizing and normalization of the image data. The preprocessed text data is input into the text encoder to extract the context embedding features of the text, and the preprocessed image data is input into the visual encoder to extract the global features and local patch features of the image. By utilizing a cross-attention mechanism to perform multimodal alignment and fusion of text and image features, and by capturing the correlation between image and text modalities, interactive features containing cross-modal semantic information are generated. The original modal features and interactive features are then fused to generate fused features with multimodal contextual information. The image encoder and text encoder parts in the pre-trained model are frozen, and only the classifier part is fine-tuned with low-rank parameters. By adding low-rank matrix adjustment terms, the classifier's ability to classify multimodal fusion features is optimized. Using a dynamic memory mechanism, the similarity between the fusion features of a new input sample and the features of historical samples in the memory bank is calculated and a weighted aggregated feature is generated. The fusion features of the current sample are then dynamically fused with the weighted aggregated feature. The final fused representation is classified by a classifier, and the result is output as either discriminatory or non-discriminatory.
[0010] Optionally, the step of inputting the preprocessed text data into a text encoder to extract the context embedding features of the text, and inputting the preprocessed image data into a visual encoder to extract the global features and local patch features of the image, includes: Preprocessed text data Word embedding is performed, mapping each word to a high-dimensional vector, and contextual embedding features of the text are generated using an XLM-R text encoder. The formula is:
[0011] In the formula, , The length of the text sequence. For the dimension of the embedded features, For the first Contextual representation of each word; The preprocessed image data is segmented to divide the image... Divided into A fixed-size patch, denoted as Linear mapping and positional encoding are performed on each patch to obtain the image features. ; The image features are input into the ViT image encoder, and a self-attention mechanism is used to perform multi-layer feature extraction to generate global features of the image. The result is expressed as:
[0012] In the formula, , The dimension of image features. These are global features of the image.
[0013] Optionally, the method of using a cross-attention mechanism to perform multimodal alignment and fusion of text features and image features, by capturing the correlation between image and text modalities, generates interactive features containing cross-modal semantic information, and fuses the original modal features with the interactive features to generate fused features with multimodal contextual information, including: Through image features Guide text features Generate image-guided text features ; Through text features Generate guiding image features Generate text-guided image features ; Image features Text features Image-guided text features Image features guided by text The features are concatenated to generate fused features with multimodal contextual information. , is represented as:
[0014] In the formula, This indicates a feature splicing operation.
[0015] Optionally, the method using image features Guide text features Generate image-guided text features ,include: Utilizing image features Form a query matrix Utilizing text features Generate key matrix Sum matrix Specifically, it is expressed as: , ,
[0016] in, , , The weight matrix is a learnable matrix; The relevance weight of the image to the text is calculated using the dot product attention formula, expressed as:
[0017] Image-guided text features are generated based on relevance weights. , is represented as:
[0018] In the formula, Text features guided by images.
[0019] Optionally, the text features Generate guiding image features Generate text-guided image features ,include: Utilizing text features Form a query matrix Utilizing image features Generate key matrix Sum matrix Specifically, it is expressed as: , ,
[0020] in, , , The weight matrix is a learnable matrix; The relevance weight of text to image is calculated using the dot product attention formula, expressed as:
[0021] Generate text-guided image features based on relevance weights. , is represented as:
[0022] In the formula, Image features guided by text.
[0023] Optionally, freezing the image encoder and text encoder parts in the pre-trained model and only fine-tuning the classifier part with low-rank parameters, by adding low-rank matrix adjustment terms, optimizes the classifier's ability to classify multimodal fusion features, including: Initialize the classifier weight matrix low-rank matrix adjustment term and ,in: ,
[0024] Freeze the weights of the image encoder and text encoder to ensure that the parameters of the pre-trained model remain unchanged during fine-tuning, only adjusting the low-rank matrix of the classifier weights. and Update; Multimodal fusion features of input Prediction is performed using a classifier with an added low-rank matrix adjustment term. The classification result is represented as follows:
[0025] in, The weights of the original classifier. For bias terms; The classification loss is calculated using the cross-entropy loss function, and the formula for the loss function is:
[0026] in, For the real category, To predict class probabilities; Using the backpropagation algorithm, only low-rank matrices are processed. and Perform parameter updates, including updating the adjustment terms of the low-rank matrix.
[0027] Optionally, the step of utilizing a dynamic memory mechanism to calculate the similarity between the fusion features of a new input sample and the features of historical samples in the memory bank, and generating weighted aggregate features, and then dynamically fusing the fusion features of the current sample with the weighted aggregate features, includes: During training, a memory is built to store the features and corresponding labels of historical samples, represented as:
[0028] In the formula, Historical Sample Feature representation, Historical Sample The tag, To predict the entropy value; Feature extraction is performed on the newly input image and text data to obtain the fused features of the current sample, using the following formula:
[0029] In the formula, This represents the multimodal fusion features of the input; Use current sample features Retrieve from the memory and calculate the fusion features of the current sample using the cosine similarity formula. Find the most similar sample by comparing its features with those of each historical sample in the memory. Features of historical samples , represented as:
[0030]
[0031] In the formula, This is the function for calculating cosine similarity. Found in the memory bank Features of historical samples Perform weighted aggregation to generate weighted aggregated features. , represented as:
[0032] In the formula, the weights Calculated based on similarity distribution; The current sample's features are fused using a gated loop unit. With weighted aggregation features Perform dynamic fusion and output the final fused representation. .
[0033] Optionally, the current sample is fused features through the gated loop unit. With weighted aggregation features Perform dynamic fusion and output the final fused representation. ,include: Computational update gate Used to control the current input features Hidden state at the previous time step The weighting ratio in this time step is calculated using the following formula:
[0034]
[0035] In the formula, This involves concatenating the current sample's fused features with the weighted aggregated features. This is the hidden state from the previous time step. and The weight matrix is a learnable matrix. For bias terms, Use the Sigmoid activation function; Calculate the reset door Used to determine historical information The retention rate is determined by the formula:
[0036] In the formula, and The weight matrix is a learnable matrix. For bias terms; Generate candidate states This is used to combine the current input features with adjusted historical information, and the formula is:
[0037] in, , The weight matrix is a learnable matrix. For bias terms, This indicates the historical state after the gate adjustment, and the candidate state. It is a potential hidden state in the current time step; Compute the final fused representation By updating the gate to dynamically balance the current input and historical information, the formula is:
[0038] in, The contribution representing the hidden state of history, This indicates the contribution of the currently input information.
[0039] Optionally, the step of classifying the final fused representation using a classifier and outputting a discriminatory or non-discriminatory judgment result includes: The final fusion representation The input is fed into the classifier to generate the final prediction result, using the following formula:
[0040] in, As the weight of the classification head, The final prediction result represents the probability that the input sample belongs to the discriminatory or non-discriminatory category.
[0041] To achieve the above objectives, a second aspect of this application provides a discrimination detection device for multimodal and multilingual information, comprising: The data preprocessing module is used to preprocess the original image and text data. The preprocessing process includes word segmentation and stop word removal of the text data, and resizing and normalization of the image data. The feature extraction module is used to input preprocessed text data into the text encoder to extract the context embedding features of the text, and to input preprocessed image data into the visual encoder to extract the global features and local patch features of the image. The multimodal alignment and fusion module is used to perform multimodal alignment and fusion of text features and image features using a cross-attention mechanism. By capturing the correlation between image and text modalities, it generates interactive features containing cross-modal semantic information, and fuses the original modal features with the interactive features to generate fused features with multimodal contextual information. The classifier fine-tuning module is used to freeze the image encoder and text encoder parts in the pre-trained model and only perform low-rank parameter fine-tuning on the classifier part. By adding low-rank matrix adjustment terms, the classifier's ability to classify multimodal fusion features is optimized. The dynamic memory fusion module is used to calculate the similarity between the fusion features of a new input sample and the features of historical samples in the memory bank using a dynamic memory mechanism, and generate weighted aggregate features. The fusion features of the current sample are then dynamically fused with the weighted aggregate features. The classification and judgment module is used to classify the final fused representation using a classifier and output discriminatory or non-discriminatory judgment results.
[0042] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects: (1) This application significantly improves the feature extraction capability of images and multilingual text information by employing the ViT image encoder and the XLM-R text encoder, enabling better capture of key information in visual and linguistic modalities. The introduction of the cross-attention mechanism effectively solves the problem of capturing fine-grained interaction information between image and text modalities, enhances the semantic association between modalities, and makes the fused features more accurate and have deep semantic expression capabilities. The combination of the dynamic memory mechanism further strengthens the model's ability to capture long-distance dependencies and complex contextual details, effectively improving the recognition accuracy of complex discrimination signals.
[0043] (2) The method proposed in this application overcomes the limitations of existing discrimination detection methods in multilingual environments. Through a multilingual processing mechanism, it effectively identifies discriminatory remarks from different linguistic and cultural backgrounds, filling the technical gap in multimodal discrimination detection in multilingual scenarios. By integrating images and text in a multimodal manner, this method fully utilizes the complementary advantages of vision and language, effectively improving the detection capability of potential discriminatory content related to gender, race, and culture.
[0044] (3) This application introduces LoRA fine-tuning technology, which adjusts the parameters of the classifier only, thereby reducing the computational resource requirements while retaining the knowledge of the pre-trained model and ensuring a balance between performance and resource consumption. This optimized design not only improves detection efficiency but also makes the method more practical and valuable for promotion. It can be widely applied to content moderation and information security management on social platforms, providing an efficient and reliable solution for detecting discriminatory remarks across languages and cultures.
[0045] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0046] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A schematic flowchart illustrating a discrimination detection method for multimodal and multilingual information provided in an embodiment of this application; Figure 2 A schematic flowchart illustrating a discrimination detection method for multimodal and multilingual information provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a discrimination detection device for multimodal and multilingual information provided in an embodiment of this application. Detailed Implementation
[0047] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0048] To address the problems existing in the prior art, this application provides a discrimination detection method for multimodal and multilingual information. Figure 1 and Figure 2 This is a flowchart illustrating a discrimination detection method for multimodal and multilingual information provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps: Step S1 involves preprocessing the original image and text data. The preprocessing process includes segmenting the text data into words and removing stop words, and resizing and normalizing the image data.
[0049] In the embodiments of this application, such as Figure 2 As shown, raw image and text data related to discriminatory content can be extracted through social media platforms. Taking social media networks as an example, users can utilize the platform's open APIs or obtain data through web scraping techniques. By setting specific keywords (such as "racism" or "sexism"), it is possible to collect post content containing these keywords, including text, comments, and associated image data. For example, on Twitter, the tweet body and user comments can be extracted; on Instagram, image captions and user comments can be extracted, along with uploaded images associated with the post. This application does not impose specific limitations on this.
[0050] After obtaining the original image and text data, they are then preprocessed.
[0051] Specifically, the text data undergoes operations such as word segmentation and stop word removal. Word segmentation divides the input text data into word or phrase units, facilitating subsequent feature extraction. Simultaneously, a stop word list is used to remove common but meaningless words (such as "the," "and," and "of"), thereby improving the semantic expressiveness of the text features.
[0052] The image data is resized and normalized. The original input image data is resized to unify the image resolution and ensure the consistency of the model input. At the same time, the image pixel values are normalized to a specific range (such as [0,1]) to reduce the impact of excessively large numerical ranges on model training and improve model convergence efficiency.
[0053] Furthermore, to ensure high-quality and consistent data, this application embodiment further cleans the image and text data, including removing invalid data (such as blank images and text without content) and correcting data anomalies (such as garbled text and corrupted image files). The cleaned data undergoes standardization and format conversion to ultimately generate preprocessed data adapted to the model's input format, providing high-quality input for the subsequent feature extraction stage.
[0054] Step S2: Input the preprocessed text data into the text encoder to extract the context embedding features of the text, and input the preprocessed image data into the visual encoder to extract the global features and local patch features of the image.
[0055] This step extracts high-dimensional feature representations of text and images using a text encoder and an image encoder, respectively, providing a foundation for subsequent multimodal alignment and fusion.
[0056] First, for the preprocessed text data, perform the following feature extraction process: Preprocessed text data Word embedding is performed, mapping each word to a high-dimensional vector, and contextual embedding features of the text are generated using an XLM-R text encoder. Specifically, it is expressed as:
[0057] In the formula, , The length of the text sequence. For the dimension of the embedded features, For the first Contextual representation of each word.
[0058] Through the feature extraction process of the text encoder, the embodiments of this application are able to capture the contextual semantic information of the text in order to generate a text embedding representation with deep semantic features.
[0059] Second, for the preprocessed image data, the following feature extraction process is performed: First, the preprocessed image data is segmented to divide the image... Divided into A fixed-size image patch The system performs linear mapping and positional encoding on each image patch to obtain local patch features of the image. .
[0060] Then, the local patch features of the image are input into the ViT image encoder, and multi-layer feature extraction is performed using a self-attention mechanism. Global features of the image are generated through global encoding by the ViT image encoder. The result is expressed as:
[0061] In the formula, , The dimension of image features. These are global features of the image, representing the overall semantic information of the image.
[0062] This step extracts textual and image features, providing deep semantic information for subsequent multimodal alignment and fusion. Textual features capture semantic contextual relationships, while image features contain both global and local visual features, laying the foundation for the model to understand complex semantic relationships in multimodal data.
[0063] Step S3: Use cross-attention mechanism to perform multimodal alignment and fusion of text features and image features. By capturing the correlation between image and text modalities, generate interactive features containing cross-modal semantic information. Fuse the original modal features with the interactive features to generate fused features with multimodal contextual information.
[0064] In this embodiment, a cross-attention mechanism is used to perform multimodal alignment and fusion of text features and image features. By capturing the semantic correlation between image and text modalities, interactive features containing cross-modal semantic information are generated. The original modal features and the generated interactive features are then fused to form fused features containing multimodal contextual information, providing a more comprehensive semantic representation for subsequent classification tasks.
[0065] Specifically, step S3 also includes the following steps: S31, through image features Guide text features Generate image-guided text features .
[0066] In this embodiment of the application, through image features Semantic information guiding text features This enables deep interaction between modalities.
[0067] Specifically, firstly, image features are utilized. Generate query matrix This indicates that the query content requires guidance from image features and utilizes text features. Generate key matrix Sum matrix Its mathematical expression is as follows: , ,
[0068] in, , , It is a learnable weight matrix used to map input features to a new feature space.
[0069] Next, the relevance weight of the image to the text is calculated using the dot product attention formula. The formula for calculating the weight is as follows:
[0070] Finally, image-guided text features are generated by multiplying the weights with the value matrix. , represented as:
[0071] In the formula, Image-guided text features can capture the performance of text features under the guidance of image semantics, and better reflect the interactive information between images and text.
[0072] S32, through text features Generate guiding image features Generate text-guided image features .
[0073] In this embodiment of the application, text features are used. Semantic information guides image features This further enables information exchange between modalities.
[0074] Specifically, firstly, text features are utilized. Generate query matrix This indicates that the query content requires guidance from text features and utilizes image features. Generate key matrix Sum matrix Its mathematical expression is as follows: , ,
[0075] in, , , The weight matrix is a learnable matrix; Next, the relevance weight of the text to the image is calculated using the dot product attention formula. The formula for calculating the weight is as follows:
[0076] Finally, text-guided image features are generated by multiplying the weights by the value matrix. , represented as:
[0077] In the formula, Text-guided image features can capture the expression of text features under the guidance of image semantics, and better reflect the interactive information between images and text.
[0078] S33, Image Features Text features Image-guided text features and text-guided image features The features are concatenated to generate fused features with multimodal contextual information. , represented as:
[0079] In the formula, This indicates a feature splicing operation.
[0080] Features after fusion Including raw information from images and text, as well as inter-modal interaction information, this provides a more comprehensive and accurate multimodal semantic representation capability for subsequent classification tasks. By mapping these features, the model represents them in a shared semantic space, thus providing a unified multimodal representation. Through this fusion process, the model can better capture the complex semantic relationships in multimodal data and enhance its ability to identify discriminatory content.
[0081] Step S4: Freeze the image encoder and text encoder parts in the pre-trained model, and only fine-tune the classifier part with low-rank parameters. By adding low-rank matrix adjustment terms, optimize the classifier's ability to classify multimodal fusion features.
[0082] This step involves freezing the image encoder and text encoder parts in the pre-trained model and fine-tuning only the classifier part with low-rank parameters to reduce computational resource requirements and optimize the classifier's ability to classify multimodal fusion features.
[0083] In this embodiment of the application, the classifier weight matrix is first... Add low-rank matrix adjustment terms and And initialize the adjustment item to a random value, specifically: ,
[0084] in, and It is a low-rank matrix used to flexibly adjust the weights of the classifier to adapt to classification tasks that fuse multimodal features.
[0085] Next, the weights of the image encoder and text encoder are frozen to ensure that the parameters of the pre-trained model remain unchanged during fine-tuning, with only the low-rank matrix of the classifier being adjusted. and Update accordingly. Freezing preserves the pre-trained model's knowledge of image and text feature extraction while reducing computational overhead.
[0086] Then, the multimodal fusion feature The input classifier performs classification prediction by adding a low-rank matrix adjustment term to the weights, as shown in the following formula:
[0087] in, The weights are those of the original classifier. For bias terms, This represents the category probability distribution of the predicted results.
[0088] To optimize the classification results, this embodiment of the application uses the cross-entropy loss function to calculate the classification loss, as shown in the formula:
[0089] in, For the real category, To predict class probabilities.
[0090] Finally, the low-rank matrix is processed using the backpropagation algorithm. and Updating the parameters can be represented as:
[0091]
[0092] It is understandable that, due to the weights of the classifier being determined by... and The adjustment, throughout the optimization process, involves updating only these two sets of parameters, while While keeping the image encoder and text encoder unchanged, this reduces the computational cost of fine-tuning the model, preserves the knowledge of the pre-trained model, and achieves accurate classification of multimodal fusion features.
[0093] After training, adjust the low-rank matrix terms. and Fix it. Freeze it during the reasoning phase. and Use the following formula to analyze the input features. Make a prediction:
[0094] The frozen classifier can efficiently process multimodal fusion features and output the final classification result.
[0095] Through this low-rank parameterized fine-tuning and inference process, the classifier can fully adapt to the complexity of multimodal data while maintaining resource efficiency, providing strong support for the classification of discriminatory content.
[0096] Step S5: Using a dynamic memory mechanism, calculate the similarity between the fusion features of the new input sample and the features of historical samples in the memory bank, and generate weighted aggregate features. Then, dynamically fuse the fusion features of the current sample with the weighted aggregate features.
[0097] This step is used to fully leverage knowledge from historical samples, enhance the model's understanding of complex contexts and long-distance dependencies, and thus improve the model's accuracy in detecting discriminatory content.
[0098] During the training process in this embodiment, a memory is constructed to store the features and corresponding labels of historical samples. This memory can be represented as:
[0099] In the formula, Historical Sample Feature representation, Historical Sample The tag, The entropy value is used to predict the model's confidence in predicting samples. Dynamic construction of the memory bank ensures that historical sample features accumulated during training can be continuously used in the inference phase, improving the predictive ability for new samples.
[0100] Among them, the predicted entropy value The calculation expression is:
[0101] In the formula, It is the number of categories (e.g., "discriminatory" versus "non-discriminatory"). It is the probability that the model predicts for a certain category.
[0102] Specifically, the embodiments of this application perform dynamic memory retrieval and feature fusion through the following steps: First, feature extraction is performed on the newly input image and text data to obtain the fused features of the current sample. The calculation of fused features is achieved through a self-attention mechanism, and the formula is:
[0103] In the formula, The input represents multimodal fusion features, including joint representations of images and text. Through a self-attention mechanism, the model is able to capture global contextual information from the multimodal features.
[0104] Then, by calculating the similarity between the current sample features and the features of each historical sample in the memory, the sample most similar to the current sample is selected. Features of historical samples Similarity is calculated using cosine similarity, and the formula is:
[0105] In the formula, This is the cosine similarity calculation function. Based on the similarity, select... The most similar samples are represented as:
[0106] This step can filter out historical samples that are semantically closest to the current sample, providing highly relevant semantic support for the dynamic memory mechanism.
[0107] Next, the memory bank was found The most similar historical sample features Perform weighted aggregation to generate weighted aggregated features. The formula for weighted aggregation is:
[0108] In the formula, the weights The weights are calculated based on the similarity distribution and are used to highlight historical features that are more similar to the current sample. The specific formula for the weights is as follows:
[0109] This process ensures that more historical features that are highly relevant to the semantics of the current sample are retained in the aggregated features, while reducing the impact of noisy samples.
[0110] Finally, the current sample features are fused through a gated recurrent unit. With weighted aggregation features Perform dynamic fusion and output the final fused representation. The gated recurrent unit's dynamic fusion mechanism can adjust the contribution weights of current sample features and historical features according to task requirements. For example, if the semantic information of the current sample features is strong enough, the contribution of historical features will be weakened; if the current sample features are semantically ambiguous, historical features will play a greater supporting role.
[0111] Specifically, the fusion logic of the gated loop unit is as follows: (1) Calculate the update gate Update Gate Used to control the current input features Hidden state at the previous time step The weight ratio in this time step. The formula for updating the gate is:
[0112]
[0113] In the formula, This involves concatenating the current sample's fused features with the weighted aggregated features. This is the hidden state from the previous time step. and The weight matrix is a learnable matrix. For bias terms, This is the Sigmoid activation function.
[0114] Update Gate The value range is within Between, it has the following meanings: when →1 indicates that the model relies more on historical information. .
[0115] when →0 indicates that the model relies more on the current input information. .
[0116] Intuitively, the update gate is like a "switch" that controls the relative importance of the model to historical memory and current input.
[0117] (2) Calculate the reset door Reset the door Used to control historical information In candidate state The degree of preservation within. The formula for resetting the door is:
[0118] In the formula, and The weight matrix is a learnable matrix. This is a bias term.
[0119] Similar to the update gate, the reset gate The value range is within Between, it has the following meanings: when →1, Historical Information It was completely preserved.
[0120] when →0, historical information is completely ignored, current input information Dominant.
[0121] Intuitively, the reset gate determines whether the model "forgets" the historical information of the previous time step.
[0122] (3) Generate candidate states Candidate state It is an intermediate result of the current input and adjusted historical information, used to calculate the final fused representation. The formula for the candidate state is:
[0123] in, , The weight matrix is a learnable matrix. For bias terms, This indicates the historical state after the gate adjustment, and the candidate state. It is a potential hidden state in the current time step.
[0124] Candidate state This represents a potential hidden state, which combines the current input and some historical information, and has the following meanings: when =1 indicates that historical information is completely ignored and the candidate state is determined solely by the current input.
[0125] when =0 indicates that the candidate state fully combines historical information and current input.
[0126] (4) Calculate the final fused representation Final fusion representation By dynamically balancing the current input and historical information through updating the gate, the formula is as follows:
[0127] in, The contribution representing the hidden state of history, This indicates the contribution of the currently input information.
[0128] Intuitively, the final fusion representation Dynamically integrates historical memory and current input At the same time, the weight ratio of the two is adjusted according to the task requirements.
[0129] Through the dynamic fusion mechanism of the gated recurrent unit, the model can flexibly handle complex contexts and dynamically adjust the contribution weights of historical information and current input in the final representation, providing more accurate and comprehensive feature representation capabilities for subsequent classification tasks.
[0130] Step S6: Classify the final fused representation using a classifier and output a discriminatory or non-discriminatory judgment result.
[0131] In the embodiments of this application, the final fusion representation The data is input into a classifier to classify samples and output a judgment result indicating whether the sample is "discriminatory" or "non-discriminatory". Simultaneously, to enhance the model's memory capacity, the features and prediction results of the current sample are saved to a dynamic memory bank to support subsequent predictions.
[0132] Specifically, the classification task is performed by a fully connected layer. The classifier receives the fused representation of the input. Through the weight matrix After a linear transformation, the softmax function is used to generate the prediction result. The specific calculation formula for the classifier is as follows:
[0133] in, As the weight of the classification head, The final prediction result represents the probability that the input sample belongs to the discriminatory or non-discriminatory category. Ultimately, based on... The category corresponding to the highest probability value in the data is used to determine the final classification result.
[0134] After classification, in order to fully utilize the information of the current sample and improve the model's performance in subsequent predictions, the dynamic memory is updated in this embodiment. Specifically, the fused features of the current sample are... and predicted labels Add to the memory bank and update the formula as follows:
[0135] This update process can incorporate new samples and prediction results into the memory bank, enhancing the memory bank's ability to accumulate historical information.
[0136] By combining classification with memory updates, this embodiment not only completes the final classification task of the samples but also provides support for continuous optimization of the model in dynamic scenarios. This linkage design between classification and memory can effectively capture long-term semantic relationships between samples, enhancing the robustness and adaptability of the model. It is particularly significant for detecting discriminatory content in complex multimodal and multilingual scenarios.
[0137] To verify the effectiveness of the method proposed in this application, comparative experiments were conducted using the unimodal visual model ViT, the unimodal text model mBERT, and the multimodal models CLIP and ALBEF as benchmark models. The experiments were performed on two multilingual multimodal datasets, BHM and Multi3Hate, with the objective of detecting whether the datasets contain discriminatory content. The F1 score was used as the evaluation metric to measure the model's overall performance in terms of precision and recall.
[0138] The experimental results of different models on different datasets are shown in Table 1.
[0139] Table 1
[0140] Experimental results show that the proposed method outperforms other benchmark models on both datasets. Specifically, by introducing a cross-attention mechanism, the proposed method significantly improves the effectiveness of modality alignment and fusion, better capturing the deep interaction relationships between image and text modalities, thereby enhancing the accuracy of discriminatory content detection. On the BHM dataset, the proposed method achieves a significantly higher F1 score than mainstream multimodal models such as CLIP and ALBEF; on the Multi3Hate dataset, thanks to stronger support for multilingual text, the proposed method also demonstrates superior detection performance in multilingual and multimodal scenarios.
[0141] Experimental results fully validate the effectiveness of the proposed method in multimodal discriminatory content detection tasks, particularly demonstrating higher robustness and applicability in multilingual and multimodal environments. This indicates that the proposed method has significant practical value and academic significance in addressing the problem of detecting implicit discriminatory content in social media.
[0142] To achieve the above embodiments, this application also proposes a discrimination detection device for multimodal and multilingual information. Figure 3 This is a schematic diagram of a discrimination detection device for multimodal and multilingual information provided in an embodiment of this application. Figure 3 As shown, the device includes: The data preprocessing module 100 is used to preprocess the original image and text data, perform word segmentation and stop word removal on the text data, and resize and normalize the image data. The feature extraction module 200 is used to input preprocessed text data into the text encoder to extract the context embedding features of the text, and to input preprocessed image data into the visual encoder to extract the global features and local patch features of the image. The multimodal alignment and fusion module 300 is used to perform multimodal alignment of text features and image features using a cross-attention mechanism. By capturing the correlation between image and text modalities, it generates interactive features containing cross-modal semantic information and fuses the original modal features with the interactive features to generate fused features with multimodal contextual information. The classifier fine-tuning module 400 is used to freeze the image encoder and text encoder parts in the pre-trained model and only perform low-rank parameter fine-tuning on the classifier part. By adding low-rank matrix adjustment terms, the classifier's ability to classify multimodal fusion features is optimized. The dynamic memory fusion module 500 is used to calculate the similarity between the fusion features of a new input sample and the features of historical samples in the memory bank using a dynamic memory mechanism, and generate weighted aggregate features, and dynamically fuse the fusion features of the current sample with the weighted aggregate features. The classification and judgment module 600 is used to classify the final fused representation through a classifier and output discriminatory or non-discriminatory judgment results.
[0143] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0144] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A discrimination detection method for multimodal and multilingual information, characterized in that, Includes the following steps: The original image and text data are preprocessed. The preprocessing process includes word segmentation and stop word removal of the text data, and resizing and normalization of the image data. The preprocessed text data is input into the text encoder to extract the context embedding features of the text, and the preprocessed image data is input into the visual encoder to extract the global features and local patch features of the image. By utilizing a cross-attention mechanism to perform multimodal alignment and fusion of text and image features, and by capturing the correlation between image and text modalities, interactive features containing cross-modal semantic information are generated. The original modal features and interactive features are then fused to generate fused features with multimodal contextual information. The image encoder and text encoder parts of the pre-trained model are frozen, and only the classifier part is fine-tuned with low-rank parameters. By adding low-rank matrix adjustment terms, the classifier's ability to classify multimodal fusion features is optimized, specifically including: Initialize the classifier weight matrix low-rank matrix adjustment term and ,in: , Freeze the weights of the image encoder and text encoder to ensure that the parameters of the pre-trained model remain unchanged during fine-tuning, only adjusting the low-rank matrix of the classifier weights. and Update; Multimodal fusion features of input Prediction is performed using a classifier with an added low-rank matrix adjustment term. The classification result is represented as follows: in, The weights of the original classifier. For bias terms; The classification loss is calculated using the cross-entropy loss function, and the formula for the loss function is: in, For the real category, To predict class probabilities; Using the backpropagation algorithm, only low-rank matrices are processed. and Perform parameter updates, including updating the adjustment terms of the low-rank matrix; Using a dynamic memory mechanism, the similarity between the fusion features of a new input sample and the features of historical samples in the memory bank is calculated and a weighted aggregated feature is generated. The fusion features of the current sample are then dynamically fused with the weighted aggregated feature. The final fused representation is classified by a classifier, and the result is output as either discriminatory or non-discriminatory.
2. The method according to claim 1, characterized in that, The process of inputting preprocessed text data into a text encoder to extract contextual embedding features of the text, and inputting preprocessed image data into a visual encoder to extract global features and local patch features of the image, includes: Preprocessed text data Word embedding is performed, mapping each word to a high-dimensional vector, and contextual embedding features of the text are generated using an XLM-R text encoder. The formula is: In the formula, , The length of the text sequence. For the dimension of the embedded features, For the first Contextual representation of each word; The preprocessed image data is segmented to divide the image... Divided into A fixed-size patch, denoted as Linear mapping and positional encoding are performed on each patch to obtain the image features. ; The image features are input into the ViT image encoder, and a self-attention mechanism is used to perform multi-layer feature extraction to generate global features of the image. The result is expressed as: In the formula, , The dimension of image features. These are global features of the image.
3. The method according to claim 2, characterized in that, The method utilizes a cross-attention mechanism to perform multimodal alignment and fusion of text and image features. By capturing the correlation between image and text modalities, it generates interactive features containing cross-modal semantic information. The original modal features and interactive features are then fused to generate fused features with multimodal contextual information, including: Through image features Guide text features Generate image-guided text features ; Through text features Generate guiding image features Generate text-guided image features ; Image features Text features Image-guided text features Image features guided by text The features are concatenated to generate fused features with multimodal contextual information. , represented as: In the formula, This indicates a feature splicing operation.
4. The method according to claim 3, characterized in that, The image features Guide text features Generate image-guided text features ,include: Utilizing image features Form a query matrix Utilizing text features Generate key matrix Sum matrix Specifically, it is expressed as: , , in, , , The weight matrix is a learnable matrix; The relevance weight of the image to the text is calculated using the dot product attention formula, expressed as: Image-guided text features are generated based on relevance weights. , represented as: In the formula, Text features guided by images.
5. The method according to claim 4, characterized in that, The text features Generate guiding image features Generate text-guided image features ,include: Utilizing text features Generate query matrix Utilizing image features Generate key matrix Sum matrix Specifically, it is expressed as: , , in, , , The weight matrix is a learnable matrix; The relevance weight of text to image is calculated using the dot product attention formula, expressed as: Generate text-guided image features based on relevance weights. , represented as: In the formula, Image features guided by text.
6. The method according to claim 5, characterized in that, The process of utilizing a dynamic memory mechanism to calculate the similarity between the fusion features of a new input sample and the features of historical samples in the memory bank, generating weighted aggregate features, and dynamically fusing the fusion features of the current sample with the weighted aggregate features includes: During training, a memory is built to store the features and corresponding labels of historical samples, represented as: In the formula, Historical Sample Feature representation, Historical Sample The tag, To predict the entropy value; Feature extraction is performed on the newly input image and text data to obtain the fused features of the current sample, using the following formula: In the formula, This represents the multimodal fusion features of the input; Use current sample features Retrieve from the memory and calculate the fusion features of the current sample using the cosine similarity formula. Find the most similar sample by comparing its features with those of each historical sample in the memory. Features of historical samples , is represented as: In the formula, This is the function for calculating cosine similarity. Found in the memory bank Features of historical samples Perform weighted aggregation to generate weighted aggregated features. , is represented as: In the formula, the weights Calculated based on similarity distribution; The current sample's features are fused using a gated loop unit. With weighted aggregation features Perform dynamic fusion and output the final fused representation. .
7. The method according to claim 6, characterized in that, The current sample is fused using a gated loop unit. With weighted aggregation features Perform dynamic fusion and output the final fused representation. ,include: Computational update gate Used to control the current input features Hidden state at the previous time step The weighting ratio in this time step is calculated using the following formula: In the formula, This involves concatenating the current sample's fused features with the weighted aggregated features. This is the hidden state from the previous time step. and The weight matrix is a learnable matrix. For bias terms, Use the Sigmoid activation function; Calculate the reset door Used to determine historical information The retention rate is determined by the formula: In the formula, and The weight matrix is a learnable matrix. For bias terms; Generate candidate states This is used to combine the current input features with adjusted historical information, and the formula is: in, , The weight matrix is a learnable matrix. For bias terms, This indicates the historical state after the gate adjustment, and the candidate state. It is a potential hidden state in the current time step; Compute the final fused representation By updating the gate to dynamically balance the current input and historical information, the formula is: in, The contribution representing the hidden state of history. This indicates the contribution of the currently input information.
8. The method according to claim 7, characterized in that, The step of classifying the final fused representation using a classifier and outputting a discriminatory or non-discriminatory judgment result includes: The final fusion representation The input is fed into the classifier to generate the final prediction result, using the following formula: in, As the weight of the classification head, The final prediction result represents the probability that the input sample belongs to a discriminatory or non-discriminatory category.
9. A discrimination detection device for multimodal and multilingual information, characterized in that, include: The data preprocessing module is used to preprocess the original image and text data. The preprocessing process includes word segmentation and stop word removal of the text data, and resizing and normalizing of the image data. The feature extraction module is used to input preprocessed text data into the text encoder to extract the context embedding features of the text, and to input preprocessed image data into the visual encoder to extract the global features and local patch features of the image. The multimodal alignment and fusion module is used to perform multimodal alignment of text features and image features using a cross-attention mechanism. By capturing the correlation between image and text modalities, it generates interactive features containing cross-modal semantic information and fuses the original modal features with the interactive features to generate fused features with multimodal contextual information. The classifier fine-tuning module freezes the image encoder and text encoder parts of the pre-trained model, performing low-rank parameter fine-tuning only on the classifier part. By adding low-rank matrix adjustment terms, it optimizes the classifier's ability to classify multimodal fusion features. Specifically, it includes: Initialize the classifier weight matrix low-rank matrix adjustment term and ,in: , Freeze the weights of the image encoder and text encoder to ensure that the parameters of the pre-trained model remain unchanged during fine-tuning, only adjusting the low-rank matrix of the classifier weights. and Update; Multimodal fusion features of input Prediction is performed using a classifier with an added low-rank matrix adjustment term. The classification result is represented as follows: in, The weights are those of the original classifier. For bias terms; The classification loss is calculated using the cross-entropy loss function, and the formula for the loss function is: in, For the real category, To predict class probabilities; Using the backpropagation algorithm, only low-rank matrices are processed. and Perform parameter updates, including updating the adjustment terms of the low-rank matrix; The dynamic memory fusion module is used to calculate the similarity between the fusion features of a new input sample and the features of historical samples in the memory bank using a dynamic memory mechanism, and generate weighted aggregate features. The fusion features of the current sample are then dynamically fused with the weighted aggregate features. The classification and judgment module is used to classify the final fused representation using a classifier and output discriminatory or non-discriminatory judgment results.