Bias risk assessment method and system for multi-modal large model
By constructing a bias lexicon and image library, and combining the attention weights of a multimodal large model, the bias risks of the multimodal large model are identified and quantified, solving the problems of cross-modal bias identification and latent risk detection, and realizing the monitoring and early warning of potential bias in the power customer service scenario.
Patent Information
- Application Number
- CN202511145845.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-04
AI Technical Summary
Existing multimodal large models have problems during training and application, such as the inability to identify biases across modalities, the lack of ability to detect implicit risks, and the inability to quantify the degree of bias.
We construct a bias lexicon and a bias image lexicon, identify potential biases through cross-modal matching of text and images, calculate the text-image bias value by combining the attention weights of a multimodal large model, determine the source of implicit bias, and classify the risk level of bias through weighted scoring.
It enables interpretable analysis of hidden biases in complex interaction scenarios, real-time monitoring of potential bias risks in power customer service language and recommended images, automatic identification and early warning of potential bias issues, and reduction of user complaints and public opinion risks.
Smart Images

Figure CN120894589A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology and relates to a method for assessing bias risk in multimodal large models. Background Technology
[0002] With the rapid development of artificial intelligence technology, multimodal large models have demonstrated powerful capabilities in scenarios such as image-text joint understanding, cross-modal retrieval, and content generation, and are widely used in fields such as social media analysis, intelligent advertising, virtual assistant interaction, and medical image diagnosis. However, these models may implicitly absorb social biases in the data during training and application, leading to unfairness in generated content or decision-making results. Therefore, developing efficient multimodal bias detection technology has become a key requirement for improving model fairness and ensuring technological ethical compliance.
[0003] Currently, there are some related technologies and methods for assessing the bias risk of multimodal large models. For example, patent CN117437507A provides a bias assessment method for evaluating image recognition models. By constructing an evaluation model, optimizing training data, and calculating fairness indicators, it achieves bias assessment of object detection models. Meanwhile, Chinese patent CN117786054A discloses a generative text visual question answering method and system based on prompting learning. This method utilizes a Transformer encoder and decoder, introduces a cloze test task, and performs multimodal data processing to optimize the reasoning ability of text visual question answering models and reduce language bias. Furthermore, patent CN118820093A provides a bias detection method for natural language reasoning of large language models based on metamorphic testing. By constructing metamorphic relationships, generating test cases, and analyzing model output results, it can detect bias in large language models in natural language reasoning tasks.
[0004] However, existing bias detection methods still have the following problems:
[0005] 1) Existing methods only assess bias risk for single-modal data and do not build cross-modal association mechanisms; among them, text detection relies on manual rules or single-modal word vectors and has no image association; image detection only analyzes visual features and has no text semantic alignment, so it cannot quantify the bias risk generated by image-text combination; it cannot identify cross-modal inconsistencies and lacks the ability to detect implicit bias generated by image-text combination.
[0006] 2) Existing methods only analyze input data bias and lack bias detection in the model generation process. They do not consider the risk of bias in other modalities when the model is overly dependent on a particular modality.
[0007] 3) Existing methods can only detect whether there is bias in the data, but cannot quantify the degree of bias in the model. Summary of the Invention
[0008] To address the shortcomings of existing technologies, such as the inability to identify cross-modal bias, the inability to identify latent risks, and the lack of bias detection in models, this invention provides a bias risk assessment method based on a multimodal large model. The method includes the following steps: constructing a bias lexicon and a bias image lexicon; comparing the actual output text of a power company during customer service communication with the baseline vectors in the bias lexicon to generate a bias risk value for the actual output text and identify biased text; calculating the image stereotype similarity between the actual output image and the image template to identify biased images; calculating the cross-modal sensitivity score between the actual output text and the actual output image to identify latent bias; calculating the image-text bias value based on the attention weights of the multimodal large model for text and image output to identify the abnormal layer of the model and determine the source of latent bias; and calculating a comprehensive bias risk score to classify multiple image-text bias risk levels. This invention's bias risk assessment enhances the interpretability of latent biases in complex interaction scenarios, thereby enabling real-time monitoring of potential bias risks in power company customer service language and recommended images.
[0009] This invention adopts the following technical solution. One aspect of this invention provides a method for assessing bias risk in a multimodal large model, comprising the following steps:
[0010] Step 1: Collect bias label words from public datasets, obtain candidate bias words manually labeled from historical data of power customer service documents, form different bias categories based on candidate bias words and bias label words, and construct a bias word library; collect bias images from public datasets, cluster the bias images to obtain image templates for each type of bias image cluster, and construct a bias image library.
[0011] Step 2: Compare the actual output text of the power company during customer service communication with the baseline vector of the bias category in the bias lexicon to generate the bias risk value of the actual output text and identify biased text;
[0012] Step 3: Calculate the stereotypical similarity between the actual output image and the image template in the biased image library within the predetermined time interval above and below the actual output text to identify biased images;
[0013] Step 4: Calculate the cross-modal sensitivity scores of the actual output text and the actual output image, and identify implicit bias by combining the identification of biased text and biased images; based on the attention weights of the multimodal large model on the text and image outputs, calculate the anomaly layer of the image-text bias value identification model to determine the source of implicit bias.
[0014] Step 5: Combine the bias risk value of the actual output text, the image stereotype similarity of the actual output image, the cross-modal sensitivity score, and the image-text bias value, and use weighted calculation to calculate the comprehensive bias risk score, and divide it into multiple image-text bias risk levels.
[0015] Preferably, the process of constructing the bias lexicon in step 1 includes:
[0016] Import publicly available bias vocabulary datasets to obtain bias label words; perform data mining on historical data from electricity customer service documents to collect candidate bias words;
[0017] A text encoder based on a multimodal large model generates context-dependent word vectors for each bias label word and candidate bias word; all word vectors are clustered to form a bias word cluster.
[0018] Calculate the mean vector of word vectors within the biased word cluster and use it as the semantic center of the corresponding biased word cluster; extract the keywords of the biased label words corresponding to the biased category from the biased label words, generate the word vectors of the keywords, and calculate the mean of the word vectors of the keywords as the reference vector of the corresponding category; match the semantic center with the reference vector of the biased category to determine the biased category label of each biased word cluster.
[0019] A bias lexicon is constructed based on bias label words, candidate bias words, bias word cluster numbers, and bias category labels.
[0020] Preferably, the process of constructing the biased image library in step 1 includes:
[0021] Import publicly available biased image databases, collect and label biased images; extract local detail features and global semantic features from biased images to generate comprehensive feature vectors; cluster the comprehensive feature vectors to divide all biased images into multiple biased image clusters and determine the center vector of each cluster.
[0022] Select a predetermined number of biased images that are closest to the center vector from the cluster as the image templates of the corresponding cluster; calculate the correlation between each image template and the bias words corresponding to each bias category in the bias vocabulary, and take the bias category with the highest correlation as the classification result of the corresponding image template to determine the bias category label of the biased image cluster;
[0023] A biased image library is constructed based on image templates of biased image clusters and their corresponding bias category labels, comprehensive feature vectors, and center vectors.
[0024] Preferably, in step 2, after preprocessing the actual output text, dynamic semantic encoding is performed to obtain semantic encoding; the baseline vector of each bias category is extracted from the bias lexicon, and for each word in the actual output text, the semantic encoding and the bias intensity of the baseline vector of each bias category are calculated.
[0025] Calculate the average attention weight of each word to other words in the actual output text as the corresponding word's context dependency score;
[0026] Based on the bias intensity and context dependence score of each word, the bias risk value of the corresponding actual output text is calculated by combining the word frequency-inverse document frequency in the actual output text.
[0027] If the bias risk value is not less than the predetermined text bias risk threshold, it means that the corresponding actual output text matches the text bias library and there is bias.
[0028] Preferably, in step 3, the image feature vector of the actual output image is extracted, and its similarity is calculated with the comprehensive feature vector in the bias image library to obtain the image stereotype similarity.
[0029] If the stereotype similarity of an image is not less than a predetermined image bias risk threshold, it means that the corresponding actual output image matches the biased image library and there is bias.
[0030] Preferably, in step 4, the process of identifying implicit bias includes:
[0031] Extract the text semantic vector of the actual output text and the image semantic vector of the actual output image, and calculate the cross-modal sensitivity score of both.
[0032] If the cross-modal sensitivity score is not less than the predetermined first cross-modal bias threshold and is less than the predetermined second cross-modal bias threshold, the existence of implicit bias is determined based on the bias risk value of the actual output text and the image stereotype similarity of the actual output image; if both the actual output text and the actual output image have bias, then the actual output text and the actual output image have implicit bias.
[0033] Preferably, in step 4, for each cross-modal attention layer of the multimodal large model, the weights pointing to the text and the weights pointing to the image are extracted from the attention matrix output by each attention head of each layer, respectively, to obtain the attention weights of the text and the image as well as the cross-modal attention weights, and the cross-modal attention ratio is calculated;
[0034] The results of all heads in each cross-modal attention layer are summed to obtain the total attention weights for text and image in that layer. The average cross-entropy of the cross-modal attention ratio is used as the modal interaction strength coefficient, and it is combined with the total attention weights for text and image to obtain the text-image bias value.
[0035] Based on the text-image bias value, the abnormal layer in the cross-modal attention layer is located, and the text-high attention content and image-high attention content in this layer are identified. They are then matched with the bias categories in the bias lexicon and bias image lexicon, respectively. If both matches are successful, the source of the latent bias is the abnormal amplification of bias information by the model, and the abnormal layer is located as the bias source layer. Otherwise, it indicates that the latent bias is caused by model noise.
[0036] Preferably, the process of locating the anomaly layer and identifying textual and image content of high interest includes:
[0037] Calculate the image-text bias value for each cross-modal attention layer in the multimodal large model to form an image-text bias sequence; calculate the mean and standard deviation of the image-text bias sequence; weight the mean and standard deviation to obtain the first bias threshold, and take the maximum value of the first bias threshold and the predetermined second bias threshold as the predetermined anomaly threshold; when there is an image-text bias value in the image-text bias sequence that is greater than the predetermined anomaly threshold, the layer is considered an anomaly layer;
[0038] In the anomaly layer, the attention weight distribution of text words and image regions is obtained, and words and regions with attention weights greater than a predetermined weight threshold are identified as the corresponding text high-attention content and image high-attention content.
[0039] Preferably, the process of performing bias category matching includes:
[0040] Extract the vector of the highly relevant content in the text, calculate the cosine similarity with the baseline vectors of different bias categories in the bias lexicon, and if the similarity exceeds a predetermined similarity threshold, the word is matched with the corresponding bias category.
[0041] Extract the feature vector of the high-interest content region in the image, calculate the cosine similarity with the comprehensive feature vector of the image template in the bias image library, and if the similarity exceeds a predetermined similarity threshold, the region is determined to match the corresponding bias category.
[0042] Another aspect of the present invention provides a bias risk assessment system for a multimodal large model, comprising:
[0043] The image and text bias library construction module collects bias label words from public datasets, extracts candidate bias words in the power customer service communication process, and constructs a bias word library; it also collects bias images from public datasets and constructs a bias image library.
[0044] The text bias identification module compares the actual output text of the power company during customer service communication with the baseline vector of bias categories in the bias lexicon, generates a bias risk value for the actual output text, and identifies biased text.
[0045] The image bias recognition module calculates the stereotypical similarity between the actual output image and the image templates in the bias image library within a predetermined time interval before and after the actual output text, and identifies biased images.
[0046] The implicit bias identification module calculates the cross-modal sensitivity scores of the actual output text and the actual output image, and identifies implicit bias by combining the identification of biased text and biased images; based on the attention weights of the multimodal large model on the text and image outputs, it calculates the abnormal layer of the image-text bias value identification model to determine the source of implicit bias.
[0047] The image-text bias risk classification module integrates the bias risk value of the actual output text, the image stereotype similarity of the actual output image, the cross-modal sensitivity score, and the image-text bias value, and uses a weighted calculation to calculate a comprehensive bias risk score, classifying multiple image-text bias risk levels.
[0048] Compared with the prior art, the beneficial effects of the present invention include at least the following:
[0049] 1. This invention constructs a text bias lexicon and a bias image lexicon, which enables cross-modal bias matching of text and images, thereby effectively identifying potential correlation biases between different modalities and improving the ability to identify complex implicit biases.
[0050] 2. This invention achieves risk assessment and identification of potentially biased text by comparing the output text with the baseline vector in the bias lexicon, thus enhancing the detection capability of textual bias. Combined with a biased image library, it performs image stereotype similarity comparison on the output images to identify image bias information. Through a cross-modal sensitivity scoring mechanism, it integrates information from text and images, improving the identification of implicit bias in text-image combinations. Furthermore, it can determine the source of implicit bias based on the attention weight anomalies of a multimodal large model, enhancing the interpretable analysis capability of hidden biases in complex interactive scenarios. This allows for real-time monitoring of potential bias risks in power customer service language and recommended images, effectively reducing user complaints and public opinion risks caused by inappropriate expressions or visual presentations.
[0051] 3. This invention classifies different levels of bias risk by weighted fusion of text bias risk value, image stereotype similarity, cross-modal sensitivity score and image-text bias value, enabling the system to automatically identify and warn of potential bias problems in customer service interactions, and assisting power companies in proactively discovering discriminatory language or stereotypical image outputs in the service process. Attached Figure Description
[0052] Figure 1 This is a framework diagram for bias risk assessment of a multimodal large model provided by the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this application are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.
[0054] Example 1
[0055] This embodiment provides a method for assessing bias risk in multimodal large models, taking the Qwen-VL model as an example. (See also...) Figure 1 This includes the following steps:
[0056] Step 1: Collect bias label words from public datasets, obtain candidate bias words manually labeled from historical data of power customer service documents, form different bias categories based on candidate bias words and bias label words and construct a bias word library; collect bias images from public datasets, cluster the bias images to obtain image templates for each type of bias image cluster, and construct a bias image library.
[0057] The process of building the bias lexicon in step 1 includes:
[0058] Import publicly available bias vocabulary datasets, including the HateBase hate speech detection dataset and the WikiBias bias detection dataset, to obtain verified bias words, i.e., bias label words; acquire historical data of electricity customer service documents for data mining, collect potential candidate bias words, and manually label these candidate words, removing low-frequency words to avoid noise interference; among them, electricity customer service documents include electricity consultation, fault repair, complaints and suggestions, electricity application, texts of communication with electricity customer service, texts of communication with customers, etc.
[0059] The text encoder based on the Qwen-VL model generates context-dependent word vectors for each bias label word and candidate bias word to capture their subtle semantic differences in different contexts.
[0060] The HDBSCAN density clustering algorithm is used to cluster all word vectors, automatically identifying high-density semantic regions of candidate biased words and forming biased word clusters.
[0061] In this embodiment, the minimum density sample size is set to 10, meaning each effective cluster contains at least 10 word vectors of candidate biased words, ensuring the clustering results are semantically representative and stable. The minimum cluster size is set to 10 for kernel density estimation, distinguishing between core points and noise points, and more rigorously selecting effective semantic clusters. All candidate biased word vectors are batch-input into the algorithm, where each word vector is a numerical vector with consistent dimensions. The core distance and reachability distance between all word vectors are calculated, with the reachability distance using Euclidean distance by default. A minimum spanning tree (MST) is constructed based on density connectivity, forming a hierarchical structure of locally density-reachable nodes. Clustering structure: Merge word vectors with high local density and mutual reachability. When the distance between two word vectors is less than the current density threshold and both belong to the core point or can be connected through the core point, they are merged into the same cluster. Isolated points with insufficient local density or lack of connectivity with other word vectors are marked as noise points and excluded. If the proportion of noise samples is too high, the set parameters can be readjusted and the clustering task can be re-executed. When the density threshold increases to the point where no new clusters can be merged or the affiliation of all sample points has stabilized, the iteration ends and the clustering results are output. In the clustering results, only high-density regions with high sample density and strong semantic coherence are retained as effective semantic clusters, and word vectors corresponding to noise points are removed.
[0062] Each biased word cluster is assigned a unique number, and the mean vector of word vectors within the biased word cluster is calculated as the semantic center of the corresponding biased word cluster.
[0063] Extract keywords from the bias label vocabulary for each bias category, generate word vectors for the keywords using the Qwen-VL text encoder, and calculate the mean of the word vectors as the baseline vector for the corresponding category.
[0064] The semantic center is matched with the baseline vector of the bias category to determine the bias category label of each bias word cluster; specifically, based on the semantic center of each bias word cluster, the bias category label of each bias word cluster is matched using the cosine similarity function; the formula is as follows:
[0065]
[0066] Among them, Label(C k ) represents C k Bias category labels; C k This represents the k-th cluster. The vector representing the cluster center of the k-th cluster; μ l The reference vector represents the bias category l; CosineSimilarity() represents the cosine similarity function, used to measure the similarity between two vectors; l∈L represents the bias category l in the bias category set; This means selecting the bias category with the highest cosine similarity from the set of bias categories;
[0067] A bias lexicon is constructed based on candidate bias words, their corresponding bias word cluster numbers, and bias category labels.
[0068] In this embodiment, in the bias lexicon, each candidate bias word is treated as a row of data, accompanied by its cluster number, bias category label, and corresponding baseline vector.
[0069] The process of building the biased image library in step 1 includes:
[0070] Import publicly available biased image databases, including biased image datasets such as Gender Shades and FairFace, collect biased images and label them;
[0071] ResNet-50 is used to extract local detail features from biased images, and the image encoder in the Qwen-VL model is used to extract global semantic information and generate global semantic features. Based on the local detail features and global semantic features, a comprehensive feature vector is generated. The specific formula is as follows:
[0072]
[0073] In the formula, f tem_img f represents the comprehensive feature vector; local Indicates local detail features; f global GAP(f) represents the global semantic features; α is the local feature weight, β is the global feature weight, and the values of α and β are both within the interval (0,1) and α+β=1; local ) c f local Features on the c-th channel after global average pooling; Score c The function represents feature similarity and is used to measure the correlation between the c-th channel of a local feature and the global semantic feature. TopK() is a channel selection function used to select the K channels that are most relevant to the local detail features and the global semantic features after global average pooling.
[0074] In this embodiment, each channel corresponds to a semantic feature. The cosine similarity is used to measure the relevance of each channel to the global feature, and the 512 channels with the highest scores are retained.
[0075] Based on the comprehensive feature vector, the HDBSCAN density clustering algorithm is used to divide all biased images into multiple biased image clusters, and the center vector of each cluster is calculated. The specific formula is as follows:
[0076]
[0077] Where, μ G Let G be the center vector of cluster G, and let G be the average position of the composite feature vectors corresponding to all biased images within the cluster; ||G|| represents the number of biased images in the cluster. This represents the comprehensive feature vector of the i-th biased image within cluster G;
[0078] For each biased image cluster, a predetermined number of biased images closest to the center vector from the cluster are selected as image templates for the corresponding cluster based on Euclidean distance. The correlation between each image template and the biased vocabulary corresponding to each bias category in the bias lexicon is calculated one by one, and the bias category with the highest correlation is taken as the classification result for the corresponding image template. The specific formula for the correlation is:
[0079]
[0080] in, express Bias categories; W l This represents the set of bias words belonging to bias category l in the bias lexicon. The feature vector representing the j-th bias word of bias category l is extracted by the Qwen-VL model; in this embodiment, the predetermined number is 5.
[0081] Based on the classification results of image templates in each biased image cluster, the bias category label of the corresponding biased image cluster is determined by majority voting. If there is a tie, the classification result of the cluster's center vector is used as the bias category label for the corresponding biased image cluster, ensuring that each cluster is uniquely and interpretably classified. The image templates of all biased image clusters, along with their corresponding bias category labels, comprehensive feature vectors, and center vectors, are stored in a structured manner to form a biased image library for subsequent similarity retrieval.
[0082] Step 2: Compare the actual output text of the power company during customer service communication with the baseline vector of the bias category in the bias lexicon to generate the bias risk value of the actual output text and identify biased text.
[0083] The actual output text of the power company during customer service communication is collected and preprocessed. The actual output text includes consultation records, online chat, WeChat communication, emails, and automatic reply records of intelligent customer service. The preprocessing process includes: using spaCy to process phrases, splitting the actual output text into semantic units, removing words without actual meaning, and retaining key bias carriers such as nouns and adjectives.
[0084] The Qwen-VL model is used to dynamically semantically encode each word in the preprocessed actual output text, resulting in the semantic code of each word. The Qwen-VL model generates semantic vectors by combining the context of words with the surrounding text and uses an attention mechanism to capture the dependencies between words, so that the generated vectors not only contain the semantic information of the words themselves, but also reflect their relevance in a specific context.
[0085] The baseline vector for each bias category is extracted from the bias lexicon. For each word in the actual output text, the cosine similarity between the semantic code corresponding to each word and the baseline vectors of all bias categories is calculated, and the maximum value is taken as the bias strength of the corresponding word. The specific formula is as follows:
[0086]
[0087] Among them, S(w x ) indicates the word w x The intensity of bias; The word w x Semantic encoding; μ l Let L be the base vector of bias category l; L represents the set of all bias categories.
[0088] For each actual output text, obtain the word w. x Attention weights for other words; calculate w x The average attention weight (or the maximum value) for other words is used as their context-dependent score; the specific formula for the average attention weight is as follows:
[0089]
[0090] Where T represents the number of words in the actual output text; Attention_Score(w x ,w y ) indicates the word w x For w x The attention weights come from the last layer of the Qwen-VL text encoder;
[0091] Based on the bias intensity and context dependence score of each word, the bias risk value of the corresponding actual output text is calculated using the following formula:
[0092]
[0093] Among them, TF_IDF(w x () represents the term frequency-inverse document frequency, indicating the frequency of the word w x Importance in actual output text; S(w x ) indicates the word wx The intensity of bias; D(w) x ) indicates the word w x The context-dependent score reflects the semantic influence of the word in the actual output text, avoiding misjudgment of isolated words; The word w x Corresponding bias category l x The weights are used to distinguish the degree of harm of different bias categories; γ1, γ2 and γ3 represent balance coefficients, satisfying γ1+γ2+γ3=1.
[0094] When VRB is not less than the predetermined text bias risk threshold, it means that the corresponding actual output text matches the text bias library and there is bias; when VRB is less than the predetermined text bias risk threshold, it means that no obvious bias features were detected in the corresponding actual output text; in this embodiment, the predetermined text bias risk threshold is 0.6.
[0095] Step 3: Calculate the stereotypical similarity between the actual output image and the image template in the bias image library within the predetermined time interval above and below the actual output text to identify biased images.
[0096] ResNet-50 is used as the image feature extraction network to perform convolutional feature extraction on the actual output images within a predetermined time interval before and after the actual output text. The 2048-dimensional feature vector of the image is extracted and reduced to 128 dimensions to obtain the image feature vector. The actual output images in the customer service communication process include images in the documents uploaded by the customer, images in the chat between the customer and the customer service, and images of the automatic reply from the intelligent customer service. In this embodiment, the images sent before and after the actual output text output time in the chat between the customer and the customer service are obtained as the actual output images.
[0097] The similarity between the image feature vector and the comprehensive feature vector in the bias image database is calculated to measure the matching degree between the stereotype features in the image and the bias feature database, thus obtaining the image stereotype similarity of the actual output image. The specific formula is as follows:
[0098]
[0099] Where IS represents the similarity between the image feature vector and the image stereotypes in the biased image library; f img Represents the image feature vector; This represents the comprehensive feature vector of the r-th image template in the biased image library;
[0100] When IS ≥ 0.6, it indicates that the corresponding actual output image matches the biased image library and there is bias; when IS < 0.6, it indicates that no obvious bias features were detected in the corresponding actual output image.
[0101] Step 4: Calculate the cross-modal sensitivity scores of the actual output text and the actual output image, and identify implicit bias by combining the identification of biased text and biased images; based on the attention weights of the large language model on the text and image outputs, calculate the anomaly layer of the image-text bias value identification model to determine the source of implicit bias.
[0102] The preprocessed actual output text is input into the text encoder of the Qwen-VL model to obtain the corresponding text semantic vector; the actual output image is input into the image encoder of the Qwen-VL model to obtain the corresponding image semantic vector, ensuring that the features of both text and image modalities are mapped to the same semantic space; based on the text semantic vector and image semantic vector, a cross-modal sensitivity score is calculated.
[0103] CSS = 1 - CosineSimilarity(u text ,u img );
[0104] Among them, u text The text semantic vector representing the actual output text; u img This represents the image semantic vector of the actual output image;
[0105] When CSS < 0.3, the actual output text and image are highly consistent, indicating good semantic alignment, no obvious conflict, and extremely low risk of bias. When 0.3 ≤ CSS < 0.5, the actual output text and image have slight inconsistencies, indicating that the text and image are generally matched but have subtle differences. The presence of implicit bias should be judged by combining the bias risk value of the actual output text and the stereotype similarity of the actual output image. If both VBR and IS indicators indicate the presence of bias risk, it means that although the text and image are superficially consistent, there are implicit stereotypes or biases in the deeper semantics, which is classified as implicit bias. When 0.5 ≤ CSS < 0.8, the text and image are significantly and moderately inconsistent, indicating a low degree of semantic matching between the actual output text and image, with a relatively obvious risk of conflict. This should be given special attention and manually reviewed. When CSS ≥ 0.8, the actual output text and image are severely inconsistent, indicating that the semantic alignment between the text and image has failed and there is a great conflict. This is directly judged as an explicit conflict, indicating a serious risk of bias in the output content.
[0106] In this embodiment, if the CSS value is negative or greater than 1, it is considered abnormal. In this case, it is necessary to jointly check whether the text and image are correctly normalized and the quality of the input preprocessing, and to verify by matching a biased vocabulary or biased image library to ensure that the data quality is correct before recalculating. When there is implicit bias in the text and image, the degree of over-reliance of the large language model on the text or image modality is quantified to detect whether the bias is amplified due to the imbalance of modality weights. If the modality bias value of a certain layer is abnormally high, it suggests that the layer may be amplifying the risk of bias due to over-focusing on a certain modality, thus providing a basis for subsequent steps.
[0107] Step 4, which calculates the anomaly layer of the image-text bias identification model and determines the source of implicit bias, includes the following steps:
[0108] In the Transformer cross-modal attention layer of the Qwen-VL model, the weights pointing to text and the weights pointing to images are extracted from the attention matrix output by each attention head of each layer, respectively, to obtain the attention weights for text and images, as well as the cross-modal attention weights, which are the sum of attention points pointing to another modality (e.g., all weights pointing to images when the query is text); the cross-modal attention percentage is calculated; the specific formula is as follows:
[0109]
[0110] In the formula, This represents the cross-modal attention percentage of the h-th attention head in the q-th attention layer of a large language model. This represents the cross-modal attention weight of the h-th cross-modal attention head in the q-th attention layer; and Let represent the text attention weight and image attention weight of the h-th attention head of the q-th attention layer, respectively;
[0111] The results of all heads in each cross-modal attention layer q are summed to obtain the total attention weight for the text and image parts of that layer. The average cross-entropy of the cross-modal attention ratio is used as the modal interaction strength coefficient. This coefficient is combined with the total attention weight for text and image to calculate the text-image bias value. The specific formula is as follows:
[0112]
[0113] in, This represents the total attention weight of the q-th layer for the text modality; This represents the total attention weight of the q-th layer for the image modality; ε = e -6 Used to prevent division by zero errors; η is an adjustment coefficient that controls the weight of the interaction strength; Let I be the modal interaction strength coefficient, representing the average cross-entropy of the q-th layer cross-modal attention head, reflecting the degree of fusion between the text and image modalities; average cross-entropy I (q) The lower the value, the more concentrated the attention between modalities, which may imply bias and amplify the risk.
[0114] Based on the image-text bias value, the abnormal layer in the cross-modal attention layer is located, and the text-high-attention content and image-high-attention content in this layer are identified. These are then matched with bias categories in the bias lexicon and bias image lexicon, respectively. The specific process includes:
[0115] The image-text bias value of each cross-modal attention layer in the Qwen-VL model is calculated to form an image-text bias sequence. The mean and standard deviation of the image-text bias sequence are calculated. The mean and standard deviation are weighted and used as a first bias threshold. The maximum value of the first bias threshold and a predetermined second bias threshold is taken as a predetermined anomaly threshold. When there is an image-text bias value in the image-text bias sequence that is greater than the predetermined anomaly threshold, the layer is considered an anomaly layer. In this embodiment, the predetermined anomaly threshold is max(0.3, μ+3σ), where μ and σ are the mean and standard deviation of the image-text bias sequence, respectively.
[0116] In the anomaly layer, the attention weight distribution for text words and image regions in this layer is obtained by extracting the cross-modal attention weight matrix. Words and regions with attention weights greater than a predetermined weight threshold are identified as the corresponding high-attention text content and high-attention image content. The high-attention text content is used to extract vectors through the Qwen-VL text encoder, and the cosine similarity is calculated with the baseline vectors of different bias categories in the bias lexicon. If the similarity exceeds a predetermined similarity threshold, the word matches the corresponding bias category. The high-attention image content is used to extract region feature vectors through the Qwen-VL image encoder, and the cosine similarity is calculated with the comprehensive feature vectors of image templates in the bias image lexicon. If the similarity exceeds a predetermined similarity threshold, the region is determined to match the corresponding bias category.
[0117] If highly-watched content is determined to contain biased information—that is, if highly-watched text or image content areas are determined to contain biased information—then this anomaly layer is designated as the bias source layer, indicating that this location has an abnormal amplification effect on potential biased information. If the highly-watched content does not match a bias category label, the sudden increase in the text-image bias value of this anomaly layer may be caused by semantic noise unrelated to bias and is not treated as a source of bias. The matched words, image templates, and corresponding attention weights in the source layer are written into the bias assessment report to generate the final interpretable risk score and content moderation recommendations, while also providing traceable evidence to support subsequent model iterations.
[0118] Step 5: Combine the bias risk value of the actual output text, the image stereotype similarity of the actual output image, the cross-modal sensitivity score, and the image-text bias value, and use weighted calculation to calculate the comprehensive bias risk score, and divide it into multiple image-text bias risk levels.
[0119] The specific formula for calculating the weighted comprehensive bias risk score is as follows:
[0120]
[0121] Where BiasRisk is the comprehensive bias risk score; VBR represents bias risk; IS represents image stereotypes; CSS represents cross-modal conflict; MB represents modal bias; z1, z2, z3, and z4 are weighting coefficients, all of which are within the interval (0,1) and z1+z2+z3+z4=1; k=10 and θ=0.5 represent the steepness and offset of the Sigmoid function, respectively.
[0122] Based on the comprehensive bias risk score, four levels of graphic bias risk are defined, as follows:
[0123]
[0124] When the risk value is greater than or equal to 0 and less than 0.3, the RiskLevel is classified as Level 0, which is unbiased (green), meaning the content has no significant risk of bias and can be published or used directly.
[0125] When the risk value is greater than or equal to 0.3 and less than 0.6, the RiskLevel is classified as Level 1, with slight bias (yellow), indicating that the content has a potential bias tendency. It is recommended to manually review and assess the reasonableness of the context.
[0126] When the risk value is greater than or equal to 0.6 and less than 0.8, RiskLevel is classified as Level 2, moderate bias (orange), which contains explicit bias and can only be used after mandatory modification or the addition of a warning label.
[0127] When the risk value is greater than or equal to 0.8 but not more than 1.0, RiskLevel is classified as Level 3, indicating severe bias (red), and the content seriously violates the principle of fairness, and its dissemination is prohibited.
[0128] Example 2
[0129] This embodiment provides a bias risk assessment system for a multimodal large model, including:
[0130] The image and text bias library construction module collects bias label words from public datasets, extracts candidate bias words in the power customer service communication process, and constructs a bias word library; it also collects bias images from public datasets and constructs a bias image library.
[0131] The text bias identification module compares the actual output text of the power company during customer service communication with the baseline vector of bias categories in the bias lexicon, generates a bias risk value for the actual output text, and identifies biased text.
[0132] The image bias recognition module calculates the stereotypical similarity between the actual output image and the image templates in the bias image library within a predetermined time interval before and after the actual output text, and identifies biased images.
[0133] The implicit bias identification module calculates the cross-modal sensitivity scores of the actual output text and the actual output image, and identifies implicit bias by combining the identification of biased text and biased images; based on the attention weights of the multimodal large model on the text and image outputs, it calculates the abnormal layer of the image-text bias value identification model to determine the source of implicit bias.
[0134] The image-text bias risk classification module integrates the bias risk value of the actual output text, the image stereotype similarity of the actual output image, the cross-modal sensitivity score, and the image-text bias value, and uses a weighted calculation to calculate a comprehensive bias risk score, classifying multiple image-text bias risk levels.
[0135] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A method for assessing bias risk in a multimodal large model, characterized in that, Includes the following steps: Step 1: Collect bias label words from public datasets, obtain candidate bias words manually labeled from historical data of power customer service documents, form different bias categories based on candidate bias words and bias label words, and construct a bias lexicon. Collect biased images from publicly available datasets, cluster the biased images to obtain image templates for each biased image cluster, and construct a biased image library; Step 2: Compare the actual output text of the power company during customer service communication with the baseline vector of the bias category in the bias lexicon to generate the bias risk value of the actual output text and identify biased text; Step 3: Calculate the stereotypical similarity between the actual output image and the image template in the biased image library within the predetermined time interval above and below the actual output text to identify biased images; Step 4: Calculate the cross-modal sensitivity scores of the actual output text and the actual output image, and identify implicit bias by combining the identification of biased text and biased images; based on the attention weights of the multimodal large model on the text and image outputs, calculate the anomaly layer of the image-text bias value identification model to determine the source of implicit bias. Step 5: Combine the bias risk value of the actual output text, the image stereotype similarity of the actual output image, the cross-modal sensitivity score, and the image-text bias value, and use weighted calculation to calculate the comprehensive bias risk score, and divide it into multiple image-text bias risk levels.
2. The bias risk assessment method for multimodal large models according to claim 1, characterized in that: The process of building the bias lexicon in step 1 includes: Import publicly available bias vocabulary datasets to obtain bias label words; perform data mining on historical data from electricity customer service documents to collect candidate bias words; A text encoder based on a multimodal large model generates context-dependent word vectors for each bias label word and candidate bias word; all word vectors are clustered to form a bias word cluster. Calculate the mean vector of word vectors within the biased word cluster and use it as the semantic center of the corresponding biased word cluster; extract the keywords of the biased label words corresponding to the biased category from the biased label words, generate the word vectors of the keywords, and calculate the mean of the word vectors of the keywords as the reference vector of the corresponding category; match the semantic center with the reference vector of the biased category to determine the biased category label of each biased word cluster. A bias lexicon is constructed based on bias label words, candidate bias words, bias word cluster numbers, and bias category labels.
3. The bias risk assessment method for multimodal large models according to claim 1, characterized in that: The process of building the biased image library in step 1 includes: Import publicly available biased image databases, collect and label biased images; extract local detail features and global semantic features from biased images to generate comprehensive feature vectors; cluster the comprehensive feature vectors to divide all biased images into multiple biased image clusters and determine the center vector of each cluster. Select a predetermined number of biased images that are closest to the center vector from the cluster as the image templates of the corresponding cluster; calculate the correlation between each image template and the bias words corresponding to each bias category in the bias vocabulary, and take the bias category with the highest correlation as the classification result of the corresponding image template to determine the bias category label of the biased image cluster; A biased image library is constructed based on image templates of biased image clusters and their corresponding bias category labels, comprehensive feature vectors, and center vectors.
4. The bias risk assessment method for multimodal large models according to claim 1, characterized in that: In step 2, the actual output text is preprocessed and then subjected to dynamic semantic encoding to obtain semantic encoding; the baseline vector of each bias category is extracted from the bias lexicon, and the bias intensity of the semantic encoding and the baseline vector of different bias categories is calculated for each word in the actual output text. Calculate the average attention weight of each word to other words in the actual output text as the corresponding word's context dependency score; Based on the bias intensity and context dependence score of each word, the bias risk value of the corresponding actual output text is calculated by combining the word frequency-inverse document frequency in the actual output text. If the bias risk value is not less than the predetermined text bias risk threshold, it means that the corresponding actual output text matches the text bias library and there is bias.
5. The method for assessing bias risk in a multimodal large model according to claim 1, characterized in that: In step 3, the image feature vector of the actual output image is extracted, and its similarity is calculated with the comprehensive feature vector in the bias image library to obtain the image stereotype similarity. If the stereotype similarity of an image is not less than a predetermined image bias risk threshold, it means that the corresponding actual output image matches the biased image library and there is bias.
6. The method for assessing bias risk in a multimodal large model according to claim 1, characterized in that: In step 4, the process of identifying implicit bias includes: Extract the text semantic vector of the actual output text and the image semantic vector of the actual output image, and calculate the cross-modal sensitivity score of both. If the cross-modal sensitivity score is not less than the predetermined first cross-modal bias threshold and is less than the predetermined second cross-modal bias threshold, the existence of implicit bias is determined based on the bias risk value of the actual output text and the image stereotype similarity of the actual output image; if both the actual output text and the actual output image have bias, then the actual output text and the actual output image have implicit bias.
7. The method for assessing bias risk in a multimodal large model according to claim 1 or 6, characterized in that: In step 4, for each cross-modal attention layer of the multimodal large model, the weights pointing to the text and the weights pointing to the image are extracted from the attention matrix output by each attention head of each layer to obtain the attention weights of the text and the image as well as the cross-modal attention weights, and the cross-modal attention ratio is calculated. The results of all heads in each cross-modal attention layer are summed to obtain the total attention weights for the text and image in that layer, respectively. The average cross-entropy of cross-modal attention ratio is used as the modal interaction strength coefficient, and it is combined with the total attention weight of text and image to obtain the image-text bias value. Based on the text-image bias value, the abnormal layer in the cross-modal attention layer is located, and the text-high attention content and image-high attention content in this layer are identified. They are then matched with the bias categories in the bias lexicon and bias image lexicon, respectively. If both matches are successful, the source of the latent bias is the abnormal amplification of bias information by the model, and the abnormal layer is located as the bias source layer. Otherwise, it indicates that the latent bias is caused by model noise.
8. The bias risk assessment method for multimodal large models according to claim 7, characterized in that: The process of locating anomaly layers and identifying textual and image content of high interest includes: Calculate the image-text bias value for each cross-modal attention layer in the multimodal large model to form an image-text bias sequence; calculate the mean and standard deviation of the image-text bias sequence; weight the mean and standard deviation to obtain the first bias threshold, and take the maximum value of the first bias threshold and the predetermined second bias threshold as the predetermined anomaly threshold; when there is an image-text bias value in the image-text bias sequence that is greater than the predetermined anomaly threshold, the layer is considered an anomaly layer; In the anomaly layer, the attention weight distribution of text words and image regions is obtained, and words and regions with attention weights greater than a predetermined weight threshold are identified as the corresponding text high-attention content and image high-attention content.
9. The method for assessing bias risk in a multimodal large model according to claim 7, characterized in that: The process of performing bias category matching includes: Extract the vector of the highly relevant content in the text, calculate the cosine similarity with the baseline vectors of different bias categories in the bias lexicon, and if the similarity exceeds a predetermined similarity threshold, the word is matched with the corresponding bias category. Extract the feature vector of the high-interest content region in the image, calculate the cosine similarity with the comprehensive feature vector of the image template in the bias image library, and if the similarity exceeds a predetermined similarity threshold, the region is determined to match the corresponding bias category.
10. A multimodal large-scale bias risk assessment system, comprising the method described in any one of claims 1-9, characterized in that, include: The image and text bias library construction module collects bias label words from public datasets, extracts candidate bias words in the process of power customer service communication, and constructs a bias word library; Collect biased images from publicly available datasets and build a biased image library; The text bias identification module compares the actual output text of the power company during customer service communication with the baseline vector of bias categories in the bias lexicon, generates a bias risk value for the actual output text, and identifies biased text. The image bias recognition module calculates the stereotypical similarity between the actual output image and the image templates in the bias image library within a predetermined time interval before and after the actual output text, and identifies biased images. The implicit bias identification module calculates the cross-modal sensitivity scores of the actual output text and the actual output image, and identifies implicit bias by combining the identification of biased text and biased images; based on the attention weights of the multimodal large model on the text and image outputs, it calculates the abnormal layer of the image-text bias value identification model to determine the source of implicit bias. The image-text bias risk classification module integrates the bias risk value of the actual output text, the image stereotype similarity of the actual output image, the cross-modal sensitivity score, and the image-text bias value, and uses a weighted calculation to calculate a comprehensive bias risk score, classifying multiple image-text bias risk levels.
Citation Information
Patent Citations
Predictive evaluation method for evaluating image recognition model
CN117437507A
Generative text visual question-answering method and system based on prompt learning
CN117786054A
Large language model natural language reasoning prejudice detection method based on metamorphic test
CN118820093A