Image-text content auditing method based on multi-modal large model
By building a pyramid-shaped labeling system and a phased dynamic adjustment strategy, the problems of pseudo-label generation and labeling system construction in graphic and text content review are solved, efficient and accurate review of the model during training is achieved, the multimodal feature fusion capability is improved, and it adapts to the ever-changing network environment.
Patent Information
- Application Number
- CN202510815715.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Existing image and text content review technologies have defects in pseudo-label generation and screening, label system construction, and multimodal feature fusion, which makes it difficult for the model to quickly establish basic classification capabilities in the early stages of training, and is unable to further improve the review accuracy in the later stages of training. It also fails to fully utilize the complex relationship between image and text features, resulting in insufficient recognition capabilities.
Build a pyramid-style labeling system and implement a phased dynamic adjustment strategy. Through multimodal pseudo-label generation, pseudo-label quality assessment and screening, labeling system construction and multi-granularity comparative learning and training, use a large multimodal model to fuse image and text features, adjust the label granularity in stages, and improve the model's audit performance from coarse to fine.
It significantly improves the accuracy and efficiency of image and text content review, shortens the training cycle, reduces computing resource consumption, and enhances the model's ability to identify comprehensive illegal content in images and texts, adapting to the image and text content review needs of different fields and scenarios.
Smart Images

Figure CN120673211A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graphic content auditing, and specifically to a graphic content auditing method based on a multimodal large model. Background Art
[0002] With the development of the Internet, the speed and scale of the dissemination of graphic and text content have increased exponentially. However, a large amount of graphic and text content may contain illegal information, posing a serious threat to the physical and mental health of users, social order and platform ecology. Therefore, efficient and accurate graphic and text content review technology has become an urgent need for Internet platforms and regulatory agencies. Traditional manual review methods are inefficient, costly, and difficult to cope with massive amounts of data. Automated review methods based on rules or a single modality have problems with poor generalization ability and high error rates. At present, large multimodal models have shown significant advantages in graphic and text content understanding tasks by integrating image and text features.
[0003] When faced with unlabeled data, existing image and text content review technologies usually directly use pre-trained models to generate pseudo labels, but lack an effective pseudo-label quality assessment mechanism, resulting in low-quality pseudo labels being mixed into the training data, affecting model performance. In terms of label system construction, most methods use single-granularity labels and cannot dynamically adjust the degree of label refinement according to the model training stage, making it difficult for the model to quickly establish basic classification capabilities in the early stages of training, and unable to further improve the review accuracy in the later stages of training. In addition, existing multimodal training methods fail to fully utilize the advantages of labels of different granularities and find it difficult to effectively integrate the complex correlation between image and text features, resulting in insufficient model recognition of comprehensive illegal content in images and texts.
[0004] In summary, the existing graphic and text content review technology has obvious defects in pseudo-label generation and screening, label system construction, and multimodal feature fusion training, and cannot meet the growing demand for network content review. Therefore, there is a need for a graphic and text content review method that can efficiently process unlabeled data, dynamically adjust label granularity, and fully integrate multimodal features to improve review accuracy and efficiency and adapt to the ever-changing network environment. Summary of the Invention
[0005] The purpose of the present invention is to make up for the shortcomings of the existing technology and provide a graphic content review method based on a multimodal large model. It can significantly improve the accuracy and efficiency of graphic content review by constructing a pyramid label system and implementing a phased dynamic adjustment strategy. In the early stage of training, coarse-grained labels are used to quickly establish the basic classification ability of the model, avoiding the low training efficiency and overfitting problems caused by the premature introduction of fine-grained labels in traditional methods. When medium-grained labels are introduced in the middle stage, the model can further refine the identification of violation types, especially improving the recognition accuracy of easily confused violation categories. When fine-grained labels are used in the later stage, the model's ability to identify specific violation elements is significantly enhanced, and it can accurately locate subtle violation content. This label system adjustment mechanism from coarse to fine not only improves the review performance of the model, but also shortens the model training cycle and reduces computing resource consumption.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for reviewing graphic content based on a multimodal large model, wherein the specific steps of the method are as follows:
[0007] S100, multimodal pseudo-label generation: Generate pseudo-labels for unlabeled image and text data using a multimodal large model, fusing image features with text features to generate an initial pseudo-label set;
[0008] S200, Pseudo-label quality assessment and screening: Design confidence assessment indicators, perform quality scoring on pseudo-labels, screen high-confidence pseudo-labels, and build a high-quality pseudo-label dataset for subsequent training;
[0009] S300, labeling system construction: Divide the S200 high-quality pseudo-labeled dataset into three levels and dynamically adjust the label granularity in stages;
[0010] S400, Multi-granularity Contrastive Learning Training: Leveraging the high-quality pseudo-labeled dataset and label system adjustment information from S200, combined with contrastive learning strategies, semi-supervised training of large multimodal models is performed.
[0011] S500, model iterative optimization: Input the graphic and text content to be reviewed into the trained multimodal large model. The model outputs the review results based on the learned knowledge and features to complete the graphic and text content review.
[0012] Furthermore, the S100 specifically includes:
[0013] The pre-trained visual encoder is used to extract features from images in unlabeled image and text data, extracting the global and local features of the image to form the visual feature vector of the image;
[0014] The pre-trained text encoder is used to extract features from the unlabeled text data, extract the syntactic features of the text, and form a semantic feature vector of the text;
[0015] The visual feature vector of the image and the semantic feature vector of the text are weightedly fused through the attention mechanism. Based on the multimodal features after weighted fusion, an initial pseudo-label set is generated through the prediction layer of the multimodal large model. Each pseudo-label in the initial pseudo-label set contains the image violation category, the text violation category, and the comprehensive violation judgment result of the image and text.
[0016] Furthermore, the S200 designs a confidence evaluation index to score the quality of pseudo labels: C(L p )=α×P p (L p )+β×Sim f (L p )+γ×Sim c (L p ), where C(L p ) represents the pseudo label L p The confidence score of the pseudo label is in the range of [0,1]. The higher the score, the higher the credibility of the pseudo label. p (L p ) is the pseudo label L predicted by the multimodal large model p The probability value range is [0,1]. This value is output by the prediction layer of the multimodal large model and reflects the degree of certainty of the model's prediction of the pseudo label. Sim f (L p ) is the pseudo label L p The similarity between the corresponding image and text features and the annotated similar image and text features is calculated using the cosine similarity formula, with a value range of [0,1] and the formula is: F p is the pseudo label L p The corresponding image-text fusion feature vector, F s is the fusion feature vector of the annotated similar images and texts, ||·|| represents the modulus of the vector, Sim c (L p ) is the pseudo label L p The cross-modal consistency score of the image violation category and the text violation category in the range of [0,1] is obtained by calculating the cosine similarity of the feature vectors of the image and text violation categories. α, β, γ are weight coefficients, satisfying α+β+γ=1 and α, β, γ≥0, which are used to adjust the importance of different evaluation dimensions. The confidence score C(L p )>0.8 to construct a high-quality pseudo-label dataset.
[0017] Furthermore, the three levels of S300 are: coarse granularity, medium granularity, and fine granularity, where:
[0018] Coarse-grained label generation: For high-quality pseudo-label data, use a multi-modal large model to predict the major categories of violations as coarse-grained high-confidence samples. Input them into the multi-modal large model to establish the model foundation and calculate the model accuracy rate.
[0019] Medium-grained label generation: For the coarse-grained high-confidence samples, refine the prediction of violation behaviors as medium-grained high-confidence samples. When the accuracy rate of the model in the contrastive learning training reaches the first preset value, introduce medium-grained labels to refine the violation types.
[0020] Fine-grained label generation: For the medium-grained high-confidence samples, locate the specific violation elements. When the accuracy rate reaches the second preset value and tends to be stable, use fine-grained labels to improve the audit accuracy.
[0021] Furthermore, in S300, set the current training round as n, the accuracy rate of the model after the nth round of training as A(n), define the first preset accuracy rate threshold as A1, the second preset accuracy rate threshold as A2, and 0 < A1 < A2 < 1; at the same time, set the accuracy rate fluctuation coefficient k to measure the stability of the accuracy rate. The smaller k is, the more stable it is. The accuracy rate The And when n = 1, then k(1) = 0, where A(n) represents the accuracy rate of the model after the nth round of training, and its value range is [0, 1]. The higher this value is, the higher the accuracy of the model's judgment on the graphic and text content audit after this round of training. A(n - 1) represents the accuracy rate of the model after the (n - 1)th round of training. TP(n) represents the number of graphic and text samples that the model correctly predicts as violations after the nth round of training, that is, the number of samples that are actually violations and the model also judges as violations. TN(n) represents the number of graphic and text samples that the model correctly predicts as compliant after the nth round of training, that is, the number of samples that are actually compliant and the model also judges as compliant. FP(n) represents the number of graphic and text samples that the model wrongly predicts as violations after the nth round of training, that is, the number of samples that are actually compliant but the model misjudges as violations. FN(n) represents the number of graphic and text samples that the model wrongly predicts as compliant after the nth round of training, that is, the number of samples that are actually violations but the model misjudges as compliant. k(n) is the accuracy rate fluctuation coefficient at the nth round of training. The label granularity level is represented by Level, and n ∈ {1, 2, 3}, where 1 represents coarse-grained, 2 represents medium-grained, and 3 represents fine-grained.
[0022] Furthermore, the specific dynamic adjustment of the label granularity in S300 is as follows:
[0023] When n ≤ N1 and A(n) < A1: Level(n) = 1, and at this time, use coarse-grained labels to establish the basic classification ability of the model. N1 is the set initial number of training rounds.
[0024] When A1 ≤ A(n) < A2 and k(n) ≤ K1: Level(n) = 2, introduce medium-grained labels to refine the identification of violation types. K1 is the stable threshold for medium-grained switching and K1 = 0.05. And when k(n) > K1, then maintain Level(n) = 1;
[0025] When A(n) ≥ A2 and k(n) ≤ K2: Level(n) = 3, use fine-grained labels to improve the audit accuracy. And when K2 < k(n) < K1, then maintain Level(n) = 2 to ensure the model improves performance in a stable state. K2 is the stable threshold for fine-grained switching and K2 = 0.03;
[0026] After each completion of the adjustment of the label granularity level, the current Level(n) is fed back as label system adjustment information to S400.
[0027] Furthermore, the semi-supervised training steps of the S400 are as follows:
[0028] Obtain the labeled sample set as S l , each sample in this set has an artificially labeled tag, which is used to provide supervision information for the model. The high-quality pseudo-labeled data set is S p , as supplementary data for the labeled samples, receive the label system adjustment information from S300, obtain the current label granularity level Level(n). For any sample x l in the labeled sample set S i , respectively obtain the visual feature vector of its image and the semantic feature vector of the text . After feature fusion, obtain the comprehensive feature vector . The corresponding accurate label is y i . For any sample x p in the high-quality pseudo-labeled data set S j , also obtain its fused comprehensive feature vector . The corresponding pseudo-label is y j ;
[0029] Adopt the contrastive learning loss function L c to perform semi-supervised training on the multi-modal large model. The loss function where N represents the number of samples included in each training batch, k represents the k-th sample in the current batch, sim(·,·) is the cosine similarity calculation function, which is used to measure the similarity between two feature vectors, f <000p Find the same as x k Tag y k For the same sample, obtain its feature vector as Representative and sample x k Negative sample feature vectors of different categories, from S l and S p Filter out labels and y k Different samples, obtain their feature vectors as negative sample feature vectors M represents the number of negative samples, m represents the sequence number of negative samples, ω(y k ,Level) is related to the sample label y k The weight coefficient related to the current label granularity level Level, τ is a hyperparameter with an initial value of 0.1, and is dynamically adjusted according to the label granularity level Level: when Level(n) = 1, τ = 0.1, when Level(n) = 2, τ = 0.05, when Level(n) = 3, τ = 0.03.
[0030] Furthermore, during the multimodal large model training process, S300 assigns different weights to samples according to the current label granularity level Level(n):
[0031] When Level(n)=1, all samples are given the same weight, that is, each sample contributes the same amount to the loss function, that is, ω(y k ,Level)=1;
[0032] When Level(n)=2, y k Belongs to the medium-granularity label category, then ω(y k ,Level)=1.5, otherwise 1;
[0033] When Level=3, y k belongs to the fine-grained label category, then ω(y k ,Level)=2.0, otherwise it is 1.
[0034] Compared with existing technologies, this method for reviewing text and image content based on a multimodal large model has the following beneficial effects:
[0035] 1. The present invention significantly improves the accuracy and efficiency of graphic content review by constructing a pyramid label system and implementing a phased dynamic adjustment strategy. In the early stage of training, coarse-grained labels are used to quickly establish the basic classification capability of the model, avoiding the low training efficiency and overfitting problems caused by the premature introduction of fine-grained labels in traditional methods. When medium-grained labels are introduced in the middle stage, the model can further refine the identification of violation types, especially improving the recognition accuracy of easily confused violation categories. When fine-grained labels are used in the later stage, the model's ability to identify specific violation elements is significantly enhanced, and it can accurately locate subtle violation content. This coarse-to-fine label system adjustment mechanism not only improves the review performance of the model, but also shortens the model training cycle and reduces computing resource consumption.
[0036] 2. The multi-granularity comparative learning training method proposed in the present invention effectively solves the shortcomings of the existing technology in multimodal feature fusion, and greatly improves the model's ability to identify comprehensive illegal content in pictures and texts. Through the comparative learning strategy, the model can learn the deep correlation information between image and text features, so that the picture and text content can be more accurately represented in the feature space. When processing unlabeled data, the method of the present invention fully utilizes the value of a large amount of unlabeled data through high-quality pseudo-label screening and dynamic weight adjustment. When the labeled data is limited, the model performance can still maintain a stable improvement, and the label granularity and training strategy can be flexibly adjusted according to actual needs. It is suitable for picture and text content review in different fields and scenarios.
[0037] Other advantages, objects and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be learned from the practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0039] Figure 1 This is an operational flowchart of the graphic content review method based on a multimodal large model;
[0040] Figure 2 This is a step-by-step framework diagram of the graphic content review method based on a multimodal large model. DETAILED DESCRIPTION
[0041] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0042] Example 1
[0043] This embodiment provides a method for reviewing text content based on a multimodal large model, such as Figure 1 As shown in the figure, this method achieves efficient and accurate review of graphic and text content through the steps of multimodal pseudo-label generation, pseudo-label quality assessment and screening, label system construction, multi-granularity comparative learning training, and model iterative optimization. Its core lies in using a large multimodal model to fuse image and text features to generate pseudo-labels, screening high-quality pseudo-labels through confidence evaluation indicators, constructing a three-level label system with dynamic adjustment in stages, and combining it with a comparative learning strategy for semi-supervised training, thereby significantly improving the model's recognition accuracy and generalization ability for illegal graphic and text content, and is suitable for complex graphic and text content review scenarios on Internet platforms.
[0044] First, enter the multimodal pseudo-label generation stage (S100). In this stage, with the help of pre-trained visual encoders and text encoders, the images and texts in the unlabeled graphic and text data are respectively extracted. Then, the features are fused through the attention mechanism and pseudo labels are generated by the multimodal large model. For images in the unlabeled graphic and text data, the pre-trained visual encoder is used to extract their global features and local features. The global features can reflect the overall semantic information of the image. These two types of features together constitute the visual feature vector of the image, which comprehensively describes the visual information of the image. For text data, the pre-trained text encoder is used to extract its syntactic features. The syntactic features can reflect the language information of the text. The legal structure and semantic relationship are combined to form the semantic feature vector of the text to represent the semantic content of the text. The attention mechanism is used to perform weighted fusion on the visual feature vector of the image and the semantic feature vector of the text. The attention mechanism can automatically learn the degree of correlation between image and text features, and assign different weights to different feature components, so that the model pays more attention to the features that are important for violation judgment. Based on the multimodal features after weighted fusion, the initial pseudo-label set is generated through the prediction layer of the multimodal large model. Each pseudo-label contains the image violation category, text violation category and the comprehensive violation judgment result of the image and text, thereby realizing multi-dimensional prediction of violations of image and text content.
[0045] Then, the pseudo-label quality evaluation and screening stage (S200) is entered. In this stage, a confidence evaluation index is designed to score the quality of the pseudo-label. The credibility of the pseudo-label is comprehensively evaluated by considering multiple dimensions such as model prediction probability, image-text feature similarity, and cross-modal consistency. The pseudo-label quality scoring formula is C(L p )=α×P p (Lp )+β×Sim f (L p )+γ×Sim c (L p ), where C(L p ) represents the pseudo label L p The confidence score of P is in the range of [0, 1]. The higher the score, the more reliable the pseudo label. p (L p ) is the pseudo label L predicted by the multimodal large model p The probability value is output by the prediction layer of the model, which reflects the degree of certainty of the model's prediction of the pseudo label. The higher the probability value, the more accurate the model's prediction result. f (L p ) is the pseudo label L p The similarity between the corresponding image and text features and the annotated similar image and text features is calculated using the cosine similarity formula Calculated, where F p is the pseudo label L p The corresponding image-text fusion feature vector, F s is the fusion feature vector of the labeled similar images and texts, ||·|| represents the modulus of the vector. The similarity reflects the similarity between the image and text features corresponding to the pseudo-label and the known labeled data features. The higher the similarity, the higher the reliability of the pseudo-label. Sim c (L p ) is the pseudo label L p The cross-modal consistency score between the image violation category and the text violation category is calculated by calculating the cosine similarity of the feature vectors of the image and text violation categories. It is used to measure the consistency between the image and text violation judgments. The higher the consistency, the more reasonable the pseudo-label. α, β, and γ are weight coefficients, satisfying α+β+γ=1 and α, β, γ ≥ 0. They are used to adjust the importance of different evaluation dimensions and can be set according to actual needs. This confidence evaluation metric is used to calculate the score of each pseudo-label. Pseudo-labels with confidence scores greater than the set threshold are screened out to construct a high-quality pseudo-label dataset, providing reliable samples for subsequent training.
[0046] Next, enter the label system construction stage (S300). Divide the high-quality pseudo-label dataset into three levels: coarse-grained, medium-grained, and fine-grained, and dynamically adjust the label granularity in stages. The principle is to gradually refine the label granularity according to the accuracy and stability during the model training process, so that the model can adapt to different learning difficulties at different training stages, thereby improving the training effect. Among them, when generating coarse-grained labels, for high-quality pseudo-label data, use a multi-modal large model to predict the major categories of violations, and use the coarse-grained high-confidence samples as input to the multi-modal large model to establish the model foundation and calculate the model accuracy. In the initial stage of training, the model's ability to identify violation content is weak. The coarse-grained major categories of violations can enable the model to quickly master the basic classification framework and avoid the model being difficult to learn due to overly fine labels. When generating medium-grained labels, refine the prediction of violation behaviors for the coarse-grained high-confidence samples as medium-grained high-confidence samples. When the accuracy of the model reaches the first preset value during the contrastive learning training, introduce medium-grained labels to refine the violation types. At this time, the model already has a certain basic classification ability. Medium-grained labels can enable the model to further distinguish more specific violation behaviors and improve the recognition accuracy of violation types. When generating fine-grained labels, locate specific violation elements for the medium-grained high-confidence samples. When the accuracy reaches the second preset value and tends to be stable, use fine-grained labels to improve the audit accuracy. At this time, the performance of the model is relatively stable. Fine-grained labels can enable the model to accurately locate specific violation elements in the text and image content and achieve more refined audits. To more accurately judge when to adjust the label granularity, set the current training round as n, the accuracy of the model after the nth round of training as A(n), define the first preset accuracy threshold as A1, the second preset accuracy threshold as A2, and 0 < A1 < A2 < 1. At the same time, set the accuracy fluctuation coefficient k to measure the stability of the accuracy. The smaller k is, the more stable it is. Among them, the accuracy TP(n) represents the number of text and image samples that the model correctly predicts as violations after the nth round of training, that is, the number of samples that are actually violations and the model also judges as violations. TN(n) represents the number of text and image samples that the model correctly predicts as compliant after the nth round of training, that is, the number of samples that are actually compliant and the model also judges as compliant. FP(n) represents the number of text and image samples that the model incorrectly predicts as violations after the nth round of training, that is, the number of samples that are actually compliant but the model misjudges as violations. FN(n) represents the number of text and image samples that the model incorrectly predicts as compliant after the nth round of training, that is, the number of samples that are actually violations but the model misjudges as compliant. This formula accurately reflects the accuracy of the model's audit judgment after this round of training by comprehensively considering the correct and incorrect predictions of the model. The accuracy fluctuation coefficient When n = 1, k(1) = 0, which is used to measure the change range of accuracy between adjacent training rounds, so as to judge the stability of model training; the specific rule for dynamically adjusting the label granularity in stages is as follows: when n ≤ N1 and A(n) < A1, Level(n) = 1, and coarse-grained labels are used to establish the basic classification ability of the model. Here, N1 is the set number of initial training rounds. In the initial stage of training, the model needs to first master the basic classification framework, and coarse-grained labels can meet this requirement, avoiding the model being trapped in complex learning tasks due to overly fine labels, resulting in low training efficiency. When A1 ≤ A(n) < A2 and k(n) ≤ K1, Level(n) = 2, and medium-grained labels are introduced to refine the identification of violation types. Here, K1 is the stability threshold for medium-grained switching. At this time, the accuracy of the model has reached a certain level, and the training tends to be stable. Introducing medium-grained labels allows the model to further learn more specific violation types and improve the identification ability. If k(n) > K1, then maintain Level(n) = 1 because the accuracy fluctuates greatly at this time and the model training is not stable enough. Continuing to use coarse-grained labels can make the model stable first, avoiding fluctuations in model performance caused by introducing more fine-grained labels. When A(n) ≥ A2 and k(n) ≤ K2, Level(n) = 3, and fine-grained labels are used to improve the audit accuracy. Here, K2 is the stability threshold for fine-grained switching and K2 < K1. At this time, the accuracy of the model is relatively high, and the training is very stable. Fine-grained labels can enable the model to accurately locate violation elements and achieve more refined audits. If K2 < k(n) < K1, then maintain Level(n) = 2 to ensure that the model improves its performance in a stable state and avoid rashly increasing the label granularity due to fluctuations in accuracy, which may affect the training effect of the model. After each adjustment of the label granularity level, the current Level(n) is fed back as label system adjustment information to the subsequent multi-granularity contrast learning training stage (S400), so that this stage can perform corresponding training according to the current label granularity.
[0047] After that, enter the multi-granularity contrast learning training stage (S400). Using the labeled sample set and the high-quality pseudo-label data set, combined with the contrast learning strategy and label system adjustment information, perform semi-supervised training on the multi-modal large model, so as to improve the model's ability to fuse text and image features and the ability to identify violation content. Specifically: obtain the labeled sample set S l , each sample in this set has an artificial annotation label, which is used to provide supervision information for the model to ensure that the model learns in the correct direction. The high-quality pseudo-label data set S p is used as supplementary data for the labeled samples, making full use of the value of a large amount of unlabeled data to expand the training sample size of the model. Receive the label system adjustment information from the label system construction stage, obtain the current label granularity level Level(n). For any sample x in the labeled sample set S l i , respectively obtain the visual feature vector of its image and the semantic feature vector of the text After feature fusion, the comprehensive feature vector is obtained The corresponding accurate label is y i , for high-quality pseudo-label dataset S p Any sample x in j , and also obtain the integrated feature vector after fusion The corresponding pseudo label is y j , using contrastive learning loss function L c Semi-supervised training of the multimodal large model is performed, and the loss function formula is specifically: Where N is the number of samples in each training batch, k is the kth sample in the current batch, sim(·,·) is the cosine similarity calculation function, which is used to measure the similarity between two feature vectors. By calculating the cosine value of the feature vector, we can judge their directional consistency and thus reflect the similarity of the features. k is the feature vector of the kth sample, Representative and sample x k The positive sample feature vector of the same category is constructed from S l and S p Find the same as x k Tag y k For the same sample, obtain its feature vector as This allows the model to learn the common features of similar samples. Representative and sample x k Negative sample feature vectors of different categories, from S l and S p Filter out labels and y k Different samples, obtain their feature vectors as negative sample feature vectors By comparing with negative samples, the model can distinguish the feature differences between different categories. M represents the number of negative samples, ω(y k ,Level(n)) is related to the sample label y kThe weight coefficient related to the current label granularity level Level is used to adjust the contribution of samples in the loss function according to different label granularities and label categories. τ is a hyperparameter with an initial value of 0.1 and is dynamically adjusted according to the label granularity level Level: when Level(n)=1, τ=0.1, when Level(n)=2, τ=0.05, when Level(n)=3, τ=0.03. During the training of the multimodal large model, different weights are given to samples according to the current label granularity level Level(n): when Level(n)=1, the model is in the early stage of training, and the main task is to establish basic classification capabilities. Therefore, all samples are given the same weight, that is, ω(y k ,Level)=1, each sample contributes equally to the calculation of the loss function, ensuring that the model can fully learn the basic characteristics of various samples. When Level(n)=2, the model has a certain basic classification ability and needs to further refine the violation type identification. Therefore, for samples y belonging to the medium-granularity label category, k , giving ω(y k ,Level)=1.5, and for other samples it is still 1, so that the model can pay more attention to the samples of medium-grained label categories and strengthen the learning of violation types. When Level=3, the model needs to accurately locate the specific violation elements. For samples belonging to the fine-grained label category y k Assign ω(y k ,Level)=2.0, and other samples are 1, so that the model focuses more on samples with fine-grained label categories and improves the ability to identify specific illegal elements. Through this dynamic weight adjustment, the model can learn the characteristics of key samples in a targeted manner according to the label granularity and task requirements at different training stages, thereby improving the training effect.
[0048] Finally, the model iteration optimization stage (S500) is entered, and the image and text content to be reviewed is input into the trained multimodal large model. The model analyzes and processes the input image and text content based on the knowledge and features learned in the multi-granularity comparative learning training stage. The model will first extract the visual features of the image and the semantic features of the text, and then fuse these features through the attention mechanism. Then, based on the fused multimodal features and the learned violation pattern, the violation category of the image and text content is predicted, and the image violation category, text violation category and comprehensive image and text violation judgment result are output to complete the image and text content review. In the actual application process, if it is found that the review effect of the model does not meet expectations, new image and text content to be reviewed can be collected, and the multimodal pseudo-label generation stage (S100) can be re-entered to generate new pseudo-labels. The above training and optimization process is repeated to iteratively update the model, and continuously improve the review accuracy and performance of the model to adapt to the ever-changing image and text content violation forms and review requirements.
[0049] To sum up, this embodiment constructs a complete method for text and image content review based on a multimodal large model through the steps of multimodal pseudo-label generation, pseudo-label quality assessment and screening, label system construction, multi-granularity comparative learning training, and model iterative optimization. This method screens high-quality pseudo-labels through multi-dimensional confidence assessment indicators to ensure the reliability of training data; the constructed three-level label system with phased dynamic adjustment enables the model to gradually improve the review accuracy at different training stages, avoiding low training efficiency and overfitting problems, and combines the comparative learning strategy with the semi-supervised training method of dynamic weight adjustment to effectively integrate image and text features, thereby improving the model's ability to identify comprehensive illegal content in pictures and texts. Through the method of this embodiment, the text and image content review task can be completed efficiently and accurately.
[0050] Example 2
[0051] like Figure 1 As shown, this embodiment provides a workflow for image and text content review based on a multimodal large model. The specific steps of the process are:
[0052] Input unlabeled image and text data, use the pre-trained visual encoder to process the image part of the unlabeled image and text data, and extract the global and local features of the image;
[0053] Use pre-trained text encoders to process the text portion of unlabeled image and text data and extract syntactic features of the text;
[0054] The image features and text features are weighted and fused through the attention mechanism to generate a comprehensive multimodal feature representation;
[0055] Based on the fused multimodal features, the prediction layer of the multimodal large model is used to generate an initial set of pseudo-labels. Each pseudo-label includes the image violation category, the text violation category, and the combined image and text violation judgment result.
[0056] Design a confidence evaluation metric that comprehensively considers the certainty of model predictions, feature similarity, and cross-modal consistency;
[0057] Assign a quality score to each pseudo-label based on the model’s predicted probability, feature similarity with the annotated data, and consistency between image and text violation categories.
[0058] Filter out pseudo labels with a confidence score higher than 0.8 as high-quality pseudo label datasets for subsequent training;
[0059] Discard low-confidence pseudo-labels to ensure the reliability of training data;
[0060] Divide high-quality pseudo-label datasets into three levels: coarse-grained, medium-grained, and fine-grained;
[0061] In the early stages of training, coarse-grained labels are used to establish the model's basic classification capabilities and monitor the model's accuracy.
[0062] When the model accuracy reaches the first preset value and the accuracy fluctuation is small, switch to medium-granularity labeling to refine the violation type identification;
[0063] When the model accuracy reaches the second preset value and the accuracy is stable, switch to fine-grained labeling to improve audit accuracy;
[0064] Dynamically adjust the label granularity based on the number of training rounds, accuracy, and fluctuation coefficient, and feed the current granularity level information back to the training phase;
[0065] Collect labeled sample datasets (manually labeled) and high-quality pseudo-label datasets as training input;
[0066] Receive the current granularity level information (coarse, medium or fine granularity) from the label system construction phase;
[0067] Extract image and text features for each sample and fuse them into a comprehensive feature vector;
[0068] Construct positive samples (samples with the same violation category as the current sample) and negative samples (samples with different violation categories from the current sample);
[0069] Assign sample weights based on the current granularity level. When the granularity is coarse, all samples have the same weight. When the granularity is medium, the weight of medium-grained samples is increased. When the granularity is fine, the weight of fine-grained samples is increased.
[0070] Use contrastive learning strategies for semi-supervised training to optimize large multimodal models and strengthen the association between image and text features;
[0071] Input the image and text content to be reviewed into the trained multimodal model;
[0072] The model automatically analyzes image and text content based on learned multimodal features and knowledge;
[0073] Output audit results, including violation categories (such as image violations, text violations, or combined image and text violations) or compliance judgments;
[0074] Use the results in the actual content review system to complete the review process.
[0075] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A graphic content review method based on a multimodal large model, characterized by: The specific steps of this method are as follows: S100. Multi-modal pseudo-label generation: Use a multi-modal large model to generate pseudo-labels for unlabeled image-text data, fuse image features and text features, and generate an initial pseudo-label set; S200. Pseudo-label quality evaluation and screening: Design a confidence evaluation index, score the quality of pseudo-labels, screen high-confidence pseudo-labels, and construct a high-quality pseudo-label data set for subsequent training; S300. Label system construction: Divide the high-quality pseudo-label data set in S200 into three levels, and dynamically adjust the label granularity in stages; S400. Multi-granularity contrastive learning training: Use the high-quality pseudo-label data set in S200 and the label system adjustment information, and combine the contrastive learning strategy to perform semi-supervised training on the multi-modal large model; S500. Model iterative optimization: Input the image-text content to be reviewed into the trained multi-modal large model, and the model outputs the review result according to the learned knowledge and features to complete the review of the image-text content.
2. The method for reviewing graphic content based on a multimodal large model according to claim 1 is characterized in that: The specific content of S100 includes: Extract the features of the images in the unlabeled image-text data through a pre-trained visual encoder, extract the global features and local features of the images, and form the visual feature vector of the images; Extract the features of the text in the unlabeled text data through a pre-trained text encoder, extract the syntactic features of the text, and form the semantic feature vector of the text; Weight and fuse the visual feature vector of the image and the semantic feature vector of the text through an attention mechanism, and generate an initial pseudo-label set through the prediction layer of the multi-modal large model based on the weighted-fused multi-modal features. Each pseudo-label in the initial pseudo-label set includes an image violation category, a text violation category, and a comprehensive image-text violation judgment result.
3. The method for reviewing graphic content based on a multimodal large model according to claim 1 is characterized in that: The S200 designs a confidence evaluation index to score the quality of pseudo labels: C(L p )=α×P p (L p )+β×Sim f (L p )+γ×Sim c (L p ), where C(L p ) represents the pseudo label L p The confidence score ranges from [0,1], P p (L p ) is the pseudo label L predicted by the multimodal large model p The probability value range is [0,1], Sim f (L p ) is the pseudo label L p The similarity between the corresponding image and text features and the annotated similar image and text features is calculated using the cosine similarity formula, with a value range of [0,1] and the formula is: F p is the pseudo label L p The corresponding image-text fusion feature vector, F s is the fusion feature vector of the annotated similar images and texts, ||·|| represents the modulus of the vector, Sim c (L p ) is the pseudo label L p The cross-modal consistency score of the image violation category and the text violation category in the range of [0,1] is obtained by calculating the cosine similarity of the feature vectors of the image and text violation categories. α, β, γ are weight coefficients, satisfying α+β+γ=1 and α, β, γ≥0, which are used to adjust the importance of different evaluation dimensions. The confidence score C(L p )>0.8 to construct a high-quality pseudo-label dataset.
4. The method for reviewing graphic content based on a multimodal large model according to claim 1 is characterized in that: The three levels of S300 are: coarse granularity, medium granularity, and fine granularity, where: Coarse-granularity label generation: For high-quality pseudo-label data, use the multi-modal large model to predict the major violation categories as coarse-granularity high-confidence samples, input them into the multi-modal large model to establish the model foundation, and calculate the model accuracy; Medium-granularity label generation: Refine the predicted violation behaviors for the coarse-granularity high-confidence samples as medium-granularity high-confidence samples. When the accuracy of the model reaches the first preset value in the contrastive learning training, introduce medium-granularity labels to refine the violation types; Fine-granularity label generation: Locate the specific violation elements for the medium-granularity high-confidence samples. When the accuracy reaches the second preset value and tends to be stable, use fine-granularity labels to improve the review accuracy.
5. The method for reviewing graphic content based on a multimodal large model according to claim 1 is characterized in that: The S300 sets the current training round as n, the accuracy of the model after the n-th round of training is A(n), the first preset accuracy threshold is defined as A1, the second preset accuracy threshold is defined as A2, and 0 < A1 < A2 < 1; at the same time, an accuracy fluctuation coefficient k is set to measure the stability of the accuracy. The accuracy The And when n = 1, then k(1) = 0, where A(n) represents the accuracy of the model after the n-th round of training, and its value range is [0, 1], A(n - 1) represents the accuracy of the model after the (n - 1)-th round of training, TP(n) represents the number of graphic and text samples that the model correctly predicts as illegal after the n-th round of training, that is, the number of samples that are actually illegal and the model also judges as illegal, TN(n) represents the number of graphic and text samples that the model correctly predicts as compliant after the n-th round of training, that is, the number of samples that are actually compliant and the model also judges as compliant, FP(n) represents the number of graphic and text samples that the model wrongly predicts as illegal after the n-th round of training, that is, the number of samples that are actually compliant but the model misjudges as illegal, FN(n) represents the number of graphic and text samples that the model wrongly predicts as compliant after the n-th round of training, that is, the number of samples that are actually illegal but the model misjudges as compliant, k(n) is the accuracy fluctuation coefficient at the n-th round of training, the label granularity level is represented by Level, and n ∈ {1, 2, 3}, where 1 represents coarse granularity, 2 represents medium granularity, and 3 represents fine granularity.
6. The method for reviewing graphic content based on a multimodal large model according to claim 5 is characterized in that: The specific content of dynamically adjusting the label granularity in stages in S300 is: When n ≤ N1 and A(n) < A1: Level(n) = 1. At this time, use coarse-granularity labels to establish the basic classification ability of the model. N1 is the set initial number of training rounds; When A1 ≤ A(n) < A2 and k(n) ≤ K1: Level(n) = 2. Introduce medium-granularity labels to refine the recognition of violation types. K1 is the stable threshold for medium-granularity switching and K1 = 0.
05. And when k(n) > K1, maintain Level(n) = 1; When A(n) ≥ A2 and k(n) ≤ K2: Level(n) = 3, use fine-grained tags to improve the audit accuracy, and when K2 < k(n) < K1, then maintain Level(n) = 2, K2 is the stable threshold for fine-grained switching and K2 = 0.03; After each completion of the adjustment of the label granularity level, the current Level(n) is fed back to S400 as label system adjustment information.
7. The method for reviewing graphic content based on a multimodal large model according to claim 1 is characterized in that: The semi-supervised training steps of the said S400 are: Get the labeled sample set S l , each sample in the set has a manual annotation label to provide supervision information for the model, and the high-quality pseudo-label dataset is S p , as supplementary data for the labeled samples, receives label system adjustment information from S300, obtains the current label granularity level Level(n), and for the labeled sample set S l Any sample x in i , respectively obtain the visual feature vector of its image and the semantic feature vector of the text After feature fusion, the comprehensive feature vector is obtained The corresponding accurate label is y i , for high-quality pseudo-label dataset S p Any sample x in j , and also obtain the integrated feature vector after fusion The corresponding pseudo label is y j ; Using contrastive learning loss function L c Semi-supervised training of multimodal large models, the loss function Where N is the number of samples in each training batch, k is the kth sample in the current batch, sim(·,·) is the cosine similarity calculation function, which is used to measure the similarity between two feature vectors, and f k is the feature vector of the kth sample, Representative and sample x k The positive sample feature vector of the same category is constructed from S l and S p Find the same as x k Tag y k For the same sample, obtain its feature vector as Representative and sample x k Negative sample feature vectors of different categories, from S l and S p Filter out labels and y k Different samples, obtain their feature vectors as negative sample feature vectors M represents the number of negative samples, m represents the sequence number of negative samples, ω(y k ,Level) is related to the sample label y k The weight coefficient related to the current label granularity level Level, τ is a hyperparameter with an initial value of 0.1, and is dynamically adjusted according to the label granularity level Level: when Level(n) = 1, τ = 0.1, when Level(n) = 2, τ = 0.05, when Level(n) = 3, τ = 0.
03.
8. The method for reviewing graphic content based on a multimodal large model according to claim 7 is characterized in that: During the training process of the multi-modal large model by the said S300, different weights are assigned to the samples according to the current label granularity level Level(n): When Level(n)=1, all samples are given the same weight, that is, each sample contributes the same amount to the loss function, that is, ω(y k ,Level)=1; When Level(n)=2, y k Belongs to the medium-granularity label category, then ω(y k ,Level)=1.5, otherwise 1; When Level=3, y k belongs to the fine-grained label category, then ω(y k ,Level)=2.0, otherwise it is 1.
Citation Information
Patent Citations
Text auditing method and device, electronic equipment and storage medium
CN113886573A
Sensitive information identification method and device, equipment and storage medium
CN117332090A
Transform structure-based intelligent evaluation report generation method and system
CN118261163A
Video highlight detection method based on weak supervision multi-mode large model
CN119785257A
Method and device for training semi-supervised object detection model, and object detection method and device
WO2024222444A1
Cited By
Surveying and mapping result classification model fusing time dynamic weight and use method
CN121030475A
A classification system for surveying and mapping results incorporating time-dynamic weights and its application methods
CN121030475B