Image-text content auditing method based on multi-modal large model
Through the multimodal large model-based graphic content review method, the existing system's insufficient semantic understanding and insufficient cross-modal data integration capabilities are solved, and high-precision content review and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202411980982.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The existing news content audit system lacks deep semantic understanding when processing unstructured data, it is difficult to accurately identify complex situations and information relevance in images and texts, and lacks cross-modal data integration capabilities, resulting in limited audit accuracy and excessive burden on audit personnel.
The multimodal large model-based graphic content review method is adopted to perform semantic analysis through the unstructured data processing module to generate semantic correlation diagrams, and extract the core situation and information correlation in text and images; the multimodal data integration module integrates graphic text, audio and video data in a unified context to generate graphic text correlation vector space; the content analysis module performs in-depth analysis, generates situation factors, sensitive factors and multimodal correlation factors, calculates the audit weight value and content anomaly index; the intelligent control module performs compliance and risk assessment, and selectively triggers manual review or automatic release.
It improves the identification accuracy of complex situations and information associations, improves audit accuracy, realizes context consistency and situational correlation of cross-modal data, reduces the burden on auditors, and significantly improves audit efficiency.
Smart Images

Figure CN119941157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of content auditing, and in particular to a method for auditing graphic and text content based on a multimodal large model. Background Art
[0002] Content review is an important part of the information dissemination process of news media. Standardized dissemination of text, images, audio and video content can avoid misleading the public, enhance the guidance and credibility of news media, and promote social stability. Existing manual content review has high accuracy but low efficiency. The supervised deep learning model for a specific scenario places a heavy burden on manual reviewers during the review reasoning process. Existing intelligent content review is based on a specific lack of active learning ability, and it is difficult to obtain features other than prior knowledge and cannot understand multimodal content. Based on the above problems, we first designed a content review method based on multimodal large-scale for specific content review scenarios. Through the ability of large models to process unstructured data, such as graphic data and context understanding, we can enhance the deep understanding of text images. Secondly, we constructed news graphic content review methods for 4 scenarios and conducted comparative experiments. The experimental results show that based on multimodal datasets, large models are used for training and testing, which improves the accuracy of intelligent content review and reduces the burden on content reviewers.
[0003] Based on the limitations of the existing news content review system and combined with the advantages of multimodal large models, the following technical shortcomings still exist in practical applications:
[0004] 1. Difficulty in understanding unstructured data: Traditional news content review systems lack deep semantic understanding when processing unstructured data such as images, text, and audio. They find it difficult to accurately identify complex contexts and information relevance in images and texts, which limits the accuracy of the review.
[0005] 2. Lack of multimodal data integration: Most existing systems process different types of data independently and lack cross-modal integration capabilities. For example, the relationship between images and text is difficult to understand in the same context, resulting in information fragmentation in the review of combined image and text content.
[0006] 3. Reliance on manual review, heavy review burden: Although traditional systems have certain intelligent review capabilities, they still rely on manual review for multiple rounds of screening and judgment, which increases the workload of reviewers and makes it difficult to effectively improve review efficiency. Summary of the invention
[0007] The purpose of the present invention is to provide a method for reviewing graphic content based on a multimodal large model to solve the above-mentioned problems.
[0008] The present invention is achieved through the following technical solutions:
[0009] A method for reviewing text and image content based on a multimodal large model is applied to a system for reviewing text and image content based on a multimodal large model, wherein the system comprises an unstructured data processing module, a multimodal data integration module, a content analysis module, a content anomaly quantification module, and an intelligent control module;
[0010] The unstructured data processing module is used to perform semantic analysis on unstructured data including news graphics, audio and video using a multimodal large model, generate a semantic association graph, extract the core context and information relevance in text and images through in-depth analysis, and construct a preliminary review data set;
[0011] The multimodal data integration module is used to integrate the graphic data, audio data and video data in a unified context, convert them into a graphic-text association vector space, and transmit them to the content analysis module for in-depth analysis of content context and emotion;
[0012] The content analysis module is used to preliminarily determine the degree of matching between the text and the image content based on the semantic association graph generated by the unstructured data processing module, generate a context factor Qycs, and analyze the potential sensitive content in each modal data to generate a sensitivity factor Gmyz; then, based on the context factor Qycs and the sensitivity factor Gmyz, perform hierarchical analysis on the potential emotional factors, social sensitive information and potential misleading factors in the image and text content to generate a multimodal association factor Mmgyz;
[0013] The content anomaly quantification module calculates the audit weight value Hqwz based on the situational factor Qycs, the sensitivity factor Gmyz and the multimodal association factor Mmgyz; then obtains the system audit risk factor Rxz by calculation and associates it with the audit weight value Hqwz, thereby generating the final content anomaly index Cxzs;
[0014] Among them, the calculation formula of the audit weight value Hqwz is:
[0015] Hqwz=b1×Qycs+b2×Gmyz+b3×Mmgyz+B;
[0016] Among them, b1, b2 and b3 represent the preset weight coefficients of the situational factor Qycs, the sensitive factor Gmyz and the multimodal association factor Mmgyz respectively, and B represents the audit correction coefficient;
[0017] The intelligent control module is used to set the compliance threshold T and the audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standards, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger the manual review or automatic release control plan based on the audit result.
[0018] The method comprises the following steps:
[0019] Step 1: Use a multimodal big model to perform semantic analysis on unstructured data including news graphics, audio, and video, generate a semantic association graph, extract the core context and information relevance in text and images through in-depth analysis, and build a preliminary review data set;
[0020] Step 2: After integrating the text, audio and video data in a unified context, converting them into a text-image association vector space, and transmitting them to the content analysis module for in-depth analysis of the content context and emotions;
[0021] Step 3: Based on the semantic association graph generated by the unstructured data processing module, the matching degree between the text and the image content is preliminarily determined to generate the context factor Qycs, and the potential sensitive content in each modal data is analyzed to generate the sensitivity factor Gmyz; then, based on the context factor Qycs and the sensitivity factor Gmyz, the potential emotional factors, social sensitive information and potential misleading factors in the image and text content are hierarchically analyzed to generate the multimodal association factor Mmgyz;
[0022] Step 4: Calculate the audit weight value Hqwz based on the situational factor Qycs, the sensitive factor Gmyz, and the multimodal correlation factor Mmgyz; then obtain the system audit risk factor Rxz by calculation and associate it with the audit weight value Hqwz, thereby generating the final content anomaly index Cxzs;
[0023] Step 5. Set the compliance threshold T and the audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standards, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger manual review or automatic release control plans based on the audit results.
[0024] Preferably, the unstructured data processing module is first used to receive the unstructured data of news graphics, audio and video to be reviewed, and perform format standardization processing on the unstructured data, including image resolution adjustment, text language standardization processing and audio and video transcoding;
[0025] Then, the pre-trained network of the multimodal large model is called to extract multi-level features of news pictures, audio and video, including visual features, semantic features and emotional features, and the extracted multi-level features are normalized to generate a unimodal feature matrix; based on the unimodal feature matrix, the multi-level features of news pictures, audio and video are mapped in the multimodal association vector space, and a cross-modal semantic association graph is generated through the semantic embedding mechanism of the multimodal large model;
[0026] Secondly, we use the cross-modal semantic association graph to conduct a detailed analysis of the deep-level association between text and images, extract the core contextual elements in the text and images, and simultaneously analyze the emotional state and expression trends in the audio and video, so as to identify potential sensitive information, misleading information and social influencing factors in the news content in real time, and form multi-modal association semantic indicators;
[0027] Finally, a preliminary review dataset is generated.
[0028] Preferably, the multimodal data integration module is used to receive the preliminary review data set generated by the unstructured data processing module, and perform uniform alignment processing of the timing and content of the graphic data, audio data and video data therein through a frame synchronization algorithm;
[0029] Then, a multimodal feature mapping algorithm is applied to fuse the context, emotion and semantic features of text, audio and video into a multimodal feature vector, and a correlation matrix is generated based on the correlation between each modality to perform hierarchical mapping of information density, context relevance and semantic consistency.
[0030] Subsequently, the image-text association vector space is constructed based on the association matrix, and the multimodal data is quantified in a unified semantic space to generate an integrated semantic association representation vector. At the same time, the semantic deviation caused by modality differences is corrected through the context adjustment algorithm.
[0031] Finally, the integrated image-text association vector space is transmitted to the content analysis module.
[0032] Preferably, the content analysis module includes a preliminary discrimination unit, a multi-layer sensitivity analysis unit and a correlation factor generation unit;
[0033] The preliminary discrimination unit identifies semantic deviation related data and content matching related data by analyzing the semantic similarity and context consistency between the image and the text; extracts the semantic similarity Ycy, context consistency Sjy, context association deviation Qjl and image and text emotion consistency Twq from the semantic deviation related data and the content matching related data and performs dimensionless processing, and then calculates the context factor Qycs through the following formula:
[0034]
[0035] The preset situation threshold Q is compared with the situation factor Qycs to evaluate the situational relevance of text and images in the audit scenario. The specific comparison and evaluation contents are as follows:
[0036] If the context factor Qycs ≥ the context threshold Q, it indicates that the contexts of the text and the image match, and the semantics, emotions, and contexts of the two are consistent. The review is passed and marked as "normal context association";
[0037] If the context factor Qycs < context threshold Q, it indicates that the context between the text and the image does not match, and the semantics, emotions and context of the two are inconsistent. The review will fail and further analysis of the specific reasons for the deviation will be conducted, including potential sensitive content.
[0038] Preferably, the multi-layer sensitivity analysis unit is used to further perform in-depth analysis on the potential sensitive content of the graphic data, including analyzing the emotion-related data, social sensitivity-related data and potential misleading-related data in the text and image, and constructing a sensitive content data set after summarizing and dimensionless processing the emotion-related data, social sensitivity-related data and potential misleading-related data in the analyzed text and image; extracting the sensitive content data set and generating the sensitivity factor Gmyz through the following formula:
[0039]
[0040] Where Qmq represents the emotion intensity coefficient in the sensitive content data set, Mgc represents the sensitive word index in the sensitive content data set, and Xwg represents the possibility of information misleading in the sensitive content data set.
[0041] Preferably, the correlation factor generating unit is used to quantify the correlation degree between emotion, context and sensitive content based on the context factor Qycs and the sensitivity factor Gmyz by fusing multi-level information in the image-text correlation vector space, and generate the multimodal correlation factor Mmgyz by calculating through the following formula;
[0042]
[0043] Preferably, the content anomaly quantification module includes an audit calculation unit and an anomaly index acquisition unit.
[0044] The content evaluation unit is used to construct a system audit risk factor Rxz, obtain the semantic deviation degree Ycp and the content consistency index Nry by extracting semantic deviation related data and content matching related data, and obtain the high sensitivity trigger rate Gmg by extracting the sensitive content data set, and calculate the system audit risk factor Rxz in combination with the following formula:
[0045]
[0046] Preferably, the abnormality index obtaining unit is used to calculate and obtain the final content abnormality index Cxzs, and its specific calculation formula is as follows:
[0047]
[0048] Preferably, the intelligent control module compares and evaluates the audit weight value Hqwz with the compliance threshold T, and compares and evaluates the audit risk threshold R with the content anomaly index Cxzs, and specifically generates the following evaluation content:
[0049] Compliance standards comparison:
[0050] If the audit weight value Hqwz ≥ the compliance threshold T, it means that the content meets the audit standards and the system marks it as "compliant"
[0051] If the audit weight value Hqwz < the compliance threshold T, it means that the content does not meet the audit standards and is insufficiently compliant, requiring further review or adjustment.
[0052] Abnormal risk assessment:
[0053] If the content anomaly index Cxzs ≥ the audit risk threshold R, it indicates that the content has an abnormal risk, and the system generates an "abnormal content" mark, triggering the manual review process;
[0054] If the content anomaly index Cxzs < the audit risk threshold R, it means that there is no abnormal risk in the content, and the system generates a "normal content" mark, and then enters the automatic publishing process;
[0055] When the evaluation results of the audit weight value Hqwz and the content anomaly index Cxzs are marked as "compliant" and "normal content" at the same time, the content will be automatically marked as "passed" and published directly;
[0056] When the evaluation result of the audit weight value Hqwz and the content anomaly index Cxzs is that the content does not meet the audit standards or has abnormal risks, the content will be automatically marked as "pending review" and the selective review process will be triggered.
[0057] It is used to set the compliance threshold T and the audit risk threshold R, and to determine whether the current content meets the audit standards, obtain the final audit results, and selectively trigger manual review or automatic release control plans based on the audit results.
[0058] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0059] 1. The present invention provides a more comprehensive and in-depth semantic understanding capability of unstructured data through the collaborative work of the unstructured data processing module, the multimodal data integration module, the content analysis module, the content anomaly quantification module and the intelligent control module, solves the problem of inaccurate semantic understanding of the existing system when processing unstructured data including pictures, texts, audio and video, generates a semantic association graph using a multimodal large model, and extracts the core context and information relevance in texts and images through in-depth semantic analysis, thereby improving the system's recognition accuracy of complex contexts and information associations, and significantly improving the accuracy of auditing;
[0060] 2. The present invention realizes the unified integration of cross-modal data through a multimodal data integration module, performs unified alignment processing of time sequence and content of graphic data, audio data and video data, generates multimodal feature vectors with a multimodal feature mapping algorithm, and hierarchically maps information density, context relevance and semantic consistency in a unified semantic space through an association matrix, thereby realizing the construction of a graphic-text association vector space, ensuring the context consistency and context relevance of different types of data, and enabling the system to have information integration capabilities in the review of cross-modal content, solving the problem of fragmentation in traditional systems when reviewing graphic-text combined content;
[0061] 3. The present invention realizes an intelligent audit process through the quantitative calculation and comparison of the audit weight value Hqwz, the system audit risk factor Rxz and the content anomaly index Cxzs through the content anomaly quantification module and the intelligent control module, sets the compliance threshold T and the audit risk threshold R to make multi-level audit decisions, and when the content anomaly index Cxzs is higher than the audit risk threshold R, it is marked as "abnormal content" to trigger the manual review process; when the audit weight value Hqwz is higher than the compliance threshold T, it is marked as "compliant" and enters the automatic release process. This mechanism avoids tedious manual review, reduces the burden on auditors, and significantly improves the audit efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:
[0063] Figure 1 A schematic diagram of the framework structure of a graphic content review system based on a multimodal large model of the present invention;
[0064] Figure 2 A schematic diagram of the steps of a method for reviewing text and image content based on a multimodal large model according to the present invention; DETAILED DESCRIPTION
[0065] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and drawings. The schematic implementation modes and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention. It should be noted that the present invention is already in the actual development and use stage.
[0066] Example 1
[0067] like Figure 1 As shown, this embodiment includes a graphic content review system based on a multimodal large model, including an unstructured data processing module, a multimodal data integration module, a content analysis module, a content anomaly quantification module and an intelligent control module;
[0068] The unstructured data processing module is used to perform semantic analysis on unstructured data including news graphics, audio and video using a multimodal large model, generate a semantic association graph, extract the core context and information relevance in text and images through in-depth analysis, and construct a preliminary review data set;
[0069] The multimodal data integration module is used to integrate the graphic data, audio data and video data in a unified context, convert them into a graphic-text association vector space, and transmit them to the content analysis module for in-depth analysis of content context and emotion;
[0070] The content analysis module is used to preliminarily determine the degree of matching between the text and the image content based on the semantic association graph generated by the unstructured data processing module, generate a context factor Qycs, and analyze the potential sensitive content in each modal data to generate a sensitivity factor Gmyz; then, based on the context factor Qycs and the sensitivity factor Gmyz, perform hierarchical analysis on the potential emotional factors, social sensitive information and potential misleading factors in the image and text content to generate a multimodal association factor Mmgyz;
[0071] The content anomaly quantification module calculates the audit weight value Hqwz based on the situational factor Qycs, the sensitivity factor Gmyz and the multimodal association factor Mmgyz; then obtains the system audit risk factor Rxz by calculation and associates it with the audit weight value Hqwz, thereby generating the final content anomaly index Cxzs;
[0072] Among them, the calculation formula of the audit weight value Hqwz is:
[0073] Hqwz=b1×Qycs+b2×Gmyz+b3×Mmgyz+B;
[0074] Among them, b1, b2 and b3 represent the preset weight coefficients of the situational factor Qycs, the sensitive factor Gmyz and the multimodal association factor Mmgyz respectively, and B represents the audit correction coefficient;
[0075] The intelligent control module is used to set the compliance threshold T and the audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standards, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger manual review or automatic release control plans based on the audit results.
[0076] In this embodiment, the unstructured data processing module uses a multimodal large model to perform semantic analysis on unstructured data including news graphics, audio and video to generate a semantic association graph, which can deeply extract the core context and information correlation in text and images, thereby improving the accuracy of understanding complex contexts and helping to build a more complete preliminary review data set; the multimodal data integration module integrates graphic data, audio data and video data in a unified context, converts them into a graphic-text association vector space and transmits them to the content analysis module, so that the system has the context consistency and information integration capabilities of cross-modal data, and solves the problem of information fragmentation in the existing review system in the combination of graphics and text; the content analysis module can preliminarily determine the degree of matching between text and image content based on the semantic association graph generated by the unstructured data processing module, and generate context factor Qycs, sensitive factor Gmyz and multimodal relationship The multimodal correlation factor Mmgyz is used to realize hierarchical analysis of potential emotional factors, social sensitive information and potential misleading factors, thereby improving the system's ability to conduct detailed review of contexts and sensitive content; the content anomaly quantification module calculates the review weight value Hqwz based on the context factor Qycs, the sensitive factor Gmyz and the multimodal correlation factor Mmgyz, and generates the final content anomaly index Cxzs by calculating the system review risk factor Rxz and associating it with the review weight value Hqwz, which can effectively quantify the degree of content anomaly and provide a quantitative reference for the review; the intelligent control module sets the compliance threshold T and the review risk threshold R, and compares and evaluates the review weight value Hqwz with the compliance threshold T and the content anomaly index Cxzs with the review risk threshold R, selectively triggers manual review or automatically releases the plan, thereby realizing automation and intelligent control of the review process, effectively reducing the burden of manual review and improving review efficiency.
[0077] Example 2
[0078] The unstructured data processing module is first used to receive the unstructured data of news graphics, audio and video to be reviewed, and perform format standardization processing on the unstructured data, including image resolution adjustment, text language standardization processing and audio and video transcoding;
[0079] Then, the pre-trained network of the multimodal large model is called to extract multi-level features of news pictures, audio and video, including visual features, semantic features and emotional features, and the extracted multi-level features are normalized to generate a unimodal feature matrix; based on the unimodal feature matrix, the multi-level features of news pictures, audio and video are mapped in the multimodal association vector space, and a cross-modal semantic association graph is generated through the semantic embedding mechanism of the multimodal large model;
[0080] Secondly, we use the cross-modal semantic association graph to conduct a detailed analysis of the deep-level association between text and images, extract the core contextual elements in the text and images, and simultaneously analyze the emotional state and expression trends in the audio and video, so as to identify potential sensitive information, misleading information and social influencing factors in the news content in real time, and form multi-modal association semantic indicators;
[0081] Finally, a preliminary review dataset is generated.
[0082] The multimodal data integration module is used to receive the preliminary review data set generated by the unstructured data processing module, and perform uniform alignment processing of the timing and content of the graphic data, audio data and video data therein through a frame synchronization algorithm;
[0083] Then, a multimodal feature mapping algorithm is applied to fuse the context, emotion and semantic features of text, audio and video into a multimodal feature vector, and a correlation matrix is generated based on the correlation between each modality to perform hierarchical mapping of information density, context relevance and semantic consistency.
[0084] Subsequently, the image-text association vector space is constructed based on the association matrix, and the multimodal data is quantified in a unified semantic space to generate an integrated semantic association representation vector. At the same time, the semantic deviation caused by modality differences is corrected through the context adjustment algorithm.
[0085] Finally, the integrated image-text association vector space is transmitted to the content analysis module.
[0086] In this embodiment, the multimodal large-model graphic content review system realizes the standardization and feature extraction of unstructured data such as news graphics, audio and video through the hierarchical processing of the unstructured data processing module and the multimodal data integration module, so as to be able to deeply analyze the visual features, semantic features and emotional features in the multimodal data, and ensure that the reviewed content has semantic and emotional integrity; wherein, the unstructured data processing module extracts the multi-level features in the news graphics, audio and video through the pre-trained network, and generates a unimodal feature matrix, and then generates a cross-modal semantic association graph, thereby realizing the accurate analysis of the deep correlation between text, images, audio and video, and extracting the core situational elements and emotional states in the content. It can effectively identify potential sensitive information, misleading information and social influencing factors, so as to construct a complete preliminary review data set; the multimodal data integration module aligns the graphic data, audio data and video data in terms of time sequence and content through a frame synchronization algorithm, and applies a multimodal feature mapping algorithm to generate a multimodal feature vector, so that the system has contextual consistency and correlation matrix of cross-modal information, quantifies and hierarchically maps information density, contextual relevance and semantic consistency in a unified semantic space, and eliminates semantic deviations between modalities through a context adjustment algorithm, generating an integrated semantic association representation vector, thereby providing a complete and quantitative data foundation for the content analysis module and improving the accuracy and comprehensiveness of the system's review in complex situations.
[0087] Example 3
[0088] The content analysis module includes a preliminary discrimination unit, a multi-layer sensitivity analysis unit and a correlation factor generation unit;
[0089] The preliminary discrimination unit identifies semantic deviation related data and content matching related data by analyzing the semantic similarity and context consistency between the image and the text; extracts the semantic similarity Ycy, context consistency Sjy, context association deviation Qjl and image and text emotion consistency Twq from the semantic deviation related data and the content matching related data and performs dimensionless processing, and then calculates the context factor Qycs through the following formula:
[0090]
[0091] The preset situation threshold Q is compared with the situation factor Qycs to evaluate the situational relevance of text and images in the audit scenario. The specific comparison and evaluation contents are as follows:
[0092] If the context factor Qycs ≥ the context threshold Q, it indicates that the contexts of the text and the image match, and the semantics, emotions, and contexts of the two are consistent. The review is passed and marked as "normal context association";
[0093] If the context factor Qycs < context threshold Q, it indicates that the context between the text and the image does not match, and the semantics, emotions and context of the two are inconsistent. The review will fail and further analysis of the specific reasons for the deviation will be conducted, including potential sensitive content.
[0094] The multi-layer sensitivity analysis unit is used to further perform in-depth analysis on the potential sensitive content of the graphic data, including analyzing the emotion-related data, social sensitivity-related data and potential misleading-related data in the text and image, and constructing a sensitive content data set after summarizing and dimensionless processing the emotion-related data, social sensitivity-related data and potential misleading-related data in the analyzed text and image; extracting the sensitive content data set and then generating the sensitivity factor Gmyz through the following formula:
[0095]
[0096] Where Qmq represents the emotion intensity coefficient in the sensitive content data set, Mgc represents the sensitive word index in the sensitive content data set, and Xwg represents the possibility of information misleading in the sensitive content data set.
[0097] The correlation factor generation unit is used to quantify the correlation degree between emotion, context and sensitive content based on the context factor Qycs and the sensitivity factor Gmyz by fusing multi-level information in the image-text correlation vector space, and calculate and generate the multimodal correlation factor Mmgyz by the following formula;
[0098]
[0099] In this embodiment, the content analysis module enables the system to comprehensively evaluate the contextual relevance, potential sensitivity and multimodal relevance of the graphic content through multi-level analysis of the preliminary judgment unit, the multi-layer sensitivity analysis unit and the correlation factor generation unit, thereby greatly improving the accuracy and comprehensiveness of the audit; wherein, the preliminary judgment unit extracts semantic deviation-related data and content matching-related data such as semantic similarity Ycy, context consistency Sjy, contextual relevance deviation Qjl and graphic-text emotional consistency Twq by analyzing the semantic similarity and contextual consistency of the graphic and text, and performs dimensionless processing, and generates the contextual factor Qycs by calculation, effectively quantifies the contextual consistency of the graphic content, and determines whether the context matches by comparing with the preset context threshold Q, thereby realizing basic context evaluation; wherein, the semantic similarity Ycy is used to measure the consistency between the text content and the image represented. The similarity of the content at the semantic level is expressed by inputting the text into a pre-trained multimodal model to obtain the text vector representation; the context consistency Sjy mainly measures the consistency of multi-dimensional information such as time, place, person or event between the text context and the image context, and is obtained by comparing the context feature set obtained from the text with the context feature set identified in the image or video; the context association deviation Qjl is used to quantify the potential context mismatch between the text and the image or video. After aligning the text and the image, the inconsistent parts of the "core scene or main object" are identified, and the inconsistent categories or numbers are counted to obtain the result; the text-image emotional consistency Twq is a measure of the degree of consistency between the emotions conveyed by the text and the emotions conveyed by the image, and is obtained by calculating the similarity between the text emotional distribution and the image emotional distribution;
[0100] The multi-layer sensitivity analysis unit deeply analyzes the emotions, social sensitivity and potential misleading data in texts and images, extracts parameters such as the emotion intensity coefficient Qmq, the sensitive word index Mgc and the possibility of information misleading Xwg to construct a sensitive content data set, and further calculates and generates the sensitivity factor Gmyz, quantifies the sensitivity of the graphic content, and ensures the audit's ability to identify potential risks; among them, the emotion intensity coefficient Qmq is used to reflect the emotion or emotion intensity in the news content, which is obtained through text emotion intensity analysis; the sensitive word index Mgc is mainly used to measure the frequency and intensity of sensitive information, sensitive terms, banned words or sensitive symbols in the text, which is obtained through sensitive dictionary and rule matching; the possibility of information misleading Xwg is used to evaluate whether there is a potential tendency in the news content to "confuse the audience", "falsify information" or "mislead the public", which is obtained through cross-modal comparison;
[0101] The correlation factor generation unit is based on the situational factor Qycs and the sensitive factor Gmyz, and integrates the multi-level information in the image-text correlation vector space to generate a multimodal correlation factor Mmgyz. Through the correlation evaluation of emotions, situations and sensitive content, it realizes a comprehensive review of the image and text content, enabling the system to have the ability of cross-modal situational association and risk identification, effectively improving the accuracy and depth of content review.
[0102] Example 4
[0103] The content anomaly quantification module includes an audit calculation unit and an anomaly index acquisition unit.
[0104] The content evaluation unit is used to construct a system audit risk factor Rxz, obtain the semantic deviation degree Ycp and the content consistency index Nry by extracting semantic deviation related data and content matching related data, and obtain the high sensitivity trigger rate Gmg by extracting the sensitive content data set, and calculate the system audit risk factor Rxz in combination with the following formula:
[0105]
[0106] The abnormal index acquisition unit is used to calculate and obtain the final content abnormal index Cxzs, and its specific calculation formula is as follows:
[0107]
[0108] The intelligent control module compares and evaluates the audit weight value Hqwz with the compliance threshold T, and compares and evaluates the audit risk threshold R with the content anomaly index Cxzs, and specifically generates the following evaluation content:
[0109] Compliance standards comparison:
[0110] If the audit weight value Hqwz ≥ the compliance threshold T, it means that the content meets the audit standards and the system marks it as "compliant"
[0111] If the audit weight value Hqwz < the compliance threshold T, it means that the content does not meet the audit standards and is insufficiently compliant, requiring further review or adjustment.
[0112] Abnormal risk assessment:
[0113] If the content anomaly index Cxzs ≥ the audit risk threshold R, it indicates that the content has an abnormal risk, and the system generates an "abnormal content" mark, triggering the manual review process;
[0114] If the content anomaly index Cxzs < the audit risk threshold R, it means that there is no abnormal risk in the content, and the system generates a "normal content" mark, and then enters the automatic publishing process;
[0115] When the evaluation results of the audit weight value Hqwz and the content anomaly index Cxzs are marked as "compliant" and "normal content" at the same time, the content will be automatically marked as "passed" and published directly;
[0116] When the evaluation result of the audit weight value Hqwz and the content anomaly index Cxzs is that the content does not meet the audit standards or has abnormal risks, the content will be automatically marked as "pending review" and the selective review process will be triggered.
[0117] It is used to set the compliance threshold T and the audit risk threshold R, and to determine whether the current content meets the audit standards, obtain the final audit results, and selectively trigger manual review or automatic release control plans based on the audit results.
[0118] In this embodiment, the content anomaly quantification module provides the system with comprehensive content anomaly quantitative analysis capabilities through step-by-step calculations by the audit calculation unit and the anomaly index acquisition unit, thereby improving the accuracy and automation level of the audit; the audit calculation unit quantifies the degree of deviation in semantics and consistency of the graphic content by extracting the semantic deviation degree Ycp in the semantic deviation related data and the content consistency index Nry in the content matching related data, and calculates and generates the system audit risk factor Rxz in combination with the high sensitivity trigger rate Gmg in the sensitive content data set, thereby comprehensively quantifying the potential abnormal risk of the content and providing accurate risk data support for the subsequent audit process; the anomaly index acquisition unit calculates the content anomaly index Cxzs based on the system audit risk factor Rxz and the audit weight value Hqwz, It effectively quantifies the comprehensive performance of the content in terms of compliance and risk, enabling the system to accurately evaluate the suitability of content release; the intelligent control module sets the compliance threshold T and the audit risk threshold R, and compares and evaluates the audit weight value Hqwz with the compliance threshold T, and the content anomaly index Cxzs with the audit risk threshold R, thereby realizing multi-level automated review and judgment of the content. When the evaluation results of the audit weight value Hqwz and the content anomaly index Cxzs are both "compliant" and "normal content", the content is directly passed and automatically released; when the evaluation results show that the content does not meet the audit standards or there is an abnormal risk, the content is marked as "pending review" and the manual review process is triggered, realizing efficient control and intelligent decision-making of content release review, and effectively improving the accuracy and efficiency of system review.
[0119] Example 5
[0120] like Figure 2 As shown, a method for reviewing text and image content based on a multimodal large model, according to a system for reviewing text and image content based on a multimodal large model, comprises the following steps:
[0121] Step 1: Use a multimodal big model to perform semantic analysis on unstructured data including news graphics, audio, and video, generate a semantic association graph, extract the core context and information relevance in text and images through in-depth analysis, and build a preliminary review data set;
[0122] Step 2: After integrating the text, audio and video data in a unified context, converting them into a text-image association vector space, and transmitting them to the content analysis module for in-depth analysis of the content context and emotions;
[0123] Step 3: Based on the semantic association graph generated by the unstructured data processing module, the matching degree between the text and the image content is preliminarily determined to generate the context factor Qycs, and the potential sensitive content in each modal data is analyzed to generate the sensitivity factor Gmyz; then, based on the context factor Qycs and the sensitivity factor Gmyz, the potential emotional factors, social sensitive information and potential misleading factors in the image and text content are hierarchically analyzed to generate the multimodal association factor Mmgyz;
[0124] Step 4: Calculate the audit weight value Hqwz based on the situational factor Qycs, the sensitive factor Gmyz, and the multimodal correlation factor Mmgyz; then obtain the system audit risk factor Rxz by calculation and associate it with the audit weight value Hqwz, thereby generating the final content anomaly index Cxzs;
[0125] Step 5. Set the compliance threshold T and the audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standards, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger manual review or automatic release control plans based on the audit results.
[0126] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for reviewing text and image content based on a multimodal large model, characterized by: Applied to a graphic content review system based on a multimodal large model, the system includes an unstructured data processing module, a multimodal data integration module, a content analysis module, a content anomaly quantification module and an intelligent control module; The unstructured data processing module is used to perform semantic analysis on unstructured data including news graphics, audio and video using a multimodal large model, generate a semantic association graph, extract the core context and information relevance in text and images through in-depth analysis, and construct a preliminary review data set; The multimodal data integration module is used to integrate the graphic data, audio data and video data in a unified context, convert them into a graphic-text association vector space, and transmit them to the content analysis module for in-depth analysis of content context and emotion; The content analysis module is used to preliminarily determine the degree of matching between the text and the image content based on the semantic association graph generated by the unstructured data processing module, generate a context factor Qycs, and analyze the potential sensitive content in each modal data to generate a sensitivity factor Gmyz; then, based on the context factor Qycs and the sensitivity factor Gmyz, perform hierarchical analysis on the potential emotional factors, social sensitive information and potential misleading factors in the image and text content to generate a multimodal association factor Mmgyz; The content anomaly quantification module calculates the audit weight value Hqwz based on the situational factor Qycs, the sensitivity factor Gmyz and the multimodal association factor Mmgyz; then obtains the system audit risk factor Rxz by calculation and associates it with the audit weight value Hqwz, thereby generating the final content anomaly index Cxzs; Among them, the calculation formula of the audit weight value Hqwz is: Hqwz=b1×Qycs+b2×Gmyz+b3×Mmgyz+B; Among them, b1, b2 and b3 represent the preset weight coefficients of the situational factor Qycs, the sensitive factor Gmyz and the multimodal association factor Mmgyz respectively, and B represents the audit correction coefficient; The intelligent control module is used to set the compliance threshold T and the audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standards, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger the manual review or automatic release control plan based on the audit result. The method comprises the following steps: Step 1: Use a multimodal big model to perform semantic analysis on unstructured data including news graphics, audio, and video, generate a semantic association graph, extract the core context and information relevance in text and images through in-depth analysis, and build a preliminary review data set; Step 2: After integrating the text, audio and video data in a unified context, converting them into a text-image association vector space, and transmitting them to the content analysis module for in-depth analysis of the content context and emotions; Step 3: Based on the semantic association graph generated by the unstructured data processing module, the matching degree between the text and the image content is preliminarily determined to generate the context factor Qycs, and the potential sensitive content in each modal data is analyzed to generate the sensitivity factor Gmyz; then, based on the context factor Qycs and the sensitivity factor Gmyz, the potential emotional factors, social sensitive information and potential misleading factors in the image and text content are hierarchically analyzed to generate the multimodal association factor Mmgyz; Step 4: Calculate the audit weight value Hqwz based on the situational factor Qycs, the sensitive factor Gmyz, and the multimodal correlation factor Mmgyz; then obtain the system audit risk factor Rxz by calculation and associate it with the audit weight value Hqwz, thereby generating the final content anomaly index Cxzs; Step 5. Set the compliance threshold T and the audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standards, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger manual review or automatic release control plans based on the audit results.
2. According to claim 1, a method for reviewing text and image content based on a multimodal large model is characterized by: The unstructured data processing module is first used to receive the unstructured data of news graphics, audio and video to be reviewed, and perform format standardization processing on the unstructured data, including image resolution adjustment, text language standardization processing and audio and video transcoding; Then, the pre-trained network of the multimodal large model is called to extract multi-level features of news pictures, audio and video, including visual features, semantic features and emotional features, and the extracted multi-level features are normalized to generate a unimodal feature matrix; based on the unimodal feature matrix, the multi-level features of news pictures, audio and video are mapped in the multimodal association vector space, and a cross-modal semantic association graph is generated through the semantic embedding mechanism of the multimodal large model; Secondly, we use the cross-modal semantic association graph to conduct a detailed analysis of the deep-level association between text and images, extract the core contextual elements in the text and images, and simultaneously analyze the emotional state and expression trends in the audio and video, so as to identify potential sensitive information, misleading information and social influencing factors in the news content in real time, and form multi-modal association semantic indicators; Finally, a preliminary review dataset is generated.
3. The method for reviewing text and image content based on a multimodal large model according to claim 2 is characterized in that: The multimodal data integration module is used to receive the preliminary review data set generated by the unstructured data processing module, and perform uniform alignment processing of the timing and content of the graphic data, audio data and video data therein through a frame synchronization algorithm; Then, a multimodal feature mapping algorithm is applied to fuse the context, emotion and semantic features of text, audio and video into a multimodal feature vector, and a correlation matrix is generated based on the correlation between each modality to perform hierarchical mapping of information density, context relevance and semantic consistency. Subsequently, the image-text association vector space is constructed based on the association matrix, and the multimodal data is quantified in a unified semantic space to generate an integrated semantic association representation vector. At the same time, the semantic deviation caused by modality differences is corrected through the context adjustment algorithm. Finally, the integrated image-text association vector space is transmitted to the content analysis module.
4. The method for reviewing text and image content based on a multimodal large model according to claim 3 is characterized by: The content analysis module includes a preliminary discrimination unit, a multi-layer sensitivity analysis unit and a correlation factor generation unit; The preliminary discrimination unit identifies semantic deviation related data and content matching related data by analyzing the semantic similarity and context consistency between the image and the text; extracts the semantic similarity Ycy, context consistency Sjy, context association deviation Qjl and image and text emotion consistency Twq from the semantic deviation related data and the content matching related data and performs dimensionless processing, and then calculates the context factor Qycs through the following formula: The preset situation threshold Q is compared with the situation factor Qycs to evaluate the situational relevance of text and images in the audit scenario. The specific comparison and evaluation contents are as follows: If the context factor Qycs ≥ the context threshold Q, it indicates that the contexts of the text and the image match, and the semantics, emotions, and contexts of the two are consistent. The review is passed and marked as "normal context association"; If the context factor Qycs < context threshold Q, it indicates that the context between the text and the image does not match, and the semantics, emotions and context of the two are inconsistent. The review will fail and further analysis of the specific reasons for the deviation will be conducted, including potential sensitive content.
5. The method for reviewing text and image content based on a multimodal large model according to claim 4 is characterized in that: The multi-layer sensitivity analysis unit is used to further perform in-depth analysis on the potential sensitive content of the graphic data, including analyzing the emotion-related data, social sensitivity-related data and potential misleading-related data in the text and image, and constructing a sensitive content data set after summarizing and dimensionless processing the emotion-related data, social sensitivity-related data and potential misleading-related data in the analyzed text and image; extracting the sensitive content data set and then generating the sensitivity factor Gmyz through the following formula: Where Qmq represents the emotion intensity coefficient in the sensitive content data set, Mgc represents the sensitive word index in the sensitive content data set, and Xwg represents the possibility of information misleading in the sensitive content data set.
6. The method for reviewing text and image content based on a multimodal large model according to claim 5 is characterized by: The correlation factor generation unit is used to quantify the correlation degree between emotion, context and sensitive content based on the context factor Qycs and the sensitivity factor Gmyz by fusing multi-level information in the image-text correlation vector space, and calculate and generate the multimodal correlation factor Mmgyz by the following formula; 7. The method for reviewing text and image content based on a multimodal large model according to claim 6 is characterized by: The content anomaly quantification module includes an audit calculation unit and an anomaly index acquisition unit; The content evaluation unit is used to construct a system audit risk factor Rxz, obtain the semantic deviation degree Ycp and the content consistency index Nry by extracting semantic deviation related data and content matching related data, and obtain the high sensitivity trigger rate Gmg by extracting the sensitive content data set, and calculate the system audit risk factor Rxz in combination with the following formula:
8. The method for reviewing text and image content based on a multimodal large model according to claim 7 is characterized in that: The abnormal index acquisition unit is used to calculate and obtain the final content abnormal index Cxzs, and its specific calculation formula is as follows:
9. The method for reviewing text and image content based on a multimodal large model according to claim 8, characterized in that: The intelligent control module compares and evaluates the audit weight value Hqwz with the compliance threshold T, and compares and evaluates the audit risk threshold R with the content anomaly index Cxzs, and specifically generates the following evaluation content: Compliance standards comparison: If the audit weight value Hqwz ≥ the compliance threshold T, it means that the content meets the audit standards and the system marks it as "compliant" If the audit weight value Hqwz < the compliance threshold T, it means that the content does not meet the audit standards and is not compliant enough, which requires further review or adjustment. Abnormal risk assessment: If the content anomaly index Cxzs ≥ the audit risk threshold R, it indicates that the content has an abnormal risk, and the system generates an "abnormal content" mark, triggering the manual review process; If the content anomaly index Cxzs < the audit risk threshold R, it means that there is no abnormal risk in the content, and the system generates a "normal content" mark, and then enters the automatic publishing process; When the evaluation results of the audit weight value Hqwz and the content anomaly index Cxzs are marked as "compliant" and "normal content" at the same time, the content will be automatically marked as "passed" and published directly; When the evaluation result of the review weight value Hqwz and the content anomaly index Cxzs indicates that the content does not meet the review standards or has an abnormal risk, the content will be automatically marked as "pending review" and the selective review process will be triggered; It is used to set the compliance threshold T and the audit risk threshold R, and to determine whether the current content meets the audit standards, obtain the final audit results, and selectively trigger manual review or automatic release control plans based on the audit results.
Citation Information
Patent Citations
Audio auditing method, device and equipment and readable storage medium
CN114666618A
Multi-mode network content security intelligent auditing system and method thereof
CN118312922A
Clinical examination result auditing method and system based on artificial intelligence and big data
CN118629571A
Media asset intelligent auditing system based on large model
CN118965190A
Video auditing method based on multi-modal large model
CN118968380A
Cited By
Large model defense method based on multi-view image-text conversion
CN120145402A
A large model defense method based on multi-view image-text conversion
CN120145402B
Marketing video auditing method based on AI
CN120583273A
Consistency comparison method and system for multi-mode electronic signed files
CN120599630A
Multimodal data sensitivity grading method and system based on semantic risk map diffusion perception
CN120670963A