Intelligent auditing system and method for multi-modal content
By developing an intelligent review system and method for multimodal content, we extract modality generation feature vectors, construct a joint representation of generation intent, identify semantic offset relationships, and build a risk evolution model. This solves the problem of identifying risks caused by hidden semantic offsets and their evolution over time in multimodal content, and enables dynamic perception and accurate judgment of risks in multimodal content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING KAIYUAN GONGCHUANG TECH CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to identify the hidden semantic shifts between different modal generation intentions in multimodal content and the potential risks arising from their evolution over time, making it difficult to accurately identify and manage potential risks in multimodal content.
It provides an intelligent review system and method for multimodal content. The system extracts modal features to generate feature vectors, constructs and generates joint representations of intents through an intent inversion module, identifies semantic offset relationships through a cross-modal consistency violation identification module, constructs a risk evolution model through a risk evolution module, and outputs the review results.
Dynamically perceive and determine the risk status of multimodal content, improve the accuracy of review, identify cross-modal consistency violations, build a risk evolution model, and achieve dynamic perception and accurate determination of multimodal content risks.
Smart Images

Figure CN122045844A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically to an intelligent review system and method for multimodal content. Background Technology
[0002] As the forms of multimodal content generation and dissemination become increasingly diverse, different modalities such as text, images, and audio often collaborate in content expression. While their surface semantics may remain consistent, subtle and subtle differences exist in terms of generation intent, emotional guidance, or behavioral inducement. These differences do not manifest as explicit violations in a single modality, but rather as inconsistencies in semantic direction between different modalities and the gradual accumulation and amplification of risks over time. Due to the concealed and dynamically evolving nature of multimodal content generation intent, single, static review analyses struggle to capture the semantic shifts between different modalities in a timely manner, making it difficult to effectively characterize the changing trends of risks over continuous time. Consequently, the potential risks within multimodal content are difficult to accurately identify and manage. Summary of the Invention
[0003] This application provides an intelligent review system and method for multimodal content, which addresses the technical problem that existing technologies struggle to identify the hidden semantic shifts between different modal generation intentions in multimodal content and the potential risks arising from their evolution over time.
[0004] In view of the above problems, this application provides an intelligent review system and method for multimodal content.
[0005] The first aspect of this application provides an intelligent content moderation system for multimodal content, the system comprising: The feature extraction module is used to perform modal-level parsing processing on multimodal content containing text, images, or audio after receiving the multimodal content, and extract modal generation feature vectors that represent the generation tendency of each modality. The intent inversion module is used to perform generation intent inversion based on the modal generation feature vectors and construct a joint representation of multimodal generation intent. The cross-modal consistency violation identification module is used to perform cross-modal consistency analysis on the joint representation of multimodal generation intent, identify the semantic offset relationship between different modal generation intents, and construct a cross-modal consistency violation feature tensor. The risk evolution module is used to construct a multimodal content risk evolution model based on the cross-modal consistency violation feature tensor, and generate a risk evolution state quantity representing the risk evolution state of multimodal content according to the consistency violation change rate of cross-modal content in the time dimension. The review decision output module is used to perform multimodal risk limit state judgment condition matching based on the risk evolution state quantity and output the review result.
[0006] A second aspect of this application provides an intelligent moderation method for multimodal content, the method comprising: After receiving multimodal content containing text, images, or audio, modal-level parsing is performed on the multimodal content to extract modal generation feature vectors that characterize the generation tendency of each modality. Based on these modal generation feature vectors, generation intent inversion is performed to construct a joint representation of multimodal generation intent. Cross-modal consistency analysis is performed on this joint representation to identify semantic offset relationships between different modal generation intents, and a cross-modal consistency violation feature tensor is constructed. Based on this cross-modal consistency violation feature tensor, a multimodal content risk evolution model is constructed, and a risk evolution state quantity characterizing the risk evolution state of the multimodal content is generated according to the rate of change of consistency violation in the cross-modal content over time. Multimodal risk limit state judgment condition matching is performed based on the risk evolution state quantity, and the review result is output.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application, upon receiving multimodal content containing text, images, or audio, performs modal-level parsing processing on the multimodal content to extract modal generation feature vectors that characterize the generation tendency of each modality. Based on the modal generation feature vectors, it performs generation intent inversion to construct a joint representation of multimodal generation intent. It then performs cross-modal consistency analysis on the joint representation of multimodal generation intent to identify semantic offset relationships between different modal generation intents and constructs a cross-modal consistency violation feature tensor. Based on the cross-modal consistency violation feature tensor, it constructs a multimodal content risk evolution model and generates a risk evolution state quantity characterizing the risk evolution state of multimodal content according to the rate of change of consistency violation in the cross-modal content over time. Finally, it performs multimodal risk limit state judgment condition matching based on the risk evolution state quantity and outputs the review result. This invention addresses the technical problem of existing technologies' difficulty in identifying the hidden semantic shifts between different modal generation intentions in multimodal content and the potential risks brought about by their evolution over time. By inverting multimodal generation intentions and identifying cross-modal consistency violations, and combining the rate of change of consistency violations to construct a risk evolution model, this invention achieves the technical effect of dynamically perceiving and determining the risk status of multimodal content, thereby improving the accuracy of review. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A schematic diagram of the structure of an intelligent content review system for multimodal content provided in an embodiment of this application; Figure 2 This is a schematic diagram of the intelligent review method for multimodal content provided in an embodiment of this application.
[0010] Figure labeling: Feature extraction module 11, Intent inversion module 12, Cross-modal consistency violation identification module 13, Risk evolution module 14, Audit decision output module 15. Detailed Implementation
[0011] This application provides an intelligent review system and method for multimodal content, addressing the technical problem that existing technologies struggle to identify hidden semantic shifts between different modal generation intentions in multimodal content and the potential risks arising from their evolution over time. By inverting multimodal generation intentions and identifying cross-modal consistency violations, and constructing a risk evolution model based on the rate of change of consistency violations, this application achieves the technical effect of dynamically perceiving and determining the risk status of multimodal content, thereby improving the accuracy of review.
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0013] It should be noted that any variation of the terms "comprising" and "having" is intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products, or devices.
[0014] Example 1, as Figure 1 As shown in the embodiment of this application, an intelligent review system for multimodal content is provided, the system comprising: The feature extraction module 11 is used to perform modal-level parsing processing on the multimodal content after receiving multimodal content containing text, images or audio, and extract modal generation feature vectors that respectively characterize the generation tendency of each modal content.
[0015] In this embodiment, when receiving multimodal content containing text, images, or audio, the feature extraction module 11 obtains the user-uploaded multimodal content through an access API. During this process, modal-level parsing is performed on the input. By parsing the file structure and extracting the content, the text, image, and audio information in the multimodal content are identified and separated. When the multimodal content is in Markdown or HTML format, the text structure is parsed using markdown-it-py or BeautifulSoup, and the embedded image identifiers are extracted. When the multimodal content contains an audio carrier, the audio track is decapsulated, and the basic acoustic parameters are read. Thus, in this process, original modal data representations that can respectively characterize the text modality, image modality, and audio modality are obtained.
[0016] When processing text modalities in multimodal content, the feature extraction module 11 performs text cleaning and encoding unification through regular expressions during the parsing of the text content. In this process, langdetect is used for language recognition. For Chinese text, jieba combined with a custom sensitive word dictionary is used for word segmentation. For English text, WordPiece is used for sub-word segmentation. Then, the normalized text is input into the text encoding model MacBERT for contextual semantic modeling. The high-level semantic representation of the text is extracted through a multi-layer Transformer structure. In this process, a modality generation feature vector reflecting the text content generation tendency is formed. This modality generation feature vector includes discourse position distribution features to describe the directionality of the text stance and action instruction display features to describe the intensity of text behavior guidance.
[0017] When processing the image modality in multimodal content, the feature extraction module 11 performs image metadata removal, color space unification, and size normalization through Pillow and OpenCV during the parsing of the image content. In this process, it performs OCR text recognition on the image, extracts the embedded text information in the image through PaddleOCR and incorporates it as auxiliary semantics into the text modality analysis process. Then, the normalized image is input into the SigLIP image coding model for visual semantic modeling. The global semantic representation of the image is extracted through a multi-layer attention structure to form a modality generation feature vector that reflects the generation tendency of image content. This modality generation feature vector includes visual composition guidance features for describing the visual attention guidance method and emotional arousal region distribution features for describing the distribution of emotional stimulus regions.
[0018] When processing audio modalities in multimodal content, the feature extraction module 11 performs uniform sampling, amplitude normalization, and time slicing on the audio signal during the parsing process. Time slicing divides the audio signal into multiple consecutive time segments. In this process, the energy value is calculated for each time segment by squaring and summing the amplitude of the audio signal within that time segment. The energy values for each time segment are then arranged chronologically to form an energy change sequence. This sequence is then smoothed to obtain an energy envelope curve. Based on the energy envelope curve, the amplitude, rate, and continuity of energy change between adjacent time segments are calculated. These three parameters together constitute the emotional intensity envelope feature, reflecting the change in audio energy over time. The energy change amplitude index is used to characterize the intensity of audio energy change between adjacent time segments. It is obtained by calculating the absolute value of the difference between the energy values of adjacent time segments and reflects the fluctuation of emotional intensity over time. The energy change rate index is used to characterize the time-oriented characteristics of energy change. It is obtained by dividing the energy change amplitude of adjacent time segments by the corresponding time interval and reflects the speed trend of emotional intensity increasing or decreasing. The energy change continuity index is used to characterize the stability of energy change within multiple consecutive time segments. It is obtained by statistically calculating the consistency of the direction of energy change and the degree of fluctuation of the rate of change in consecutive time segments. During the calculation, within a preset consecutive time window, the direction of energy change of adjacent time segments is compared sequentially. When the direction of energy change of adjacent time segments is the same (i.e., energy continuously increases or decreases), it is recorded as a consistent direction. When the direction changes, it is recorded as a inconsistent direction. Then, the number of consistent directions within the time window is counted and divided by the total number of adjacent time segment pairs within the time window to obtain a ratio value between 0 and 1. This ratio value is used as the consistency of the direction of energy change. Next, within the same time window, the rate of energy change for each time segment is calculated, and the difference between the maximum and minimum values of the rate of energy change within that time window is obtained to represent the magnitude of the fluctuation in the rate of energy change. Then, the magnitude of the fluctuation in the rate of energy change is normalized, mapping it to a numerical range between 0 and 1, resulting in the stability of the rate of energy change; the smaller the fluctuation, the greater the corresponding stability value. Finally, the consistency of the direction of energy change and the stability of the rate of energy change are weighted and summed according to preset weights to obtain a single value. This value serves as the continuity index of energy change for that time window, used to distinguish between sustained emotional expression and transient, sudden energy fluctuations.The feature extraction module 11 concatenates the energy change amplitude index, energy change rate index, and energy change continuity index corresponding to the same time segment in a preset order to form a multi-dimensional feature vector. This multi-dimensional feature vector serves as the envelope feature representation of the emotional intensity corresponding to that time segment. Each component of this multi-dimensional feature vector corresponds to the intensity, speed, and continuity characteristics of audio energy change, respectively. These three components complement each other in depicting the overall shape of audio energy change over time from different perspectives, thus forming a unified representation of the contour of audio emotional intensity change.
[0019] Simultaneously, based on time-slicing processing, the time positions corresponding to adjacent cycles in the audio signal are detected, and the time interval values between corresponding time positions of adjacent cycles are calculated to obtain a time interval sequence. Then, the average time interval value is calculated for the time interval sequence, and the deviation of each time interval value from the average time interval value is calculated. The deviations are accumulated and normalized to obtain a rhythm stability index used to characterize the degree of fluctuation of the time interval values. At the same time, the frequency of occurrence of the same or similar time interval values in the time interval sequence is statistically analyzed, and the statistical results are normalized with the length of the time interval sequence to obtain a rhythm repeatability index. Then, the feature extraction module 11 concatenates the rhythm stability index and rhythm repeatability index corresponding to the same time window in a preset order to form a multi-dimensional feature vector. This multi-dimensional feature vector serves as the rhythm-inducing feature representation corresponding to that time window. Each component of the rhythm-inducing feature vector reflects the stability and repeatability characteristics of the audio rhythm in the time dimension in numerical form. These two components complement each other in characterizing the guiding effect of the audio rhythm structure on emotions or behaviors from different perspectives, thus forming a unified representation of the audio rhythm-inducing ability. The aforementioned emotion intensity envelope feature and rhythm-inducing feature together constitute the modality generation feature vector of the audio modality.
[0020] After completing the feature extraction processing of each modality in the multimodal content, the feature extraction module 11 encapsulates and outputs the modality generation feature vectors of the text modality, the image modality, and the audio modality, respectively, to obtain modality generation feature vectors that represent the generation tendency of each modality content.
[0021] Furthermore, in the system provided in the application embodiment, the modality generation feature vector in the feature extraction module is a high-level semantic feature vector used to characterize the generation tendency of the corresponding modality content. The generation tendency includes behavioral induction tendency, emotional arousal tendency, and stance orientation tendency, wherein: the modality generation feature vector of the text modality includes discourse stance distribution features and action instruction display features; the modality generation feature vector of the image modality includes visual composition guidance features and emotional arousal area distribution features; and the modality generation feature vector of the audio modality includes emotional intensity envelope features and rhythm induction features.
[0022] In this embodiment, the modality generation feature vector within the feature extraction module is a high-level semantic feature vector characterizing the content generation tendency of the corresponding modality. This generation tendency includes behavioral induction tendency, emotional arousal tendency, and stance orientation tendency. Specifically, for the text modality, by performing contextual semantic analysis on the text content, high-level semantic information is extracted from the semantic distribution of the text at the levels of value judgment, viewpoint expression, and behavioral cues. This forms discourse stance distribution features to characterize the directionality of the text's stance expression and action instruction visibility features to characterize the salience of behavioral instructions or operational cues in the text. These features collectively constitute the modality generation feature vector of the text modality.
[0023] For image modalities, semantic analysis is performed on the visual structure, composition layout, and emotional stimulation regions of the image to form visual composition guidance features that describe how the image composition guides visual attention and emotional arousal region distribution features that describe the spatial distribution of regions with emotional stimulation effects in the image. These features together constitute the modality generation feature vector of the image modality.
[0024] For audio modalities, by analyzing the energy changes and rhythmic characteristics of audio signals in the time dimension, we form an emotion intensity envelope feature to describe the contour of the change of audio emotion intensity over time, and a rhythmic induction feature to describe the guiding effect of audio rhythmic structure on emotions or behaviors. These features together constitute the modality generation feature vector of audio modalities, so that the modality generation feature vectors of different modalities can represent the generation tendency of corresponding modal content at a unified semantic level.
[0025] Furthermore, the system provided in the application embodiments also includes: The preprocessing module is used to standardize and clean the received multimodal content, which includes text, images, or audio, and unify its format, and encapsulate it into a structured JSON format.
[0026] In this embodiment, after receiving multimodal content containing text, images, or audio, the preprocessing module sequentially performs standardized cleaning, format unification, and structured encapsulation on the multimodal content. First, it obtains the multimodal content input through an interface access method and performs file type identification and modality parsing on the input data. During this process, text modality, image modality, and audio modality are distinguished based on file header information and content features to obtain the corresponding original modal data.
[0027] Subsequently, text cleaning is performed on the text modality by removing irrelevant markers, control characters, and redundant formatting, and standardizing character encoding to make the text content conform to the input specifications for subsequent semantic analysis. At the same time, image standardization is performed on the image modality, including removing additional metadata, standardizing color space and resolution specifications to ensure consistency in numerical representation of images from different sources. Audio standardization is performed on the audio modality, including standardizing sampling rate, channel structure, and amplitude range to ensure the comparability of audio signals in the time and amplitude domains.
[0028] After completing the standardized cleaning and format unification of each modality, the processed text content, image information, and audio information, along with the corresponding modality identifiers, timestamps, and source markers, are uniformly organized and encapsulated into a structured JSON format according to a preset field structure. The structured JSON format is used to carry the cleaning results and metadata of multimodal content in a unified data structure.
[0029] The intent inversion module 12 is used to perform intent inversion based on the modality generation feature vector to construct a joint representation of multimodal intent generation.
[0030] In this embodiment, when the intent inversion module 12 receives modality generation feature vectors representing the content generation tendencies of text modality, image modality, and audio modality respectively, it first performs semantic space unification processing on the modality generation feature vectors. During this process, considering the differences in feature dimension and numerical distribution among the different modality generation feature vectors, linear projection layers are established for text modality, image modality, and audio modality respectively. Each linear projection layer contains a corresponding weight matrix and bias term, where the weight matrix and bias term are trainable parameters and their initial values are generated by a preset initialization strategy. The initialization strategy includes using random initialization to generate initial values for the weight matrix and setting the initial values for the bias term to zero. Subsequently, the weight matrix and bias term are iteratively optimized based on training data. During the linear projection process, the modality generation feature vector of the corresponding modality is input into the linear projection layer. Dimension mapping is completed through matrix multiplication and bias superposition to obtain a projection vector of unified dimension. Normalization processing is then performed on the projection vector to constrain the numerical scale, thereby obtaining a standardized modality generation feature vector located in a unified semantic coordinate system.
[0031] Next, the intent inversion module 12 performs semantic factor parsing and decomposition on the standardized modality-generated feature vector. In this process, behavioral induction tendency, emotional arousal tendency, and stance orientation tendency are used as the set of semantic factors for generated intent. Specifically, the corresponding feature dimension ranges are pre-determined for behavioral induction tendency, emotional arousal tendency, and stance orientation tendency. The determination of the feature dimension ranges is based on statistical analysis results from historical labeled samples, achieved by comparing the differences in the numerical distribution of each feature dimension between samples with and without corresponding intent labels. For example, the feature dimension values can be normalized, and the difference in their average values between samples with and without corresponding intent labels can be calculated. When this difference exceeds a preset reference threshold (e.g., 0.25), it can be determined that the feature dimension has a strong semantic correlation with the corresponding generated intent tendency, thus classifying the feature dimension into the corresponding feature dimension range. The feature dimension range indicates which dimensions of the standardized modality-generated feature vector are used to represent the corresponding generated intent semantic factors. Subsequently, the intent inversion module 12 sequentially reads the values of each feature dimension from the standardized modality-generated feature vector, and extracts the feature dimension values belonging to behavioral induction tendency, emotional arousal tendency, and stance orientation tendency according to the feature dimension range. Then, for each extracted feature dimension value, a weighted operation is performed sequentially according to the order of the corresponding feature dimension in the vector. The weighting operation is performed by multiplying each feature dimension value with its corresponding preset weight coefficient and accumulating the product results one by one, thereby obtaining the projection intensity values of behavioral induction tendency, emotional arousal tendency, and stance orientation tendency in the standardized modality-generated feature vector. Then, the feature dimension values used to calculate the projection intensity values are combined in their original order and output as semantic factor components corresponding to behavioral induction tendency, emotional arousal tendency, and stance orientation tendency, respectively. This achieves the decomposition of the standardized modality-generated feature vector into multiple semantic factor components corresponding to different behavioral induction tendency, emotional arousal tendency, and stance orientation tendency. In the text modality, the distribution features of discourse stance and the display features of action instructions serve as semantic evidence of stance orientation and behavior induction tendencies, respectively. In the image modality, the visual composition guidance features and the distribution features of emotional arousal regions serve as semantic evidence of behavior induction and emotion arousal tendencies, respectively. In the audio modality, the emotional intensity envelope features and rhythmic induction features serve as semantic evidence of emotion arousal and behavior induction tendencies, respectively. This completes the semantic factor decomposition of the generation tendency information of different modalities.
[0032] After completing the semantic factor decomposition, the intent inversion module 12 performs cross-modal recombination processing on the semantic factor components. In this process, semantic factor components from different modalities but corresponding to the same generation tendency are aligned and aggregated. The alignment process uses the generation intention tendency type as the alignment index and normalizes and weights the semantic factor components of each modality according to their corresponding projection intensity values to form a generation intention semantic factor representation under a unified scale. This enables the behavior induction tendency, emotion arousal tendency, and stance orientation tendency to form cross-modal consistent generation intention semantic factor representations, thereby integrating the explicit generation tendency features of multiple modalities into a structured and alignable semantic factor representation.
[0033] Next, the intent inversion module 12 performs the intent generation inversion calculation. During this process, the semantic factor representation of the generated intent is input into the intent generation inversion network. The network uses a multilayer perceptron to process the semantic factor representation to output the generated intent representation. Specifically, the semantic factor representation of the generated intent is input as an input vector to the multilayer perceptron. A linear transformation operation is performed on the input vector. This linear transformation operation obtains an intermediate calculation result by multiplying the input vector by the corresponding weight parameters and adding the bias parameters. A nonlinear activation function is then applied to the intermediate calculation result to obtain the output result of the multilayer perceptron, which is then used as the generated intent representation for the corresponding modality. The weight parameters and bias parameters of the multilayer perceptron are trainable parameters obtained through iterative optimization using training data. During training, the target intent label or risk label given in the platform's historical review logs or labeled data are used as supervision signals. The generated intent representation output by the Generative Intent Inversion Network is mapped to the target intent label or risk label, and the error between the predicted value and the target value is calculated using the mean squared error method. The error is obtained by squaring the difference between the generated intent representation and the target label in the corresponding dimension and summing them. Then, based on the error, the weight parameters and bias parameters in the multilayer perceptron are updated through the backpropagation algorithm, so that the Generative Intent Inversion Network gradually reduces the difference between the predicted value and the target value, thereby learning the mapping relationship between the semantic factor representation of the generated intent and the generated intent representation. After training, the output of the Generative Intent Inversion Network is directly used as the single-modal generated intent representation output corresponding to the text modality, image modality and audio modality, respectively.
[0034] Subsequently, the intent inversion module 12 performs cross-modal attention weighted fusion to construct a joint representation of multimodal generated intent. During this process, a cross-modal attention network is constructed and its parameters are set to trainable parameters. Attention weights are obtained by calculating the similarity between the single-modal generated intent representation of one modality (as the query vector) and the single-modal generated intent representations of other modalities (as the key vectors) and the numerical vectors. Specifically, during similarity calculation, the query vector and the corresponding key vector are first multiplied element-wise along the same dimension, and the result of the element-wise multiplication is summed along the vector dimension to obtain a similarity value representing the degree of correlation between the query vector and the key vector. The similarity value measures the consistency between the corresponding single-modal generated intent and the overall multimodal content generated intent; a higher similarity value indicates a higher correlation between the corresponding modal generated intents. Next, the attention weights are subjected to Softmax normalization to ensure that the contribution weights of each modal generated intent in the overall intent fusion process form a normalized distribution. Subsequently, the numerical vectors are weighted and summed based on attention weights to obtain a fusion vector. This fusion vector is used to represent the overall generative intent representation of multimodal content in a unified semantic space, making it consistent with the overall intent of the multimodal content. Figure 1 Stronger consistency in single-modal generated intent representations results in higher fusion weights, while weaker consistency is reduced. Specifically, during computation, the attention weight corresponding to each single-modal generated intent representation is multiplied element-wise with its numerical vector to obtain a weighted single-modal generated intent vector. These weighted vectors are then summed along their vector dimensions to form a fusion vector. The training of the cross-modal attention network is also based on historical platform review logs or labeled data. By minimizing the prediction error of the fused generated intent joint representation on downstream tasks and using backpropagation to update the attention network parameters, a more consistent fused representation of generated intent across modalities is obtained.
[0035] Finally, the intent inversion module 12 performs joint representation construction processing on the generated intent fusion representation. Through vector concatenation and linear transformation, the fusion representation is mapped to a high-level semantic vector of fixed dimension, resulting in a multimodal generated intent joint representation that can comprehensively represent the multimodal content generation motivation, expression purpose and intent structure.
[0036] The cross-modal consistency violation identification module 13 is used to perform cross-modal consistency analysis on the joint representation of the multimodal generation intent, identify the semantic offset relationship between different modal generation intents, and construct a cross-modal consistency violation feature tensor.
[0037] In this embodiment, when performing cross-modal consistency violation identification module 13 performs cross-modal consistency analysis on the joint representation of multimodal generated intentions, it first performs modal-level intention parsing processing on the joint representation of multimodal generated intentions. By deconstructing the joint representation of multimodal generated intentions according to a pre-agreed semantic encoding structure, the generated intention components corresponding to the text modality, image modality, and audio modality are parsed into independent generated intention vectors. Through index mapping and dimension alignment, the generated intention vectors of the text modality, image modality, and audio modality are placed in the same unified semantic coordinate space.
[0038] Next, the cross-modal consistency violation identification module 13 quantitatively models the difference in intent direction between different modal generation intents. In this process, any two modal generation intent vectors in a unified semantic coordinate space are used as direction vectors. The cosine similarity between the two vectors is calculated to characterize the degree of direction consistency. Then, the cosine similarity is converted into a directional offset value. The conversion process adopts a differential mapping method, defining the directional offset value as one minus the cosine similarity and limiting it to a preset value range. When the directions of the two modal generation intent vectors are completely consistent, the directional offset value is zero. When the difference in the directions of the two modal generation intent vectors increases, the directional offset value increases accordingly. Thus, the intent direction difference feature can quantitatively characterize the difference in direction of different modal generation intents at the level of position orientation or expression target.
[0039] The cross-modal consistency violation identification module 13 then continues to calculate the gradient difference in intent intensity between different modal intentions. During this process, for each modal intention vector within a unified semantic coordinate space, the corresponding intent intensity index is obtained by calculating the vector norm of the intention vector. The vector norm characterizes the intensity level of the intention in the overall semantic space. The intention vector consists of multiple feature dimensions. When calculating the vector norm, the values of the intention vector in each feature dimension are read sequentially, and a square operation is performed on each feature dimension value to obtain a squared value. All squared values are then summed to obtain a sum of squares, and the square root of the sum of squares is performed to obtain the vector norm. Subsequently, a difference operation is performed on the intent intensity indices of different modal intention vectors. This is achieved by taking the difference between the intent intensity indices corresponding to any two modalities and taking the absolute value of the difference to obtain the intent intensity gradient difference feature. The intent intensity gradient difference feature increases with the increase in the intensity difference of intentions generated by different modalities, reflecting the degree of inconsistency in the intensity of behavioral induction or emotional arousal tendencies among different modalities.
[0040] Subsequently, the cross-modal consistency violation identification module 13 models the differences in the direction of intent evolution trends among different modal intentions. In this process, considering the continuous changes in the joint representation of multimodal intentions over time, a sliding window method is used to serialize and organize the joint representations of multimodal intentions at different times. Difference operations are performed on the intention vectors of each modality within adjacent time windows to obtain the corresponding intention evolution trend vectors. These intention evolution trend vectors characterize the direction of change of the generated intent over time. Then, the cosine similarity between the intention evolution trend vectors of different modalities is calculated, and a differential mapping method consistent with the difference in the direction of intent direction is used to convert the cosine similarity into a feature of the difference in the direction of intent evolution. This quantitatively characterizes whether different modal intentions exhibit synchronous changes or directional shifts at the level of evolution trend.
[0041] After obtaining the intention direction difference features, intention intensity gradient difference features, and intention evolution trend direction difference features, the cross-modal consistency violation identification module 13 organizes the above difference features in a structured manner according to the preset tensor construction rules. In this process, the intention direction difference features, intention intensity gradient difference features, and intention evolution trend direction difference features calculated between different modal pairs are spliced along the feature dimension, and the intention direction difference, intention intensity gradient difference, and intention evolution trend direction difference are encoded as independent channels, thereby constructing a cross-modal consistency violation feature tensor.
[0042] Furthermore, the system provided in the application embodiments also includes: In the cross-modal consistency violation identification module, the cross-modal consistency violation feature tensor is constructed by modeling the directional offset of different modal generation intentions in a unified semantic coordinate space. The directional offset includes differences in intention pointing direction, differences in intention intensity gradient, and differences in intention evolution trend direction.
[0043] In this embodiment, in the cross-modal consistency violation identification module 13, the cross-modal consistency violation feature tensor is constructed by modeling the directional offset of different modal generation intentions in a unified semantic coordinate space. The cross-modal consistency violation feature tensor is a multi-dimensional numerical tensor, and its element values are composed of the difference calculation results between different modal generation intentions, used to numerically represent the multi-modal generation intentions. Figure 1The degree of consistency disruption is determined by directional offset, which characterizes the semantic shift between different modal generation intentions. This includes differences in intention direction, intention intensity gradient, and intention evolution trend. Intention direction difference represents the semantic shift in the stance or target of different modal generation intentions. Intention intensity gradient difference represents the inconsistency in the gradient of behavioral induction or emotional arousal tendencies. Intention evolution trend difference represents the trend of different modal generation intentions over time in a unified semantic coordinate space. These differences are concatenated along the feature dimension according to a preset feature order and encoded as independent semantic channels. This forms a cross-modal consistency disruption feature tensor with a clear mathematical structure across the modal, feature, and semantic channel dimensions. The comprehensive modeling of these directional offsets forms a feature tensor for characterizing multimodal generation intentions. Figure 1 Cross-modal consistency violation feature tensor of consistency violation state.
[0044] The risk evolution module 14 is used to construct a multimodal content risk evolution model based on the cross-modal consistency destruction feature tensor, and generate a risk evolution state quantity that characterizes the risk evolution state of multimodal content according to the rate of change of consistency destruction of cross-modal content in the time dimension.
[0045] In this embodiment, the cross-modal consistency violation feature tensors are first sequentially organized according to time order. Serialization refers to arranging the cross-modal consistency violation feature tensors calculated at different times according to the timestamp order corresponding to the multimodal content generation or sampling, forming a feature tensor sequence ordered by time index. Each item in the sequence is used to characterize the multimodal generation intent within the corresponding time point or time segment. Figure 1The state of consistency violation is determined. A sliding window method is employed to segment the cross-modal consistency violation features within consecutive time periods. The sliding window slides across the feature tensor sequence according to a preset window length and step size, dividing the cross-modal consistency violation feature tensors corresponding to multiple consecutive time points into feature subsequences within the same time window, ensuring that the consistency violation features within each time period can form a relatively stable temporal segment representation. Next, a difference operation is performed on the cross-modal consistency violation feature tensor within each sliding window to calculate the rate of change of consistency violation in the time dimension. This rate of change reflects the change in the degree of cross-modal consistency violation between adjacent time segments. Specifically, during the difference operation, within the same sliding window, the cross-modal consistency violation feature tensors corresponding to two adjacent time segments are read sequentially according to time. The values of the two time segments are subtracted in each corresponding feature dimension to obtain the changes in each feature dimension between adjacent time segments. The changes were then normalized according to the corresponding time intervals to obtain the change values of each feature dimension per unit time, and the change values per unit time were used as the rate of change of cross-modal consistency violation in the time dimension within the sliding window.
[0046] The rate of change of consistency disruption is then recursively updated with the risk evolution state quantity corresponding to the previous time segment to form a multimodal content risk evolution model. The multimodal content risk evolution model outputs risk evolution state quantities that include risk level, risk growth rate and risk trend indicators.
[0047] Furthermore, in the system provided in the application embodiment, the risk evolution module 14 is also used for: A sliding window construction unit is used to construct a sliding window based on the cross-modal consistency violation feature tensor according to the time series, so as to capture the dynamic changes of multimodal content consistency violation in a continuous time period; a change rate calculation unit is used to calculate the time derivative of the cross-modal consistency violation tensor within the sliding window to obtain the cross-modal consistency violation change rate, which is used to quantify the dynamic trend of risk evolution; a multimodal risk evolution model unit is used to couple the cross-modal consistency violation change rate with historical risk evolution state quantities to construct a multimodal content risk evolution model, wherein the multimodal content risk evolution model is used to predict the evolution path of multimodal content risk in the future time period based on the current cross-modal consistency violation change rate; a risk evolution state quantity generation unit is used to output risk evolution state quantities characterizing the multimodal content risk evolution state according to the multimodal content risk evolution model, wherein the risk evolution state quantities include risk level, risk growth rate, and risk trend indicators.
[0048] In this embodiment, the sliding window construction unit arranges the cross-modal consistency violation feature tensors in chronological order and constructs a sliding window based on preset window length and step size parameters. In this process, the cross-modal consistency violation feature tensors at consecutive time points are divided into several consecutive time periods, so that the cross-modal consistency violation feature tensors in each sliding window can fully reflect the changes in the consistency violation state of multimodal content in the corresponding time period.
[0049] Next, the rate of change calculation unit performs time derivative calculation on the cross-modal consistency violation feature tensor within the sliding window. In this process, a discrete-time difference method is used to perform difference operations on the cross-modal consistency violation feature tensors corresponding to adjacent time points within the sliding window, and normalization is performed in combination with the time interval between adjacent time points to obtain the cross-modal consistency violation rate of change. This cross-modal consistency violation rate of change is used to quantify the magnitude and direction of change of the difference in intention direction, the difference in intention intensity gradient, and the difference in intention evolution trend direction in the time dimension, thereby forming a quantitative description of the dynamic trend of risk evolution.
[0050] Subsequently, the multimodal risk evolution model unit couples the cross-modal consistency disruption rate with historical risk evolution state variables to construct a rate-driven multimodal content risk evolution model. In this process, through a recursive state update mechanism, the historical risk evolution state variables are used as the risk state input for the previous time period, and the cross-modal consistency disruption rate corresponding to the current time period is used as the state update driver. The cross-modal consistency disruption rate is used to characterize the degree of change of cross-modal consistency disruption features in the time dimension. Its magnitude reflects the extent of enhancement or weakening of consistency disruption, and its direction of change reflects the trend direction of consistency disruption evolution. On this basis, the risk state is updated by superimposing the risk change caused by the current cross-modal consistency disruption rate on the risk state of the previous time period. The risk change is the numerical change result obtained by incrementally adjusting the risk evolution state variables according to the cross-modal consistency disruption rate, which is used to characterize the change of risk in the current time period relative to the previous time period. This process yields a risk latent state representation for the current time period. This risk latent state representation is a numerical representation of the comprehensive risk status of multimodal content within the current time period, enabling continuous updates of multimodal content risk over time and supporting the projection of risk change trends in subsequent time periods.
[0051] After completing the construction of the multimodal content risk evolution model, the risk evolution state quantity generation unit generates risk evolution state quantities to characterize the multimodal content risk evolution state based on the output results of the multimodal content risk evolution model. In this process, the risk latent state representation of the current time period is first subjected to numerical normalization processing so that the risk latent state representation falls into a predefined risk scale interval. The risk scale interval is a pre-set continuous numerical range used to map different risk states to a unified numerical scale. For example, the risk latent state representation is linearly mapped to the interval between 0 and 1 so that interval judgment and comparison can be performed later.
[0052] The normalized risk latent state representation is then matched against preset risk threshold intervals. Interval matching is achieved by determining which risk threshold interval the normalized risk latent state representation falls into. Each risk threshold interval corresponds to a different risk level in order of numerical magnitude. The corresponding risk level is determined based on the interval position of the risk latent state representation. For example, in one exemplary implementation, the risk scale interval can be divided into multiple consecutive sub-intervals. When the normalized risk latent state representation falls into the interval of 0 to 0.3, it is determined to be a level 1 risk; when it falls into the interval of 0.3 to 0.6, it is determined to be a level 2 risk; and when it falls into the interval of 0.6 to 1, it is determined to be a level 3 risk.
[0053] Next, by calculating the difference between the current risk latent state representation and the corresponding risk latent state representation in the historical risk evolution state, and normalizing it by combining the time intervals of adjacent time periods, a risk growth rate is obtained to characterize the speed of risk change. Subsequently, within a preset continuous time window, risk latent state representations corresponding to multiple time periods are obtained in chronological order, and difference operations are performed on the risk latent state representations of adjacent time periods to form a difference sequence reflecting the direction of risk state change over time. Based on the consistency of the sign and amplitude characteristics of each difference value in the difference sequence, the change pattern of the risk state is determined. Specifically, when the difference value is positive for multiple consecutive time periods, the risk state is determined to show a continuously rising change pattern; when the difference value is negative for multiple consecutive time periods, the risk state is determined to show a continuously falling change pattern; when the difference value fluctuates slightly around zero for multiple consecutive time periods and its absolute value is less than a preset stability threshold, the risk state is determined to show a relatively stable change pattern. Thus, a risk trend indicator is obtained to characterize the overall development direction of risk. Finally, the risk level, risk growth rate, and risk trend indicator are uniformly encapsulated to generate a complete risk evolution state quantity.
[0054] Furthermore, in the system provided in the application embodiments, the multimodal content risk evolution model is a risk state space model driven by the rate of change, and also includes: A state initialization subnetwork is used to temporally encode the cross-modal consistency violation feature tensor within the sliding window to generate initial values of the risk hidden states. The temporal encoding includes temporal positional encoding of the cross-modal consistency violation feature tensor and cross-modal channel attention convergence to obtain a window representation vector characterizing the consistency violation morphology. The initial values of the risk hidden states are then mapped from this window representation vector. A rate-of-change gated state update subnetwork is used to recursively update the risk hidden states using the cross-modal consistency violation feature tensor as a gate input to form a risk evolution trajectory. A multi-step risk prediction subnetwork is used to output a sequence of risk evolution state quantities based on the updated risk evolution trajectory. An online calibration and drift suppression subnetwork is used to construct calibration sample pairs based on the audit results and perform confidence constraint updates on the gate vector of the rate-of-change gated state update subnetwork, thereby suppressing abnormal transitions of the risk hidden states when there are noise fluctuations in the rate of change of cross-modal consistency violation.
[0055] In this embodiment, after receiving the cross-modal consistency violation feature tensor within the sliding window, the state initialization subnetwork performs temporal encoding processing on it to generate initial values for the hidden risk states. In this process, firstly, temporal positional encoding is introduced into the cross-modal consistency violation feature tensor along the time dimension. By assigning corresponding position vectors to different time indices and superimposing them element-wise with the original cross-modal consistency violation feature tensor, the cross-modal consistency violation features explicitly include temporal order information in the feature representation. Subsequently, cross-modal channel attention convergence processing is performed along the cross-modal channel dimension. In this process, attention weights are calculated based on the importance of different modal consistency violation features within the current time window. These attention weights are then used to weight and fuse differences in intent direction, intent intensity gradient, and intent evolution trend direction, thereby obtaining a window representation vector that can comprehensively represent the consistency violation pattern within the current sliding window. After obtaining the window representation vector, the window representation vector is converted into the initial value of the risk hidden state through vector space transformation. The vector space transformation performs dimension alignment and scale normalization on the window representation vector to make it consistent with the risk hidden state space in terms of numerical range and dimensional structure. The window representation vector is then projected onto the risk hidden state space through a learnable vector mapping relationship to obtain the initial value of the risk hidden state for subsequent recursive updates.
[0056] Next, the rate-gated state update subnetwork uses the cross-modal consistency violation rate as the gating driver to recursively update the risk hidden state to form the risk evolution trajectory. In this process, the recursive update of the risk hidden state satisfies the state update relationship driven by the rate of change. Its update form is manifested as the superposition of the risk increment modulated by the consistency violation rate of change on the risk hidden state at the previous time step. The gating vector generated by the cross-modal consistency violation rate of change is used to adaptively adjust the update amplitude of the risk hidden state, while the risk increment is generated by the combined effect of the current risk hidden state and the consistency violation feature, and is used to characterize the nonlinear projection change of consistency violation in the risk hidden state space. The calculation process for the risk increment involves first adding the current risk latent state and the cross-modal consistency violation features of the current time period one-to-one according to the same dimension to obtain a combined intermediate vector. Then, a linear transformation operation is performed on each component of the intermediate vector, which involves multiplying the component by its corresponding weight coefficient and adding a corresponding bias term. The weight coefficient characterizes the response strength of the risk dimension to the impact of consistency violation, and the bias term characterizes the basic change in the risk dimension when the consistency violation feature is zero. The weight coefficient and bias term are preset by technical experts during the initialization phase. Next, a nonlinear mapping process is performed on each component of the linearly transformed vector. The nonlinear mapping process uses a fixed Sigmoid compression operation, where each component is substituted into the Sigmoid function to obtain a value falling within the 0-1 interval. The resulting vector is the risk increment. By modulating the risk increment element-wise using a gating vector, the risk latent state can be dynamically adjusted according to the strength and direction of the consistency violation rate during the recursive update process, thus forming a continuous and controllable risk evolution trajectory in the time dimension.
[0057] Subsequently, the multi-step risk prediction subnetwork performs multi-time-step extrapolation of the risk latent state based on the updated risk evolution trajectory. In this process, the current risk latent state is used as the starting state, and multiple recursive updates are performed on the risk latent state according to a preset number of prediction steps to obtain the predicted risk latent states for multiple future time steps. Then, risk state analysis is performed on each predicted risk latent state, mapping it to a risk evolution state quantity. The risk level is determined by the position interval of the risk latent state in the risk scale space, the risk growth rate is calculated by the change amplitude between risk latent states in adjacent prediction time steps, and the risk trend indicator is judged by the consistency of the change direction of the predicted risk latent state across multiple time steps. This forms a sequence of risk evolution state quantities containing risk levels, risk growth rates, and risk trend indicators for multiple time steps.
[0058] Finally, the online calibration and drift suppression subnetwork constructs calibration sample pairs based on the feedback annotations of the audit results during model operation. By comparing the deviation between the risk evolution state variables and the corresponding audit results, confidence constraints are applied to the gating vectors used in the rate-of-change gating state update process. When the rate of change that violates consistency experiences short-term fluctuations or noise disturbances, the excessive changes in the risk hidden states are suppressed by reducing the update strength of the gating vectors, thereby avoiding abnormal transitions in the risk hidden states during the recursive update process and ensuring the stability and continuity of the risk evolution trajectory over a long timescale.
[0059] Furthermore, in the system provided in the application embodiments, the recursive update satisfies: ;in, The recursive update results characterizing the hidden risk states. The initial value is the hidden state of risk. This is a gated vector generated from the rate of change of cross-modal consistency violation, used to adaptively adjust the magnitude of risk state updates. This is the risk increment generation function, used to characterize the nonlinear projection increment of consistency violation in the risk state space.
[0060] In this embodiment, the recursive update of the hidden risk state follows a change rate-driven state update formula. The meaning and role of each item in the recursive process are reflected through the calculation steps. Specifically, the hidden risk state at the previous moment is first... As the base state for the current iterative update, the risk hidden state It is used to characterize the risk structure and risk level that multimodal content accumulates over a historical period.
[0061] Subsequently, based on the rate of change of cross-modal consistency violation... Generate gate vector The gating vector is obtained by performing numerical mapping and normalization on the rate of change. Its value range is used to characterize the modulation capability of the cross-modal consistency violation rate of change on the magnitude of risk latent state updates, thereby achieving adaptive control of the intensity of risk state updates. Then, the risk increment generation function is used... Current risk status Features of cross-modal consistency violation Joint modeling is performed, and the risk increment generation function is used to characterize the nonlinear projection result of consistency violation in the risk latent state space, generating risk increments that reflect the direction and magnitude of the impact of consistency violation on the risk state.
[0062] Then the gating vector With risk increment Perform element-wise multiplication This allows the intensity of the risk increment's effect across all risk dimensions to be finely modulated by the rate of change of cross-modal consistency disruption.
[0063] Finally, the gated-modulated risk increment is compared with the risk hidden state at the previous moment. By performing element-by-element superposition, we obtain the risk hidden state update result at the current moment. This allows for the recursive updating process of the risk latent state, which dynamically evolves with the rate of change of cross-modal consistency destruction, while maintaining the temporal continuity of the risk state.
[0064] Furthermore, the system provided in the application embodiments also includes: The multimodal content risk evolution model is trained using a joint loss function, which includes a risk state quantity prediction error term, a risk growth rate consistency constraint term, and a risk trend direction consistency constraint term.
[0065] In this embodiment, the training of the multimodal content risk evolution model employs a joint loss function for parameter optimization. The training process is based on risk evolution state quantities organized in time series form, and incorporates supervision signals constructed from the feedback annotations of review results. First, for the risk evolution state quantities output by the model at each time step, a supervised learning approach is used to calculate the risk state quantity prediction error term. During this process, the risk level and corresponding latent risk state representation output by the model are aligned with the actual risk state quantities given in the feedback annotations of the review results. The actual risk state quantities are obtained through statistical summarization of historical review results and mapped to a unified risk scale space. The risk scale space is a predefined continuous numerical interval used to numerically represent different risk levels and risk states, enabling comparison between the model output and the actual annotations on the same numerical scale. Subsequently, the risk state quantity prediction error term is obtained by calculating the numerical difference between the model's predicted risk state representation and the actual risk state quantity in the risk scale space. This risk state quantity prediction error term is used to measure the model's prediction deviation in characterizing the current risk level of multimodal content.
[0066] After calculating the risk state prediction error term, the predicted risk growth rate is calculated based on the risk evolution state quantities output by the model at adjacent time steps. This process involves performing a difference operation on the risk state quantities in the risk scale space at adjacent time steps, and then normalizing the data by combining the time intervals between corresponding time steps to obtain the model's predicted risk growth rate. Subsequently, the predicted risk growth rate is compared with the actual risk growth rate obtained statistically from historical audit data in the risk scale space. The numerical difference between the two is used to form a risk growth rate consistency constraint term. This constraint term constrains the model's ability to learn the magnitude of risk changes over time, ensuring that the model remains consistent with the actual risk evolution process in terms of the rate of change of risk states.
[0067] Subsequently, risk trend direction information is extracted from the risk evolution state quantities at multiple consecutive time steps. This process involves analyzing the direction of change of risk state quantities over time in the risk scale space to determine whether the model-predicted risk state quantities exhibit a continuous upward, continuous downward, or relatively stable change on the time axis, thus obtaining the predicted risk trend direction. The predicted risk trend direction is then compared with the actual risk trend direction obtained statistically from historical audit results in the risk scale space. A consistency constraint term is calculated to constrain the model's judgment of the overall risk development direction over a longer time span, preventing the model from responding only to short-term fluctuations while ignoring the overall trend of risk evolution.
[0068] After calculating the error term for risk state prediction, the consistency constraint term for risk growth rate, and the consistency constraint term for risk trend direction, the three types of loss terms are weighted and summed according to preset weights to form a joint loss function. The weight coefficients of each loss term are adjusted through hyperparameter configuration during the training phase. In the early stage of model training, the initial weight ratios are set according to the priority of the business side regarding the accuracy of risk state prediction, the ability to characterize the rate of risk change, and the stability of risk trend judgment. During training, the weight coefficients are iteratively adjusted based on the loss convergence results on a pre-prepared validation dataset. When a certain loss term is dominant for a long time or its convergence speed is significantly slow, its weight is reduced or increased accordingly, thereby achieving a dynamic balance constraint on the accuracy of risk state prediction, the consistency of risk growth rate, and the consistency of risk trend direction. Subsequently, a gradient descent-based optimization method is used to minimize the joint loss function, and the gradient of the joint loss function is passed to the parameters of the multimodal content risk evolution model through a backpropagation mechanism, so that the model gradually converges during training, thereby obtaining a multimodal content risk evolution modeling capability with synergistic consistency in risk state, risk change rate, and risk evolution direction.
[0069] The audit decision output module 15 is used to match the multimodal risk limit state judgment conditions based on the risk evolution state quantity and output the audit result.
[0070] In this embodiment of the application, when the audit decision output module 15 performs multimodal risk limit state determination condition matching based on the risk evolution state quantity, it first analyzes the risk level, risk growth rate and risk trend indicators contained in the risk evolution state quantity, and then performs a comprehensive matching with the preset multimodal risk limit state determination conditions. The multimodal risk limit state determination conditions are composed of risk level threshold conditions, risk growth rate threshold conditions and risk trend consistency conditions.
[0071] When the risk evolution state quantity meets the multimodal risk limit state judgment condition, the review decision output module 15 outputs the corresponding review restriction result. When the risk evolution state quantity does not meet the multimodal risk limit state judgment condition, the review decision output module 15 outputs the pass or prompt review result, thereby realizing multimodal content review decision output based on the risk evolution state quantity.
[0072] Furthermore, in the system provided in the application embodiment, the review decision output module, which performs multimodal risk limit state determination condition matching based on the risk evolution state quantity, further includes: The risk evolution state quantity is matched and judged with the preset multimodal risk limit state judgment conditions, wherein the multimodal risk limit state judgment conditions include risk level threshold conditions, risk growth rate threshold conditions, and risk trend consistency conditions; when the risk evolution state quantity meets the multimodal risk limit state judgment conditions, the corresponding review restriction result is output; when the risk evolution state quantity does not meet the multimodal risk limit state judgment conditions, the review result of approval or prompt is output.
[0073] In this embodiment, a matching judgment is first performed based on the risk evolution state quantity and the preset multimodal risk limit state determination conditions. During this process, the risk evolution state quantity is structurally analyzed to extract the risk level, risk growth rate, and risk trend indicators contained therein. Subsequently, a level matching judgment is performed according to the risk level threshold conditions in the multimodal risk limit state determination conditions. During this process, the risk level in the risk evolution state quantity is compared with the preset risk level threshold to determine whether the risk level of the current multimodal content has reached or exceeded the level boundary corresponding to the multimodal risk limit state, thereby forming a matching result for the risk level threshold conditions.
[0074] After the risk level threshold condition is determined, the rate matching judgment is performed according to the risk growth rate threshold condition in the multimodal risk limit state determination condition. In this process, the risk growth rate in the risk evolution state quantity is compared with the preset risk growth rate threshold to determine whether the rate of risk change over time exceeds the allowable range, thereby forming the matching result of the risk growth rate threshold condition.
[0075] After determining the risk growth rate threshold, stability analysis and trend matching are performed according to the risk trend consistency condition in the multimodal risk limit state determination conditions. Stability analysis is achieved by comparing the time series of risk trend indicators over multiple consecutive time periods. Specifically, the stability of the risk trend over time is determined by statistically analyzing whether the direction of the risk trend indicators remains consistent within consecutive time windows, and whether the risk level or latent risk state exhibits monotonic changes or small fluctuations within the time window. When the risk trend indicators maintain the same direction of change within a preset number of consecutive time windows, and the risk level does not exhibit a reverse jump, the risk trend is deemed to meet the consistency requirement, thus forming the matching result for the risk trend consistency condition.
[0076] After obtaining the matching results for the risk level threshold condition, risk growth rate threshold condition, and risk trend consistency condition, the matching results are comprehensively judged. When the risk evolution state quantity simultaneously meets the multimodal risk limit state determination condition, the review decision output module 15 outputs the corresponding review restriction result, which is used to restrict the publication or dissemination of multimodal content. When the risk evolution state quantity does not meet the multimodal risk limit state determination condition, the review decision output module 15 outputs a pass or prompt review result. The pass review result indicates that the multimodal content can be published normally, while the prompt review result indicates that the multimodal content needs to enter the prompt or manual review process, thereby completing the matching of the multimodal risk limit state determination condition based on the risk evolution state quantity and the output of the review result.
[0077] In summary, the embodiments of this application have at least the following technical effects: This application, upon receiving multimodal content containing text, images, or audio, performs modal-level parsing processing on the multimodal content to extract modal generation feature vectors that characterize the generation tendency of each modality. Based on the modal generation feature vectors, it performs generation intent inversion to construct a joint representation of multimodal generation intent. It then performs cross-modal consistency analysis on the joint representation of multimodal generation intent to identify semantic offset relationships between different modal generation intents and constructs a cross-modal consistency violation feature tensor. Based on the cross-modal consistency violation feature tensor, it constructs a multimodal content risk evolution model and generates a risk evolution state quantity characterizing the risk evolution state of multimodal content according to the rate of change of consistency violation in the cross-modal content over time. Finally, it performs multimodal risk limit state judgment condition matching based on the risk evolution state quantity and outputs the review result. This invention addresses the technical problem of existing technologies' difficulty in identifying the hidden semantic shifts between different modal generation intentions in multimodal content and the potential risks brought about by their evolution over time. By inverting multimodal generation intentions and identifying cross-modal consistency violations, and combining the rate of change of consistency violations to construct a risk evolution model, this invention achieves the technical effect of dynamically perceiving and determining the risk status of multimodal content, thereby improving the accuracy of review.
[0078] Example 2, based on the same inventive concept as the intelligent review system for multimodal content in the foregoing examples, such as... Figure 2 As shown in the embodiments of this application, an intelligent moderation method for multimodal content is provided, the method comprising: After receiving multimodal content containing text, images, or audio, modal-level parsing is performed on the multimodal content to extract modal generation feature vectors that characterize the generation tendency of each modality. Based on these modal generation feature vectors, generation intent inversion is performed to construct a joint representation of multimodal generation intent. Cross-modal consistency analysis is performed on this joint representation to identify semantic offset relationships between different modal generation intents, and a cross-modal consistency violation feature tensor is constructed. Based on this cross-modal consistency violation feature tensor, a multimodal content risk evolution model is constructed, and a risk evolution state quantity characterizing the risk evolution state of the multimodal content is generated according to the rate of change of consistency violation in the cross-modal content over time. Multimodal risk limit state judgment condition matching is performed based on the risk evolution state quantity, and the review result is output.
[0079] Furthermore, the method also includes: A sliding window is constructed based on the cross-modal consistency violation feature tensor in a time series manner to capture the dynamic changes of multimodal content consistency violation over a continuous time period. Within the sliding window, the time derivative of the cross-modal consistency violation tensor is calculated to obtain the rate of change of cross-modal consistency violation, which is used to quantify the dynamic trend of risk evolution. The rate of change of cross-modal consistency violation is coupled with historical risk evolution state variables to construct a multimodal content risk evolution model. This model is used to predict the evolution path of multimodal content risk in future time periods based on the current rate of change of cross-modal consistency violation. The model outputs risk evolution state variables characterizing the evolution state of multimodal content risk, including risk level, risk growth rate, and risk trend indicators.
[0080] Furthermore, the method also includes: The cross-modal consistency violation feature tensor within the sliding window is temporally encoded to generate initial values for risk hidden states. This temporal encoding includes temporal position encoding of the cross-modal consistency violation feature tensor and cross-modal channel attention convergence to obtain a window representation vector characterizing the consistency violation morphology. The initial values for risk hidden states are then mapped from this window representation vector. The cross-modal consistency violation feature tensor is used as a gating input to recursively update the risk hidden states, forming a risk evolution trajectory. A sequence of risk evolution state quantities is output based on the updated risk evolution trajectory. Calibration sample pairs are constructed based on the audit results and backflow annotations. Confidence constraints are then applied to the gating vector of the rate-of-change gating state update subnetwork to suppress abnormal transitions in risk hidden states when there are noise fluctuations in the cross-modal consistency violation rate.
[0081] Furthermore, the method also includes: The iterative update satisfies: ;in, The recursive update results characterizing the hidden risk states. The initial value is the hidden state of risk. This is a gated vector generated from the rate of change of cross-modal consistency violation, used to adaptively adjust the magnitude of risk state updates. This is the risk increment generation function, used to characterize the nonlinear projection increment of consistency violation in the risk state space.
[0082] Furthermore, the method also includes: The multimodal content risk evolution model is trained using a joint loss function, which includes a risk state quantity prediction error term, a risk growth rate consistency constraint term, and a risk trend direction consistency constraint term.
[0083] Furthermore, the method also includes: The modality generation feature vector within the feature extraction module is a high-level semantic feature vector used to characterize the generation tendency of the corresponding modality content. The generation tendency includes behavioral induction tendency, emotional arousal tendency, and stance orientation tendency. Specifically: the modality generation feature vector of the text modality includes discourse stance distribution features and action instruction display features; the modality generation feature vector of the image modality includes visual composition guidance features and emotional arousal region distribution features; and the modality generation feature vector of the audio modality includes emotional intensity envelope features and rhythm induction features.
[0084] Furthermore, the method also includes: In the cross-modal consistency violation identification module, the cross-modal consistency violation feature tensor is constructed by modeling the directional offset of different modal generation intentions in a unified semantic coordinate space. The directional offset includes differences in intention pointing direction, differences in intention intensity gradient, and differences in intention evolution trend direction.
[0085] Furthermore, the method also includes: The risk evolution state quantity is matched and judged with the preset multimodal risk limit state judgment conditions, wherein the multimodal risk limit state judgment conditions include risk level threshold conditions, risk growth rate threshold conditions, and risk trend consistency conditions; when the risk evolution state quantity meets the multimodal risk limit state judgment conditions, the corresponding review restriction result is output; when the risk evolution state quantity does not meet the multimodal risk limit state judgment conditions, the review result of approval or prompt is output.
[0086] Furthermore, the method also includes: The received multimodal content, including text, images, or audio, is standardized, cleaned, and formatted, and then uniformly encapsulated into a structured JSON format.
[0087] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0088] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0089] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. An intelligent content moderation system for multimodal content, characterized in that: The system includes: The feature extraction module is used to perform modal-level parsing processing on the multimodal content containing text, images or audio after receiving the multimodal content, and extract modal generation feature vectors that respectively characterize the generation tendency of each modal content. The intent inversion module is used to invert the generated intent based on the modality-generated feature vector and construct a joint representation of the multimodal generated intent. The cross-modal consistency violation identification module is used to perform cross-modal consistency analysis on the joint representation of the multimodal generation intent, identify the semantic offset relationship between different modal generation intents, and construct a cross-modal consistency violation feature tensor; The risk evolution module is used to construct a multimodal content risk evolution model based on the cross-modal consistency violation feature tensor, and generate a risk evolution state quantity that characterizes the risk evolution state of multimodal content according to the rate of change of consistency violation of cross-modal content in the time dimension. The audit decision output module is used to match the multimodal risk limit state judgment conditions based on the risk evolution state quantity and output the audit result.
2. The intelligent content review system for multimodal content as described in claim 1, characterized in that, The risk evolution module includes: A sliding window construction unit is used to construct a sliding window based on the cross-modal consistency violation feature tensor according to the time series, so as to capture the dynamic changes of multimodal content consistency violation in a continuous time period; The rate of change calculation unit is used to calculate the time derivative of the cross-modal consistency violation tensor within the sliding window to obtain the rate of change of cross-modal consistency violation, which is used to quantify the dynamic trend of risk evolution. A multimodal risk evolution model unit is used to couple the cross-modal consistency destruction rate with historical risk evolution state quantities to construct a multimodal content risk evolution model. The multimodal content risk evolution model is used to predict the evolution path of multimodal content risk in the future time period based on the current cross-modal consistency destruction rate. The risk evolution state quantity generation unit is used to output risk evolution state quantities that characterize the risk evolution state of multimodal content based on the multimodal content risk evolution model. The risk evolution state quantities include risk level, risk growth rate and risk trend indicators.
3. The intelligent content review system for multimodal content as described in claim 2, characterized in that, The multimodal content risk evolution model is a risk state-space model driven by the rate of change, including: A state initialization subnetwork is used to temporally encode the cross-modal consistency violation feature tensor within the sliding window to generate initial values of the risk hidden state. The temporal encoding includes temporal position encoding of the cross-modal consistency violation feature tensor and cross-modal channel attention convergence to obtain a window representation vector representing the consistency violation morphology, and the initial value of the risk hidden state is obtained by mapping the window representation vector. A rate-of-change gated state update subnetwork is used to recursively update the risk hidden state by taking the cross-modal consistency violation feature tensor as a gated input to form a risk evolution trajectory. A multi-step risk prediction subnetwork is used to output a sequence of risk evolution state variables based on the updated risk evolution trajectory; An online calibration and drift suppression subnetwork is used to construct calibration sample pairs based on the audit results and to perform confidence constraint updates on the gate vector of the rate of change gating state update subnetwork, so as to suppress abnormal transitions of risky hidden states when there are noise fluctuations in the rate of change that violates cross-modal consistency.
4. The intelligent content review system for multimodal content as described in claim 3, characterized in that, The iterative update satisfies: ; in, The recursive update results characterizing the hidden risk states. The initial value is the hidden state of risk. This is a gated vector generated from the rate of change of cross-modal consistency violation, used to adaptively adjust the magnitude of risk state updates. This is the risk increment generation function, used to characterize the nonlinear projection increment of consistency violation in the risk state space.
5. The intelligent content review system for multimodal content as described in claim 3, characterized in that, The multimodal content risk evolution model is trained using a joint loss function, which includes a risk state quantity prediction error term, a risk growth rate consistency constraint term, and a risk trend direction consistency constraint term.
6. The intelligent content review system for multimodal content as described in claim 1, characterized in that, The modality generation feature vector within the feature extraction module is a high-level semantic feature vector used to characterize the content generation tendency of the corresponding modality. This generation tendency includes behavioral induction tendency, emotional arousal tendency, and stance orientation tendency, among which: The modality generation feature vector of text modality includes discourse stance distribution features and action instruction explicitness features; The modality generation feature vector of image modality includes visual composition guidance features and emotional arousal region distribution features; The modality generation feature vector of audio modality includes emotion intensity envelope features and rhythm induction features.
7. The intelligent content review system for multimodal content as described in claim 1, characterized in that, In the cross-modal consistency violation identification module, the cross-modal consistency violation feature tensor is constructed by modeling the directional offset of different modal generation intentions in a unified semantic coordinate space. The directional offset includes differences in intention pointing direction, differences in intention intensity gradient, and differences in intention evolution trend direction.
8. The intelligent content review system for multimodal content as described in claim 1, characterized in that, In the audit decision output module, multimodal risk limit state determination condition matching is performed based on the risk evolution state quantity, including: The risk evolution state quantity is matched and judged with the preset multimodal risk limit state judgment conditions, wherein the multimodal risk limit state judgment conditions include risk level threshold conditions, risk growth rate threshold conditions, and risk trend consistency conditions. When the risk evolution state quantity meets the multimodal risk limit state determination condition, the corresponding audit restriction result is output; when the risk evolution state quantity does not meet the multimodal risk limit state determination condition, the audit result is output as either pass or prompt.
9. The intelligent content review system for multimodal content as described in claim 1, characterized in that, The system also includes: The preprocessing module is used to standardize and clean the received multimodal content, which includes text, images, or audio, and unify its format, and encapsulate it into a structured JSON format.
10. An intelligent content moderation method for multimodal content, characterized in that: The method is executed by any one of the intelligent content review systems according to claims 1 to 9, including: After receiving multimodal content containing text, images, or audio, modal-level parsing processing is performed on the multimodal content to extract modal generation feature vectors that characterize the generation tendency of each modal content. Based on the modal generation feature vectors, the generation intent is inverted to construct a joint representation of multimodal generation intent; Perform cross-modal consistency analysis on the joint representation of the multimodal generated intentions to identify the semantic offset relationship between different modal generated intentions and construct a cross-modal consistency violation feature tensor; Based on the cross-modal consistency violation feature tensor, a multimodal content risk evolution model is constructed, and a risk evolution state quantity representing the risk evolution state of multimodal content is generated according to the rate of change of consistency violation of cross-modal content in the time dimension. Based on the risk evolution state quantity, perform multimodal risk limit state determination condition matching and output the audit result.