An AI-based real-time message multi-modal content auditing system
By using real-time multimodal caching and a comprehensive scoring mechanism, relevant evidence is dynamically selected for message review, which solves the problems of missed and false judgments in real-time video scenarios and achieves accurate and stable review in high-concurrency scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI XIANGYUE JIANGFENG DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2026-06-05
- Publication Date
- 2026-07-03
Smart Images

Figure CN122333191A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence content moderation technology, specifically to an AI-based real-time multimodal content moderation system for comments. Background Technology
[0002] With the development of live streaming, short videos, and social media feeds, user comments have become an important part of content interaction. Current comment moderation typically relies on text keyword matching, single-modal classification models, or video frame extraction within a fixed time window, which can handle some content with clearly prohibited words or obvious image features. However, in real-time video scenarios, the meaning of comments is often closely related to the currently playing screen, the text on the screen, the broadcaster's voice, or other users' historical comments. For example, when a user uses referential expressions such as "look at the bottom right corner" or "what he just said," it is difficult to determine their true meaning from the comment text alone; a comprehensive analysis combining the multimodal context at the corresponding moment is necessary.
[0003] When processing such messages, related technologies using fixed frame rates or fixed time windows are prone to missing fleeting video content, changes in screen text, or audio clips, leading to incomplete review evidence. Alternatively, directly inputting a large number of irrelevant video frames, transcribed text, and historical messages into the model increases inference overhead and may introduce noise, affecting the stability of the review results. Furthermore, while some review models can output judgments, they lack the ability to locate and verify the sufficiency of specific evidence, making them prone to misjudgments or omissions when evidence is insufficient or multimodal content is asynchronous. Summary of the Invention
[0004] This application provides an AI-based real-time multimodal content review system for comments, which at least solves some of the technical problems existing in the related technologies described above.
[0005] According to a first aspect of the embodiments of this application, an AI-based real-time multimodal content review system for comments is provided, comprising: The multimodal cache generation module is used to extract video frames based on video streams and message data, and generate a real-time multimodal cache that includes video frame records, on-screen text records, speech-to-text segments, and historical message records. The comprehensive scoring module is used to convert the messages to be reviewed into message text, encode the messages to be reviewed, and generate context-dependent weight vectors and message semantic vectors. The candidate evidence generation module is used to calculate time proximity score, semantic similarity score and change significance score for cached records in real-time multimodal cache based on the release time, and obtain a comprehensive score by combining the context dependency weight vector. Based on the comprehensive score, candidate evidence records are selected, and adjacent cached records in the time neighborhood are added to the candidate evidence records and deduplicated to generate a candidate evidence set. The review and judgment module is used to organize the message text, posting time, and candidate evidence set into a structured review input containing evidence number, timestamp, and spatial location information. The input is then fed into the visual language model to obtain the review and judgment result, the list of evidence numbers, and the confidence score. The review action determination module is used to calculate the evidence sufficiency score based on the candidate evidence set, the evidence number list, and the confidence level value, and to determine the review action based on the review judgment result and the evidence sufficiency score.
[0006] According to a second aspect of the embodiments of this application, an AI-based real-time multimodal content review method for comments is also provided, including: Based on video stream and message data, video frames are extracted to generate a real-time multimodal cache that includes video frame records, on-screen text records, speech-to-text segments, and historical message records. Convert the messages to be reviewed into message text, encode the messages to be reviewed, and generate context-dependent weight vectors and message semantic vectors; Based on the release time, the cache records of the real-time multimodal cache are calculated by type for time proximity score, semantic similarity score and change significance score. Combined with the context dependency weight vector, a comprehensive score is obtained. Candidate evidence records are selected based on the comprehensive score. Adjacent cache records in the time neighborhood are added to the candidate evidence records and duplicates are removed to generate a candidate evidence set. The message text, posting time, and candidate evidence set are organized into a structured review input containing evidence number, timestamp, and spatial location information. This input is then fed into a visual language model to obtain the review judgment result, the list of evidence numbers, and the confidence score. The sufficiency score is calculated based on the candidate evidence set, the evidence number list, and the confidence level value. The review action is determined based on the review judgment result and the sufficiency score.
[0007] This application establishes a real-time multimodal cache including video frames, on-screen text, speech-to-text fragments, and historical messages. It predicts the dependence of messages on different evidence sources based on their semantic features, enabling the system to dynamically select relevant evidence based on the actual meaning of the message, rather than mechanically extracting content from a fixed time window. This improves the recall rate of key evidence while controlling the scale of review data, reducing the risk of missed judgments due to short-term changes in visuals, flashing text, or seamless transitions in speech content. Furthermore, by supplementing candidate evidence with temporal neighborhood records and organizing them into structured review inputs with evidence numbers, timestamps, and spatial location information, the visual language model can accurately understand the correspondence between referential messages and on-site content. Moreover, this invention calculates an evidence sufficiency score based on model confidence, evidence coverage, evidence consistency, and evidence source diversity, and determines review actions such as direct approval, direct rejection, delayed re-review, or manual review accordingly. This reduces the probability of misjudgments due to insufficient evidence, improves the accuracy, stability, and traceability of real-time message review, and is suitable for high-concurrency review scenarios such as live streaming, short videos, and social media updates.
[0008] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Furthermore, no embodiment in this disclosure is required to achieve all the effects described above. Attached Figure Description
[0009] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0010] Figure 1 This is a schematic diagram of an AI-based real-time multimodal content review method for comments, provided as an embodiment of this disclosure.
[0011] Figure 2 A flowchart for generating context-dependent weight vectors and message semantic vectors provided in embodiments of this disclosure.
[0012] Figure 3 A flowchart illustrating the comprehensive scoring calculation process provided in this embodiment of the disclosure.
[0013] Figure 4 A flowchart illustrating the process of generating a set of candidate evidence provided in this embodiment of the disclosure.
[0014] Figure 5 A block diagram of the visual language model structure provided in the embodiments of this disclosure.
[0015] Figure 6 This is a schematic diagram of the structure of an AI-based real-time message multimodal content review system provided in an embodiment of this disclosure.
[0016] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0017] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0018] According to embodiments of this disclosure, the present invention is applicable to real-time content review scenarios for user comments in live streams, short video playback streams, or social media feeds. The review targets are user-posted comments, which can take the form of text, emoticons, stickers, images, or voice messages. In this scenario, the comments themselves often do not contain explicitly prohibited words; their prohibited meaning arises from the combination of the comment with the currently playing screen, on-screen text, audio content, or other user comments. For example, comments like "look at the bottom right corner" or "what he just said" require a combination of the corresponding video footage or speech transcription to determine their meaning. The system is deployed in a server environment with video stream access capabilities, multimodal feature extraction capabilities, and model inference capabilities. It continuously receives video stream data from live streams or short video services and user comment data from comment services, performing review and judgment on each comment under high concurrency conditions.
[0019] The implementation process of the method described in this application will be described in detail below with reference to specific embodiments. It should be noted that this embodiment is only used to explain this application and is not intended to limit the scope of protection of this application. Conventional adjustments or substitutions of each step by those skilled in the art without departing from the concept of this application should be included in the scope of protection of this application.
[0020] Please see Figure 1 , Figure 1 The flowchart illustrates an AI-based real-time multimodal content moderation method for comments, according to an embodiment of the present invention. This method can be executed by an AI-based real-time multimodal content moderation system for comments, such as... Figure 1 As shown, the method includes steps S1-S5: In step S1, based on the video stream and message data, video frames are extracted to generate a real-time multimodal cache containing video frame records, on-screen text records, speech-to-text segments, and historical message records.
[0021] In some embodiments, the system continuously receives video streams and message data, extracts video frames from the video stream at a preset base frame rate, such as 1 to 3 frames per second. The system calculates the sum of the absolute values of the pixel-by-pixel differences in the downsampled grayscale images of adjacent frames, and uses the ratio of this value to the total number of pixels in the frame image as the image change rate. When the image change rate exceeds a preset change threshold, additional video frames are extracted at the time of the change. This threshold is determined according to the business scenario, for example, it can be configured to 5% to 20%. In addition, when the system detects the appearance of new text areas in the image or changes in the content of existing text areas, additional video frames are also triggered. The combination of base frame rate extraction and event-triggered supplementary frame extraction can reduce the possibility of missing short-flash content while controlling the total number of frames.
[0022] In some embodiments, for each extracted video frame, the system extracts the text content, the text's position coordinates in the frame, and the recognition confidence level using an Optical Character Recognition (OCR) model to generate an on-screen text record; simultaneously, it extracts the visual features of the frame using a visual encoder to generate a video frame record; for the audio stream, the system uses an Automatic Speech Recognition (ASR) model for continuous transcription, outputting a timestamped speech transcription segment; and user-posted historical messages are encoded by a text encoder to generate historical message records.
[0023] The aforementioned records collectively constitute a real-time multimodal cache. Specifically, video frame records include a record type identifier, timestamp, image data, and semantic vector, and are associated with the corresponding frame change rate and frame extraction trigger identifier. The frame extraction trigger identifier indicates whether the frame was extracted at the base frame rate, triggered by the frame change rate exceeding a preset change threshold, or triggered by changes in the text area of the image. The semantic vector is determined by the visual features extracted from the video frame by the visual encoder. The image text record includes a record type identifier, timestamp, text content, location coordinates, recognition confidence, and semantic vector. The speech-to-text segment includes a record type identifier, start timestamp, end timestamp, transcribed text content, and semantic vector. The historical message record includes a record type identifier, timestamp, message content, and semantic vector. The real-time multimodal cache is maintained using a sliding window method, with the window length configurable from the most recent few seconds to several minutes. A first-in-first-out (FIFO) strategy is used to delete cached records that exceed the window range.
[0024] Optionally, to avoid incorrect merging of different cached records at the same time point during deduplication, each type of cached record also includes a unique record identifier; among them, the unique record identifier of video frame records is determined by the video frame identifier, the unique record identifier of on-screen text records is determined by the frame identifier and the text region identifier, the unique record identifier of speech-to-text segments is determined by the start timestamp and the end timestamp, and the unique record identifier of historical message records is determined by the message identifier.
[0025] In step S2, the message to be reviewed is converted into message text and encoded, generating a context-dependent weight vector and a message semantic vector.
[0026] Specifically, please refer to Figure 2 , Figure 2 A flowchart illustrating the generation process of context-dependent weight vectors and message semantic vectors provided in embodiments of this disclosure is shown. Figure 2 As shown in box 201, when a message awaiting review arrives, the system first converts it into message text.
[0027] The conversion process is handled according to the content type of the message to be reviewed: for text messages, the text content is extracted directly; for messages containing emoticons, each emoticon is mapped to a corresponding semantic text description, such as mapping a smiley face to the word "happy"; for voice messages, speech-to-text is generated through a speech recognition model; the extracted or generated text content, semantic text description, and speech-to-text are then merged into a complete message text in their original order in the message to be reviewed.
[0028] For comments to be reviewed that contain stickers or images, the visual encoder extracts the image semantic vector of the sticker or image. This image semantic vector is used in the subsequent comment semantic vector generation step. The visual encoder uses a pre-trained visual transformer (ViT) structure to segment the input image into fixed-size patches and then perform linear embedding, add position encoding, process through a multi-layer transformer encoding layer, and take the vector corresponding to the position of the classification word in the final output as the image semantic vector.
[0029] In box 202, the converted message text is encoded to generate a context-dependent weight vector and a message semantic vector. Specifically, the context-dependent weight vector consists of four components: video image dependency weight, speech-to-text dependency weight, image-text dependency weight, and historical message dependency weight. These components represent the degree to which the current message depends on various types of multimodal evidence during the review process. The values of the four components are all between 0 and 1, and their sum is normalized to 1.
[0030] Specifically, the text encoder receives a lexical sequence of the message text after lexicalization. This encoder employs a multi-layer bidirectional Transformer structure, configurable to four to six layers. Each layer contains a multi-head self-attention sublayer and a feedforward network sublayer. The multi-head self-attention sublayer calculates the correlation between each lexical in the input sequence. Each attention head independently performs a scaled dot product attention operation on the three sets of linear mapping results (query, key, and value). The outputs of multiple heads are concatenated and linearly projected to obtain the output of this sublayer. The feedforward network sublayer consists of two fully connected layers, processed by a ReLU activation function in between. The encoder outputs a context representation matrix for the input sequence, taking the vector corresponding to the starting position of the sequence as the global semantic representation.
[0031] A dependency weight prediction head is connected above the text encoder. This prediction head consists of two fully connected layers: the first layer maps the global semantic representation to an intermediate dimension and processes it through the ReLU activation function; the second layer maps the intermediate representation to a four-dimensional output and then normalizes it through the softmax function to obtain the context dependency weight vector. For example, when the message content is "look at the bottom right corner", the dependency weight prediction head tends to assign higher values to the image text dependency weight and the video image dependency weight; when the message content is "what he just said was too outrageous", the speech transcription dependency weight will obtain a higher value. That is, the model learns the statistical association between message text patterns and evidence types from historical data without the need for manual enumeration of specific word rules.
[0032] Optionally, for training the text encoder and dependency weight prediction head, the training data comes from historical review records; for each historical comment, the reviewer or an offline large model labels the main evidence type on which the review conclusion is based, generating a four-dimensional soft label as a supervision signal; for example, when a historical comment is labeled as a dependency When considering a type of evidence, the following should be... The soft label component corresponding to each evidence type is set to... For any unlabeled evidence type, the soft label components are set to 0, and the sum of the four-dimensional soft label components is set to 1. For example, if the review of a comment is mainly based on video footage and text within the footage, then the video footage dependency weight and text dependency weight in the corresponding label are both set to 0. The other two terms are set to 0. During training, cross-entropy is used as the loss function to make the context-dependent weight vector output by the model approximate the labeled values, and mini-batch gradient descent is used to update the parameters.
[0033] The generation method of the message semantic vector depends on whether the message to be reviewed contains an image: when no image semantic vector exists, the global semantic representation output by the text encoder is directly used as the message semantic vector; when an image semantic vector exists, the global semantic representation and the image semantic vector are concatenated along the feature dimension, mapped through a linear projection layer to a vector of the same dimension as the global semantic representation, and this mapping result is used as the message semantic vector. The weights of the linear projection layer are obtained through joint training with the text encoder.
[0034] In step S3, based on the release time, the time proximity score, semantic similarity score, and change significance score are calculated for the cached records of the real-time multimodal cache according to their types, and a comprehensive score is obtained by combining the context dependency weight vector.
[0035] Specifically, please refer to Figure 3 , Figure 3 A flowchart illustrating the comprehensive score calculation process provided in an embodiment of this disclosure is shown. Figure 3 As shown in box 301, the time proximity score reflects the time distance between the cached record and the time the message was posted.
[0036] For video frame recordings, on-screen text recordings, and historical message records, the timestamp of the record is used as the recording time; for speech-to-text segments, the midpoint between the start and end timestamps is used as the representative time. ; Let the posting time of the message awaiting review be The time difference is ,and , With time decay parameter Use the same time unit.
[0037] The time difference between the cached timestamp and the publication time is decayed over time, specifically as follows:
[0038] in For the time difference, This is a time decay parameter that controls the rate at which the score decreases as the time difference increases; The value is determined based on the application scenario; for live streaming scenarios, it can be configured to range from several seconds to tens of seconds. The value ranges from 0 to 1.
[0039] In box 302, the semantic similarity score measures the degree of association between the semantic content of the cached record and the meaning of the message. The cosine similarity is calculated between the semantic vector of the message and the semantic vector of the cached record, and the result is truncated to the 0-1 range to obtain the semantic similarity score. To meet the real-time retrieval latency requirements, the semantic vectors of each record in the cache are pre-constructed as vector indexes. When a message arrives, an approximate nearest neighbor retrieval is performed using the message's semantic vector as the query, quickly narrowing down the candidate range.
[0040] In box 303, the significance score of change is determined based on the recording type. For each video frame recording, the rate of change of that frame is normalized to the 0-1 range within the buffer window to obtain the significance score of change. The normalization method involves dividing the rate of change of the current frame by the maximum rate of change of all frames within the window, ensuring that the frame with the most dramatic change receives the highest score. When the maximum rate of change of all frames within the window is zero, the change significance score of the video frame is set to zero. For speech-to-text segments, the cosine similarity between their semantic vectors and the semantic vectors of the previous speech-to-text segment is calculated. If there is no previous speech-to-text segment, the change significance score of that segment is set to zero. If there is a previous speech-to-text segment and the similarity is below a preset threshold, the change significance score is set to zero. The value should be set to 1; otherwise, the significance of the change will be scored. The value is set to 0; this preset threshold can be configured between 0.3 and 0.6.
[0041] For on-screen text recordings, the significance of change is scored based on whether the text area is newly introduced or whether the content has changed. The scores for newly introduced text areas and text areas with changed content are further divided into two categories. The score for the significance of change is determined to be 1, indicating that the content has not changed. The value is determined to be 0.
[0042] In box 304, a comprehensive score is calculated based on the type of cached record. For video frame records, speech-to-text segments, and on-screen text records, the comprehensive score is a weighted combination of the dependency weight corresponding to that type multiplied by the temporal proximity score, semantic similarity score, and change significance score:
[0043] in For the component of this type in the context-dependent weight vector, , , The combined weighting coefficients for the rating dimensions satisfy... The contribution ratios of temporal proximity, semantic similarity, and significance of change are controlled respectively. These coefficients can be calculated from historical review data. For each historical comment and its manually annotated key evidence fragments, the average score distribution of key evidence across the three scoring dimensions is used as the initial value. Then, a grid search is used to select the combination that maximizes the recall rate of candidate evidence on the validation set. All three coefficients range from 0 to 1.
[0044] For historical message records, since no changes in the image are involved, the overall score is obtained by combining the time proximity score and the semantic similarity score according to the corresponding normalized ratio.
[0045] In step S4, candidate evidence records are selected based on the comprehensive score, and adjacent cache records in the time neighborhood are added to the candidate evidence records and deduplicated to generate a candidate evidence set.
[0046] Please see Figure 4 , Figure 4 A flowchart illustrating the candidate evidence set generation process provided in an embodiment of this disclosure is shown. Figure 4 As shown in box 401, video frame records, on-screen text records, speech-to-text segments, and historical message records are sorted in descending order of comprehensive score to select candidate evidence records.
[0047] The number of selections for each type is determined jointly by the context-dependent weight vector and the upper limit of the total number of candidate evidences: let the upper limit of the total number be... ,in The same parameter applies to the upper limit of the number of pieces of evidence mentioned above; the number of pieces selected for each type is the dependency weight of that type multiplied by [the specified value]. Then round down; if the sum of the rounded quantities of each type is less than Then multiply by the dependency weights of each type. The decimal parts are added in descending order to fill the remaining quotas, until a total of [number] places are reached. Or there are no available cache records for the corresponding type.
[0048] When the dependency weight of a certain type is lower than the minimum dependency threshold, the selection quantity of that type is determined to be zero, meaning that type does not participate in the candidate evidence. The minimum dependency threshold is configurable, for example, between 0.05 and 0.1. The system retrieves candidate evidence records from the sorting results according to the selection quantity of each type; cached records that are not retrieved are not included in the candidate evidence record. This allows the system to automatically adjust the type composition of evidence based on the semantic features of the message content. For messages clearly pointing to visual content, video frame evidence is added; for messages pointing to audio content, speech-to-text evidence is added, avoiding the use of the same fixed window capture strategy for all messages.
[0049] In box 402, for the selected candidate evidence records, adjacent cached records in their temporal neighborhood are supplemented to address the situation where multimodal content is not synchronized in time. Specifically, for candidate video frame records, adjacent video frame records within a preset time range before and after their timestamp are obtained. This range can be configured to be 1 to 3 seconds. When a candidate video frame record corresponds to a scene change event, that is, when the frame extraction trigger flag of the candidate video frame record indicates that it was triggered by the scene change rate exceeding a preset change threshold, the nearest video frame record before the scene change event is added to allow the subsequent model to compare the scene differences before and after the change.
[0050] For candidate speech-to-text segments, one or two adjacent speech-to-text segments before and after them are acquired as extensions. In some embodiments, when the text length of a candidate speech-to-text segment is short, the extension range can be appropriately increased to ensure the semantic integrity of the speech context. For candidate on-screen text records, the video frame record of the frame in which it is located is acquired, and on-screen text records of the same text region in adjacent frames are also acquired, so that subsequent models can determine whether the text is continuously displayed or flashes briefly.
[0051] In box 403, the calculated comprehensive score of the adjacent cached records obtained in addition is retained, and these records are merged with the candidate evidence records. Deduplication is performed based on the record type identifier, timestamp, and unique record identifier. When the number of merged and deduplicated records exceeds the preset upper limit for the number of evidence records, records within the upper limit are retained in descending order of comprehensive score, and the remaining records are discarded. The retained records are defined as the candidate evidence set, and the discarded records are not added to the candidate evidence set. This upper limit for the number of evidence records is configurable, for example, it can be configured to 20-40 records.
[0052] In step S5, the message text, posting time, and candidate evidence set are organized into a structured review input containing evidence number, timestamp, and spatial location information. This input is then fed into the visual language model to obtain the review judgment result, the list of evidence numbers, and the confidence score.
[0053] The structured review input includes the current message field, candidate video frame field, candidate on-screen text field, candidate speech-to-text field, and historical message summary field. The current message field records the message text and posting time; video frame records, on-screen text records, speech-to-text segments, and historical message records in the candidate evidence set are written into their corresponding fields according to their evidence numbers, which are sequential integers starting from the first; in the candidate video frame field, video frames are arranged in ascending order of timestamps, and each record carries image data, a timestamp, and an evidence number; the candidate on-screen text field records the text content, location coordinates, the timestamp of the frame, and the evidence number; the candidate speech-to-text field records the transcribed text, start and end timestamps, and the evidence number; historical message records are concatenated in chronological order to form the historical message summary field, and each record is labeled with its evidence number.
[0054] The system organizes textual information into prompt text and uses video frame images as image input. The prompt text lists the type, timestamp, content summary, and spatial location description of each piece of evidence in order of evidence number. For textual evidence in the image, its location area in the image is indicated, and for speech-to-text evidence, the start and end times are indicated. The video frame images are marked with placeholders to indicate their insertion position in the prompt text. The evidence number is input into the visual language model along with the corresponding evidence item, so that the model can reference the specific evidence item in the output.
[0055] In some embodiments, specifically, such as Figure 5 As shown, Figure 5 A block diagram illustrating the visual language model structure provided in an embodiment of this disclosure is shown. Figure 5 As shown, the visual language model adopts a three-module structure. The visual encoder module receives images of each candidate video frame and uses a visual transformer structure to divide each frame image into fixed-size patches and then perform linear embedding, adding two-dimensional positional encoding. After processing by a multi-layer transformer encoding layer, a set of visual word sequences is output. Each encoding layer contains a multi-head self-attention sub-layer and a feedforward network sub-layer, and the processing method is similar to that of the text transformer encoding layer.
[0056] The visual projection module maps visual lexical units from the feature space of the visual encoder to the lexical embedding space of the language model. This module uses a multilayer perceptron consisting of two fully connected layers. The first layer maps the visual lexical units to the intermediate dimension and processes them through the GELU activation function. The second layer maps the intermediate representation to the lexical embedding dimension of the language model.
[0057] The language model module receives a mixed sequence consisting of alternating textual and mapped visual lexical units. It employs a multi-layer causal transformer structure, with each layer containing a multi-head self-attention sublayer with causal masks and a feedforward network sublayer. It generates output text word-by-word through an autoregressive approach. In some embodiments, the language model module may also employ a hybrid expert architecture, where its feedforward network sublayer consists of multiple parallel expert networks. During inference, a gating network selects and activates a small number of experts based on the input, thereby controlling the computational load of a single inference while maintaining a large model capacity. The model output includes the review decision (pass, reject, or requiring manual review), a list of evidence numbers on which the decision is based, and a confidence score.
[0058] Through the above-mentioned structured input organization method, each evidence item retains timestamp and spatial location information, enabling the model to establish a correspondence between referential expressions such as "bottom right corner" and "the sentence just now" in the message and specific evidence items, thereby completing the judgment based on clear evidence.
[0059] In step S6, the evidence sufficiency score is calculated based on the candidate evidence set, the evidence number list, and the confidence level value. The review action is then determined based on the review judgment result and the evidence sufficiency score.
[0060] In some embodiments, the sufficiency score is obtained by a weighted combination of four indicators: the first indicator is the confidence index, which is the confidence value output by the model and ranges from 0 to 1; the second indicator is the evidence coverage, which is the number of evidences in the evidence number list divided by the total number of evidences in the candidate evidence set; when the total number of evidences in the candidate evidence set is zero, the evidence coverage is determined to be zero.
[0061] The third metric is evidence consistency. The system inputs each piece of evidence corresponding to the evidence number list into a lightweight classifier to obtain the independent review tendency of each piece of evidence. This lightweight classifier uses a single-layer fully connected layer structure. The input is the semantic vector of the evidence, and the output, after softmax normalization, is the probability distribution of each review category. The category with the highest probability is taken as the review tendency of that evidence. The classifier is trained on historical review data, using the semantic vector of each piece of evidence as input and manually annotated review conclusions as labels, and employs the cross-entropy loss function for training. The evidence consistency is calculated as follows: the number of pieces of evidence whose review tendency is consistent with the model's final review judgment result is divided by the number of pieces of evidence in the evidence number list; when the evidence number list is empty, the evidence consistency is set to zero.
[0062] The fourth indicator is the diversity of evidence sources, which is calculated by dividing the number of distinct record types corresponding to the evidence number list by the total number of distinct record types in the candidate evidence set; when the total number of record types in the candidate evidence set is zero, the diversity of evidence sources is determined to be zero.
[0063] The system calculates the weighted sum of the four indicators according to a preset combination of weights to obtain the sufficiency of evidence score. The preset combined weights include four coefficients, which correspond to the confidence index, evidence coverage, evidence consistency, and evidence source diversity, respectively. The values of the four coefficients are all between 0 and 1 and their sum is 1. The combined weights are determined by grid search on the validation set, and the weight combination that maximizes the audit accuracy is selected.
[0064] In some embodiments, the review action is based on the review decision and the sufficiency of evidence score. The following rules apply: when the review result is passed and... If the message exceeds the sufficiency threshold, the review action is determined as direct approval, meaning the message can be displayed normally; if the review result is rejection and... When the threshold for rejection sufficiency is exceeded, the review action is determined as direct rejection, meaning the system automatically blocks the message. The rejection sufficiency threshold is higher than the approval sufficiency threshold to ensure that the rejection action corresponds to more substantial evidence, avoiding misjudgment as a violation when the evidence is insufficient. The approval sufficiency threshold can be configured, for example, to a range of 0.5-0.7, and the rejection sufficiency threshold can be configured, for example, to a range of 0.7-0.9, both determined by adjustment on the validation set. When the review decision indicates that manual review is required, the review action is directly determined as submitting for manual review; in this case, no further consideration is needed. The value of .
[0065] when If the evidence is not greater than the corresponding sufficiency threshold, the system enters the insufficient evidence processing branch. The system checks whether there are any new cached records that have not yet been included in the candidate evidence set after the posting time. That is, in a live streaming scenario, the video stream continues after the comment is posted, and new video frames and audio clips will continuously enter the cache. These newly arrived data may contain key evidence related to the comment. If the new cached record exists and the waiting time does not exceed the maximum waiting time, the review action is determined to be delayed review. The system re-executes the aforementioned steps after a preset waiting interval to include the newly arrived cached record in the search scope. If the new cached record does not exist or the waiting time has exceeded the maximum waiting time, the review action is determined to be submitted for manual review. The maximum waiting time is configurable, for example, it can be configured to be 1 to 3 seconds to ensure that delayed review does not excessively affect the user experience.
[0066] In some embodiments, the selection of candidate evidence records, based on the calculated comprehensive score, can also employ an end-to-end trained candidate evidence selection model to assist in re-ranking candidate cached records. The output of this candidate evidence selection model is used to assist in determining the ranking of cached records of the same type, without replacing the comprehensive score calculation or changing the process of selecting candidate evidence records based on the comprehensive score. For example, the model can adopt a dual-tower structure: the query tower receives the concatenation result of the message semantic vector and the context-dependent weight vector, which is mapped to a query representation vector through two fully connected layers, with the two layers connected by a ReLU activation function; the candidate tower receives the concatenation result of the semantic vector of the cached record, the sinusoidal position encoding vector of the time difference, and the learnable record type embedding vector, which is mapped to a candidate representation vector through two fully connected layers with the same structure. The model calculates the inner product of the query representation vector and each candidate representation vector, and maps the inner product result corresponding to each cached record using a sigmoid function to obtain the selection probability of the cached record being selected as key evidence; during training, manually annotated key evidence fragments are used as positive samples, and the remaining records are used as negative samples, trained using a binary cross-entropy loss function.
[0067] Optionally, the visual language model can be fine-tuned through knowledge distillation to ensure that the judgment capability of the smaller-scale online model, receiving only the candidate evidence set as input, is close to that of the large-scale offline model with full cached record input. Specifically, during distillation, the large-scale offline visual language model generates teacher signals, including the review judgment results and the output probability distributions for each category, using the full cached records as input; the smaller online model generates student outputs using the candidate evidence set as input. The loss function of the student model consists of two parts: one is the cross-entropy loss between the student output and the manually labeled tags, and the other is the KL divergence between the probability distribution of the student output and the probability distribution of the teacher signal. The total loss is:
[0068] in For cross-entropy loss, For KL divergence loss, This is the proportionality coefficient between cross-entropy loss and KL divergence loss. It can be configured to be 1 to 3, corresponding to a weight ratio of 1:1 to 3:1 for cross-entropy loss and KL divergence loss.
[0069] For messages that are determined to be submitted for manual review, the system provides the reviewers with the message content, the original content of each piece of evidence in the candidate evidence set, as well as its time and location on the screen, the evidence number of the review judgment result and basis output by the model, and the evidence sufficiency score and the values of each sub-indicator; the reviewers can then directly locate the relevant video clips and screen areas for review.
[0070] Therefore, this invention selects the most relevant evidence fragments from a real-time multimodal cache based on the semantics of the message. The selected evidence is organized into structured input according to time, location, and source and fed into a visual language model for judgment. The judgment result is then verified through an evidence sufficiency assessment. Compared to fixed time windows and fixed frame rate sampling, this invention exhibits higher recognition stability for short-lived flashing content and referential messages. By dynamically adjusting the selection quantity of various types of evidence according to the message dependency type, it reduces the interference of irrelevant evidence on model reasoning and computational overhead. The evidence sufficiency assessment mechanism introduces delayed re-examination or manual review when evidence is insufficient, reducing the risk of misjudgment due to missing evidence, making it suitable for deployment requirements in high-concurrency real-time review scenarios.
[0071] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of an AI-based real-time multimodal content review system for user comments, provided in an embodiment of this application. As shown in the figure, the system includes: The multimodal cache generation module 601 is used to extract video frames based on video stream and message data, and generate a real-time multimodal cache that includes video frame records, on-screen text records, speech-to-text segments and historical message records. The comprehensive scoring module 602 is used to convert the messages to be reviewed into message text, encode the messages to be reviewed, and generate context-dependent weight vectors and message semantic vectors. The candidate evidence generation module 603 is used to calculate the time proximity score, semantic similarity score and change significance score of the cached records of the real-time multimodal cache according to the type based on the release time, and obtain a comprehensive score by combining the context dependency weight vector. Based on the comprehensive score, candidate evidence records are selected, and adjacent cached records in the time neighborhood are added to the candidate evidence records and deduplicated to generate a candidate evidence set. The review and judgment module 604 is used to organize the message text, posting time and candidate evidence set into a structured review input containing evidence number, timestamp and spatial location information. The input is then fed into the visual language model to obtain the review and judgment result, the list of evidence numbers and confidence scores. The review action determination module 605 is used to calculate the evidence sufficiency score based on the candidate evidence set, the evidence number list, and the confidence level value, and to determine the review action based on the review judgment result and the evidence sufficiency score.
[0072] Optionally, the multimodal cache generation module is used to extract video frames according to a preset base frame rate; calculate the ratio of the sum of the absolute values of the pixel-by-pixel differences of the downsampled grayscale images of adjacent frames to the total number of pixels in the frame image to obtain the screen change rate; when the screen change rate exceeds a preset change threshold, supplement the extraction of video frames at the time point of the change; when a new screen text region or a change in the content of an existing screen text region is detected, supplement the extraction of video frames; generate video frame records and screen text records for the extracted video frames, and generate speech-to-text segments for the audio stream.
[0073] Optionally, the video frame record includes a record type identifier, timestamp, image data, and semantic vector, wherein the semantic vector is determined by video frame features; the on-screen text record includes a record type identifier, timestamp, text content, location coordinates, recognition confidence, and semantic vector; the speech-to-text segment includes a record type identifier, start timestamp, end timestamp, transcribed text content, and semantic vector; the historical message record includes a record type identifier, timestamp, message content, and semantic vector; and the real-time multimodal cache deletes cached records that exceed the window range using a sliding window and first-in-first-out strategy.
[0074] Optionally, it also includes a message encoding module for extracting text content from text messages; mapping emojis to semantic text descriptions for messages containing emojis; generating speech-to-text transcription for voice messages; merging text content, semantic text descriptions, and speech-to-text transcription into message text in the order they appear in the message to be reviewed; and extracting image semantic vectors from images or stickers for messages to be reviewed.
[0075] Optionally, the context-dependent weight vector includes video image dependency weights, speech-to-text dependency weights, image-text dependency weights, and historical message dependency weights; a text encoder outputs a global semantic representation of the message text; a dependency weight prediction head maps the global semantic representation to the context-dependent weight vector; wherein, when there is no image semantic vector, the global semantic representation is determined as the message semantic vector; when there is an image semantic vector, the global semantic representation and the image semantic vector are concatenated and mapped to the message semantic vector through a linear projection layer.
[0076] Optionally, the comprehensive scoring module is used to obtain a time proximity score by performing time decay processing on the time difference between the timestamp of the cached record or the representative time determined by the start and end timestamps of the speech-to-text segment and the publication time; to obtain a semantic similarity score by calculating the cosine similarity between the semantic vector of the message and the semantic vector of the cached record and truncating it to the 0-1 interval; to obtain a change significance score by normalizing the video frame record according to the rate of change of the image; to determine the change significance score of the speech-to-text segment according to the cosine similarity between its semantic vector and the semantic vector of the previous speech-to-text segment; and to determine the change significance score of the on-screen text record according to its new appearance or content change; the comprehensive score of historical message records consists of a time proximity score and a semantic similarity score.
[0077] Optionally, the candidate evidence generation module is used to sort the comprehensive scores in descending order according to video frame records, on-screen text records, speech-to-text segments, and historical message records; determine the selection quantity of each type according to the context dependency weight vector and the upper limit of the total number of candidate evidences; when the dependency weight of a certain type is lower than the minimum dependency threshold, the selection quantity of that type is set to zero, and candidate evidence records are obtained from the sorting results according to the selection quantity of each type.
[0078] Optionally, the candidate evidence generation module is further configured to: record candidate video frames and obtain adjacent video frame records within a preset time range before and after their timestamps; when a candidate video frame record corresponds to a screen change event, add the most recent video frame record before the screen change event occurs; for candidate speech-to-text segments, obtain speech-to-text segments that are adjacent in time before and after; for candidate screen text records, obtain the video frame record of the frame in which it is located and the screen text records of the same text area in adjacent frames; retain the calculated comprehensive score for the supplemented adjacent cache records, merge them with the candidate evidence records, and remove duplicates according to the record type identifier and timestamp to generate a candidate evidence set.
[0079] Optionally, the review action determination module is used to: use the confidence level value as the confidence level indicator; divide the number of evidences in the evidence number list by the total number of evidences in the candidate evidence set to obtain the evidence coverage; input each piece of evidence corresponding to the evidence number list into a lightweight classifier to obtain the review tendency of each piece of evidence, and divide the number of evidences whose review tendency is consistent with the review judgment result by the number of evidences in the evidence number list to obtain the evidence consistency; divide the number of record types corresponding to the evidence number list by the total number of record types in the candidate evidence set to obtain the evidence source diversity; and weight and sum the confidence level indicator, evidence coverage, evidence consistency, and evidence source diversity according to a preset combination weight to obtain the evidence sufficiency score.
[0080] Each processing unit and / or module in the embodiments of this application can be implemented by an analog circuit that implements the functions described in the embodiments of this application, or by software that executes the functions described in the embodiments of this application.
[0081] Please see Figure 7 It shows a schematic diagram of the structure of an electronic device according to an embodiment of this application, which can be used to implement... Figure 1 The method in the illustrated embodiment. (As shown) Figure 7 As shown, the electronic device may include: The system includes at least one processor 701, at least one network interface 704, a user interface 703, a memory 705, and at least one communication bus 702. The communication bus 702 is used to enable connection and communication between the components. The user interface 703 may include buttons, and optionally include a standard wired or wireless interface. The network interface 704 may include, but is not limited to, a Bluetooth module, an NFC module, a Wi-Fi module, etc.
[0082] The processor 701 may include one or more processing cores and connect to various parts within the electronic device via various interfaces and lines. It implements various functions and data processing of the electronic device by running or executing instructions, programs, code sets, or instruction sets stored in the memory 705, and by accessing data in the memory 705. Optionally, the processor 701 may be implemented using at least one hardware form of DSP, FPGA, or PLA. The processor 701 may also integrate one or more combinations of CPU, GPU, and modem.
[0083] The memory 705 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 705 includes a non-transitory computer-readable medium for storing instructions, programs, code, code sets, or instruction sets. The memory 705 may be divided into a program storage area and a data storage area, wherein the program storage area can be used to store instructions for implementing an operating system and instructions for implementing the foregoing method embodiments; the data storage area can be used to store data related to the relevant method embodiments. The memory 705 may also be at least one storage device located remotely from the processor 701. Figure 7 As shown, the memory 705, which serves as a computer storage medium, may contain an operating system, a network communication module, a user interface module, and program instructions.
[0084] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by processor 701, the functions defined in the methods of this application are performed.
[0085] Another embodiment of this application provides a storage medium storing computer program instructions thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application.
[0086] In the above embodiments, the descriptions of each embodiment have different focuses. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. The above descriptions are merely preferred embodiments of this application and explanations of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by specific combinations of the above technical features, but should also cover other technical solutions formed by arbitrary combinations of the above technical features or their equivalent features without departing from the inventive concept.
Claims
1. An AI-based real-time voice message multi-modal content review system, characterized in that, include: The multimodal cache generation module is used to extract video frames based on video streams and message data, and generate a real-time multimodal cache that includes video frame records, on-screen text records, speech-to-text segments, and historical message records. The comprehensive scoring module is used to convert the messages to be reviewed into message text, encode the messages to be reviewed, and generate context-dependent weight vectors and message semantic vectors. The candidate evidence generation module is used to calculate time proximity score, semantic similarity score and change significance score for cached records in real-time multimodal cache based on the release time, and obtain a comprehensive score by combining the context dependency weight vector. Based on the comprehensive score, candidate evidence records are selected, and adjacent cached records in the time neighborhood are added to the candidate evidence records and deduplicated to generate a candidate evidence set. The review and judgment module is used to organize the message text, posting time, and candidate evidence set into a structured review input containing evidence number, timestamp, and spatial location information. The input is then fed into the visual language model to obtain the review and judgment result, the list of evidence numbers, and the confidence score. The review action determination module is used to calculate the evidence sufficiency score based on the candidate evidence set, the evidence number list, and the confidence level value, and to determine the review action based on the review judgment result and the evidence sufficiency score.
2. The system of claim 1, wherein, The multimodal cache generation module is used to extract video frames according to a preset base frame rate; calculate the ratio of the sum of the absolute values of the pixel-by-pixel differences of the downsampled grayscale images of adjacent frames to the total number of pixels in the frame image to obtain the screen change rate; when the screen change rate exceeds a preset change threshold, supplementary video frames are extracted at the time point of the change; when a new screen text region or a change in the content of an existing screen text region is detected, supplementary video frames are extracted; video frame records and screen text records are generated for the extracted video frames, and speech-to-text segments are generated for the audio stream.
3. The system of claim 2, wherein: Video frame records include a record type identifier, timestamp, image data, and semantic vector, the semantic vector being determined by video frame features; on-screen text records include a record type identifier, timestamp, text content, location coordinates, recognition confidence, and semantic vector; speech-to-text segments include a record type identifier, start timestamp, end timestamp, transcribed text content, and semantic vector; historical message records include a record type identifier, timestamp, message content, and semantic vector; real-time multimodal cache deletes cached records exceeding the window range using a sliding window and first-in-first-out strategy.
4. The system of claim 1, wherein, It also includes a message encoding module for extracting text content from text messages; mapping emoticons to semantic text descriptions for messages containing emoticons; generating speech-to-text transcription for voice messages; merging text content, semantic text descriptions, and speech-to-text transcription into message text in the order they appear in the message to be reviewed; and extracting image semantic vectors from images or stickers for messages to be reviewed.
5. The system according to claim 4, characterized in that, The context-dependent weight vector includes video image dependency weights, speech-to-text dependency weights, image-text dependency weights, and historical message dependency weights. A text encoder outputs a global semantic representation of the message text. The dependency weight prediction head maps the global semantic representation to the context-dependent weight vector. Specifically, when no image semantic vector exists, the global semantic representation is determined as the message semantic vector. When an image semantic vector exists, the global semantic representation and the image semantic vector are concatenated and mapped to the message semantic vector through a linear projection layer.
6. The system according to claim 1, characterized in that, The comprehensive scoring module is used to obtain a time proximity score by processing the time difference between the cached timestamp or the representative time determined by the start and end timestamps of the speech-transcribed segment and the publication time through time decay. The semantic similarity of the semantic vectors of the comments and the semantic vectors of the cached records is calculated using cosine similarity and truncated to the 0-1 range to obtain a semantic similarity score. For video frame records, the change significance score is obtained by normalizing the rate of change of the scene. For speech-to-text segments, the change significance score is determined by the cosine similarity between their semantic vectors and the semantic vectors of the previous speech-to-text segments. For on-screen text records, the change significance score is determined by whether the text is newly appearing or has changed in content. The comprehensive score of historical comment records consists of a time proximity score and a semantic similarity score.
7. The system according to claim 6, characterized in that, The candidate evidence generation module is used to sort the comprehensive scores in descending order by video frame records, on-screen text records, speech-to-text segments, and historical message records respectively; determine the number of selections for each type based on the context dependency weight vector and the upper limit of the total number of candidate evidences; when the dependency weight of a certain type is lower than the minimum dependency threshold, the number of selections for that type is set to zero, and candidate evidence records are obtained from the sorting results according to the number of selections for each type.
8. The system according to claim 7, characterized in that, The candidate evidence generation module is also used to record candidate video frames and obtain adjacent video frame records within a preset time range before and after their timestamps; when a candidate video frame record corresponds to a screen change event, add the most recent video frame record before the screen change event occurs; for candidate speech-to-text segments, obtain speech-to-text segments that are adjacent in time before and after; for candidate screen text records, obtain the video frame record of the frame in which it is located and the screen text records of the same text area in adjacent frames; retain the calculated comprehensive score for the supplemented adjacent cache records, merge them with the candidate evidence records, and remove duplicates according to the record type identifier and timestamp to generate a candidate evidence set.
9. The system according to claim 1, characterized in that, The review action determination module is used to: use the confidence level value as the confidence level indicator; divide the number of evidences in the evidence number list by the total number of evidences in the candidate evidence set to obtain the evidence coverage; input each piece of evidence corresponding to the evidence number list into a lightweight classifier to obtain the review tendency of each piece of evidence, and divide the number of evidences whose review tendency is consistent with the review judgment result by the number of evidences in the evidence number list to obtain the evidence consistency; divide the number of record types corresponding to the evidence number list by the total number of record types in the candidate evidence set to obtain the evidence source diversity; and weight and sum the confidence level indicator, evidence coverage, evidence consistency, and evidence source diversity according to a preset combination weight to obtain the evidence sufficiency score.
10. A real-time multimodal content review method for comments based on AI, executed by the system described in any one of claims 1-9, characterized in that, include: Based on video stream and message data, video frames are extracted to generate a real-time multimodal cache that includes video frame records, on-screen text records, speech-to-text segments, and historical message records. Convert the messages to be reviewed into message text, encode the messages to be reviewed, and generate context-dependent weight vectors and message semantic vectors; Based on the release time, the cache records of the real-time multimodal cache are calculated by type for time proximity score, semantic similarity score and change significance score. Combined with the context dependency weight vector, a comprehensive score is obtained. Candidate evidence records are selected based on the comprehensive score. Adjacent cache records in the time neighborhood are added to the candidate evidence records and duplicates are removed to generate a candidate evidence set. The message text, posting time, and candidate evidence set are organized into a structured review input containing evidence number, timestamp, and spatial location information. This input is then fed into a visual language model to obtain the review judgment result, the list of evidence numbers, and the confidence score. The sufficiency score is calculated based on the candidate evidence set, the evidence number list, and the confidence level value. The review action is determined based on the review judgment result and the sufficiency score.