Multi-modal content review method and device, medium and computer program

By combining a multimodal content review method with dedicated recognition models and artificial intelligence large models, the problems of poor universality and poor recognition effects in the existing technology are solved, and accurate review of complex and variable business scenarios is achieved, especially effective understanding and recognition of long videos.

CN120493147APending Publication Date: 2025-08-15BEIJING ZHONGKE PATTEK TECH CO LTD

Patent Information

Application Number
CN202510479196.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing content review methods have problems such as poor generality and poor identification results, especially in complex and changeable business scenarios, which are difficult to effectively identify violations.

Method used

A multimodal content review method is adopted, combined with a dedicated recognition model and an artificial intelligence model, and the recognition results of violations are output by identifying subtitles and voice information, and fusion analysis and reasoning are carried out.

Benefits of technology

It improves the versatility and accuracy of content review, can cope with complex and changeable business scenarios, enhances the ability to identify content that violates public order and good customs and politically sensitive, and solves the problem of insufficient understanding and cognitive ability of long videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493147A_ABST
    Figure CN120493147A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode content review method and device, a medium and a computer program, and relates to the technical field of content review. The multi-modal content review method comprises the following steps: calling a special recognition model to recognize subtitles and voice in target review content to obtain a character recognition result and a voice recognition result; the special identification model is a model obtained based on learning of a content sample and a subtitle label and a voice label of the content sample; and calling an artificial intelligence large model, performing fusion analysis and reasoning on the character recognition result, the voice recognition result and the video picture of the target review content according to the violation review rule, and outputting a violation behavior recognition result. The universality of the content review method is improved based on the good cognitive ability of the artificial intelligence large model; the artificial intelligence large model and the special small model are organically combined, the understanding cognitive ability of the artificial intelligence large model and the perception ability of the special small model are fully exerted, and accurate and effective examination of the content is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of content review technology, and in particular to a multimodal content review method, device, medium and computer program. Background Art

[0002] Current AI technologies for content review primarily rely on specialized AI models, such as those used for sensitive speech recognition, sensitive face recognition, and sensitive keyword detection. However, specialized models for content review have certain drawbacks, primarily manifested in poor model versatility, low learning efficiency, and insufficient subjective cognitive capabilities.

[0003] (1) Poor model versatility: Dedicated models are highly customized, and a single model can only complete a single task. Moreover, the technical frameworks used by the models are also different. Different downstream tasks require the collection and construction of different training data sets. For example, in text classification, whenever the need to increase or optimize the text recognition capability of a certain type is needed, the corresponding training data needs to be collected. Therefore, the versatility and adaptability of the model are insufficient.

[0004] (2) Low learning efficiency: The training of specialized models is usually a big data + supervised approach. Although some specialized models can already achieve a certain degree of small sample set training, they still require a certain amount of training data to produce a good recognition effect, which is costly. Moreover, in the application scenarios of content review segmentation, positive sample data is often scarce, that is, effective violation data itself is rare and difficult to collect, which is not enough to support the training of specialized models.

[0005] (3) Dedicated models do not have subjective cognitive capabilities: To understand a certain intention, it is generally necessary to use a dedicated model to identify and analyze single-modal data and then make multimodal fusion decisions. However, the above methods are often not ideal for understanding intentions.

[0006] From the above, we can see that the existing content review methods have the defects of poor versatility and poor recognition effect, and are unable to cope with the current situation where new ideological and social problems are emerging in an endless stream. Therefore, there is an urgent need for a content review method that can effectively cope with complex and changing businesses. Summary of the Invention

[0007] The present invention provides a multimodal content review method, device, medium and computer program to address the defects of existing content review methods such as poor versatility and poor recognition effect, so as to achieve effective review of complex and changeable content.

[0008] The present invention provides a multimodal content review method, comprising the following steps.

[0009] Calling a dedicated recognition model to recognize subtitles and speech in the target review content to obtain text recognition results and speech recognition results; the dedicated recognition model is a model learned based on content samples and subtitle labels and speech labels of the content samples;

[0010] The artificial intelligence large model is called to perform fusion analysis and reasoning on the text recognition results, the voice recognition results and the video footage of the target review content according to the violation review rules, and output the violation recognition results.

[0011] According to a multimodal content review method provided by the present invention, a large artificial intelligence model is called to perform fusion analysis and reasoning on the text recognition results, the speech recognition results, and the video images of the target review content according to the violation review rules, including:

[0012] The artificial intelligence big model is called, and the violation review rules are input into the artificial intelligence big model. The artificial intelligence big model performs fusion analysis and reasoning on the text recognition results, the voice recognition results and the video images of the target review content based on the violation review rules; the violation review rules include prompt words related to the violation, a knowledge base and a review thinking logic chain.

[0013] According to a multimodal content review method provided by the present invention, before calling a dedicated recognition model to recognize subtitles and speech in the target review content, the method further includes:

[0014] Determine whether the length of the video to be reviewed is greater than the set threshold;

[0015] If not, the video to be reviewed is determined as the target review content;

[0016] If so, a dedicated segmentation model is called to segment the video to be reviewed into scenes to obtain multiple scene sub-videos, and the scene sub-videos are used as the target review content.

[0017] According to a multimodal content review method provided by the present invention, a large artificial intelligence model is called to perform fusion analysis and reasoning on the text recognition results, the speech recognition results, and the video images of the target review content according to the violation review rules, including:

[0018] Calling the artificial intelligence large model to describe and understand the images of each of the scene sub-videos respectively, and obtaining description and understanding results corresponding to each of the scene sub-videos;

[0019] Fusing the description understanding results in chronological order to obtain a fusion result;

[0020] The artificial intelligence large model is called to perform fusion analysis and reasoning on the text recognition results, the speech recognition results and the fusion results according to the violation review rules.

[0021] According to a multimodal content review method provided by the present invention, a dedicated recognition model is called to recognize subtitles and speech in target review content, including:

[0022] Separating the target review content from video and audio to obtain a target review video and a target review audio;

[0023] Performing frame extraction processing on the target review video to obtain a frame image;

[0024] Using a text recognition model to recognize the subtitles in the frame image to obtain the text recognition result;

[0025] A speech recognition model is used to recognize the speech information in the target review audio to obtain the speech recognition result.

[0026] According to a multimodal content review method provided by the present invention, the text recognition model is an OCR text recognition model.

[0027] According to a multimodal content review method provided by the present invention, the speech recognition model is an ASR speech recognition model.

[0028] The present invention also provides a multimodal content review method, comprising the following modules:

[0029] A dedicated model calling module is used to call a dedicated recognition model to recognize subtitles and speech in the target review content and obtain text recognition results and speech recognition results; the dedicated recognition model is a model learned based on content samples and their subtitle labels and speech labels;

[0030] The artificial intelligence large model calling module is used to call the artificial intelligence large model, and according to the violation review rules, perform fusion analysis and reasoning on the text recognition results, the voice recognition results and the video images of the target review content, and output the violation behavior recognition results.

[0031] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements any of the multimodal content review methods described above.

[0032] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the multimodal content review methods described above.

[0033] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the multimodal content review methods described above.

[0034] The present invention provides a multimodal content review method, device, medium, and computer program. By using a dedicated recognition model to identify the subtitles and voice in the target review content, and using an artificial intelligence large model to perform fusion analysis and reasoning on the text recognition results, voice recognition results, and video images of the target review content according to the review rules, the review result is finally given. The present invention improves the versatility of the content review method based on the good cognitive ability of the artificial intelligence large model, enabling it to cope with complex and changing review business; the artificial intelligence large model is organically combined with the dedicated small model, giving full play to the understanding and cognitive ability of the artificial intelligence large model and the perception ability of the dedicated small model, realizing accurate and effective review of content. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0036] Figure 1 This is one of the flow charts of the multimodal content review method provided in an embodiment of the present invention.

[0037] Figure 2 This is the second flow chart of the multimodal content review method provided by an embodiment of the present invention.

[0038] Figure 3 This is a schematic diagram of the review capability customization process provided by an embodiment of the present invention.

[0039] Figure 4 This is a schematic diagram of the long video recognition capability optimization process provided by an embodiment of the present invention.

[0040] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0042] The following combination Figure 1-Figure 4 Multimodal Content Review Method The multimodal content review method of the present invention is described.

[0043] Figure 1 This is a flow chart of the multimodal content review method provided by the present invention, such as Figure 1 As shown, the method includes the following:

[0044] Step 101: Call a dedicated recognition model to recognize subtitles and voice in the target review content to obtain text recognition results and voice recognition results; the dedicated recognition model is a model learned based on content samples and their subtitle labels and voice labels.

[0045] Step 102: Call the artificial intelligence model, and according to the violation review rules, perform fusion analysis and reasoning on the text recognition results, the voice recognition results, and the video footage of the target review content, and output the violation recognition results.

[0046] Because traditional AI-specific models (also known as specialized small models or small models) are trained based on a single task's training set, their versatility is limited. Furthermore, the sample data used to train these specialized models must include illegal content. However, in the content review industry, illegal content is a small sample size and is very rare. Therefore, specialized models trained on this sample data are limited in their recognition capabilities and cannot achieve good recognition results.

[0047] Based on this, the embodiment of the present invention introduces an artificial intelligence big model (for example, it can be Alibaba's artificial intelligence big model - Tongyi Qianwen). Since the artificial intelligence big model has a strong ability to integrate semantic understanding and analysis, it can learn and understand the context of the knowledge base data in the regulatory field very well, and can also perform logical reasoning based on the manually preset content review thinking chain. It does not require a large amount of violation data to achieve accurate and effective review and identification. It only needs to input prompt words and review logic chains into the artificial intelligence big model, and the artificial intelligence big model can use its own powerful analysis, understanding and reasoning capabilities to review and identify the target review content, which has strong versatility. There is no need to train a corresponding special model for each new review business type, thereby avoiding the problem of poor review and identification effects caused by the small number of positive samples in the special model training data set. At the same time, due to the powerful cognitive understanding ability of the artificial intelligence big model, the accuracy of review and identification is further guaranteed.

[0048] See also Figure 2 In this embodiment of the present invention, a dedicated small model is used to perceive the target content (specifically identifying text and voice information), and the results perceived by the dedicated small model are then passed to the large artificial intelligence model for integrated semantic understanding and analysis. That is, the embodiment of the present invention uses a dedicated small model to accurately perceive the content and a large artificial intelligence model to understand and reason about the content, giving full play to the respective advantages of the large and small models, thereby significantly improving the actual application effect of content review in the field of monitoring and supervision.

[0049] See also Figure 3 In some embodiments of the present invention, based on the Zero Shot and Few Shot capabilities of the AI large-scale model, a content review capability customization paradigm is defined, and a review capability customization factory is constructed to support the instant customization of new content review capabilities for complex content review scenarios and emergency task scenarios. This can be understood as providing a teaching channel for the present invention, allowing people to "hands-on" teach intelligent agents how to deal with new problems. It's like teaching a new reviewer a specific purpose for watching a video, and then how to combine the image and text content in the video with industry knowledge to analyze and review according to a specific thought process. The entire customization process is text-based, making it easy to use and even for ordinary people. It also provides an instant capability testing and verification function. Once the verification is successful, the new capability can be stored in the intelligent agent for subsequent daily review processing.

[0050] The following describes the on-the-fly customization process of content review capabilities.

[0051] When censoring and identifying a specific ideology (or otherwise illegal behavior), the AI model must be fed with relevant censorship rules (teaching it how to identify illegal behavior) and then stored in its backend. When the AI model is invoked, it identifies illegal behavior in the target content based on these censorship rules.

[0052] like Figure 3 As shown, the embodiment of the present invention specifically defines a content review capability customization paradigm to express the above-mentioned review rules, so as to cope with complex content review scenarios and emergency task scenarios, and to realize the instant customization and application of new content review capabilities. The content review capability customization paradigm includes modules such as customized prompt words (to clearly understand the purpose of the content), construction of a review industry knowledge base, and construction of a logical thinking chain. In summary, when faced with an emergency task scenario or a new task scenario, the embodiment of the present invention can realize the instant customization of new content review capabilities through the paradigms of customized prompt words, construction of a review industry knowledge base, and construction of a logical thinking chain.

[0053] It should be noted that when customizing new review capabilities in real time, a test verification step is also included. That is, after customizing the prompt words, building the review industry knowledge base, and building the logical thinking chain, a certain content to be reviewed is recognized by a dedicated recognition model for text, voice, etc. Figure 3 Input the "Review Content Input" module into the artificial intelligence big model, click the verification module, and the artificial intelligence big model will identify the content to be reviewed based on the above prompt words, knowledge base, and logical thinking chain and output the violation result. If the violation output result is correct, click the save module to solidify and save the above customized prompt words, constructed review industry knowledge base, and constructed logical thinking chain into the knowledge base in the background of the artificial intelligence big model for subsequent use by the artificial intelligence big model in the actual recognition and review process. If the violation output result is incorrect, adjust the customized prompt words, constructed review industry knowledge base, and constructed logical thinking chain until the test verification effect meets the requirements.

[0054] In some embodiments of the present invention, Figure 2 As shown, before calling the dedicated recognition model to recognize the subtitles and speech in the target review content in step 101, some preprocessing is performed on the target review content. Specifically, the video and audio of the target review content are separated to obtain the target review video and target review audio. The target review video is frame-processed to obtain frame images.

[0055] Step 101 calls a dedicated recognition model to identify the subtitles and voice in the target review content, specifically: using a text recognition model to identify the title of the target review video and the subtitles in the frame image to obtain the text recognition result; using a voice recognition model to identify the voice information in the target review audio to obtain the voice recognition result.

[0056] In some embodiments of the present invention, the text recognition model is an OCR text recognition model.

[0057] OCR (Optical Character Recognition) is a technology that converts text in an image into editable text. OCR text recognition models are widely used in computer vision and natural language processing, such as converting scanned documents into editable text files, automatically reading license plate numbers, processing handwritten text, etc. Its workflow mainly includes: image preprocessing, text detection, text recognition and post-processing. Among them, image preprocessing: includes operations such as denoising, binarization, and rotation correction to improve image quality and text readability. Text detection: uses a convolutional neural network (CNN) to detect areas containing text in an image. Text recognition: converts the image in the detected text area into editable text, and technologies such as recurrent neural networks (RNN) and long short-term memory networks (LSTM) can be used. Post-processing: includes operations such as spell checking and format correction to improve the accuracy of the final output text.

[0058] In some embodiments of the present invention, the speech recognition model is an ASR speech recognition model.

[0059] The ASR (Automatic Speech Recognition) model is a technology that converts human voice into computer-readable text. The ASR model consists of three main components: front-end processing, acoustic model, and language model. Front-end processing is responsible for processing and extracting features from input signals. It processes and extracts features from the input audio signal for subsequent acoustic recognition and language processing. The acoustic model is the core component responsible for converting the input speech signal into a text representation. By training on a large number of speech samples, the acoustic model learns and builds a model corresponding to the speech signal. The language model is used to convert the text representation into readable commands or instructions. By analyzing the language features and contextual information involved in the speech signal, the language model achieves text-to-command conversion.

[0060] Due to the unique nature of content review and regulation, large-scale models are required to be capable of reviewing long videos. However, current large-scale models, both domestic and international, are inadequate for understanding long videos, typically longer than five minutes. They currently lack sufficient support for understanding and recognizing long videos. This is primarily due to two factors: First, the most authoritative publicly available video understanding datasets primarily focus on short videos of simple scenes. Second, the vast majority of large-scale models, both domestic and international, are based on the ImageLLM approach, which sequentially feeds multiple images from a video after frame extraction. However, for long videos, this frame extraction strategy misses a lot of important information, resulting in insufficient understanding of long videos. Specifically, the ImageLLM-based frame extraction process only extracts a fixed number of frames. For short videos, the extracted frames are sufficient to represent the intended content. However, for long videos, the fixed number of frames results in a large gap between frames, which results in a loss of visual information. This results in an inaccurate description of the entire video, and the extracted frames fail to fully represent the intended content.

[0061] Therefore, in order to improve the accuracy of recognition review, the embodiment of the present invention has made corresponding optimization processing for the recognition review of long videos. The basic optimization idea is to use shots as units, understand and recognize each shot segment through a sliding window, and finally combine the results of voice recognition and video picture recognition for fusion analysis and judgment, and format the results after calibration and positioning.

[0062] Specifically, the embodiment of the present invention first determines whether the length of the video to be reviewed is greater than a set threshold. If the length of the video to be reviewed is not greater than the set threshold, the method in the above embodiment is directly used to organically combine the dedicated small model and the artificial intelligence large model to identify and review the video to be reviewed; if the length of the video to be reviewed is greater than the set threshold, the dedicated segmentation model is called to perform scene segmentation on the video to be reviewed to obtain multiple scene sub-videos, and then the method in the above embodiment is used to organically combine the dedicated small model and the artificial intelligence large model to identify and review each scene sub-video. That is, the embodiment of the present invention divides a long video into multiple short videos according to different scenes, and transforms the identification and review of the long video into the identification and review of the short video.

[0063] It should be noted that the dedicated segmentation model can be the BaSSL video scene segmentation model, or other existing video scene segmentation models. The dedicated segmentation model here is an existing model and will not be described in detail in this article.

[0064] like Figure 4As shown, after a long video is input, the long video is first divided according to the scene shots, and then the artificial intelligence big model is called to describe and understand the short videos of each scene, and the description and understanding results of the artificial intelligence big model for each scene short video are recorded. Then, based on the time sequence, the above description and understanding results are fused together, and the artificial intelligence big model is called again to combine the regulations, prompt words and thinking logic chain in the knowledge base to perform fusion analysis and judgment on the fusion results, and output the violation identification results.

[0065] The aforementioned process of segmenting long videos into short videos and then describing and understanding them yields relatively detailed and comprehensive descriptions and understanding results. Fusion of the description and understanding results from each short video for overall analysis and judgment yields relatively accurate results. This embodiment of the present invention utilizes a large artificial intelligence model to comprehensively analyze and judge the description and understanding results from each short video, enabling more accurate identification of violations.

[0066] The following is an introduction to the entire processing flow of long videos.

[0067] When a long video is input, a dedicated segmentation model is used to segment it into scene shots. A large AI model is then used to describe and understand each segmented short video. The description and understanding results for each short video are then fused chronologically to produce a fusion of description and understanding. A text recognition model is used to recognize text in the long video, and a speech recognition model is used to recognize speech in the long video. The large AI model, based on the legal knowledge, prompts, and logical chains stored in the knowledge base, integrates and analyzes these fusion results, along with the text and speech recognition results, to produce a violation identification result.

[0068] It should be noted that, in order to improve work efficiency, the segmentation and recognition operations of the above-mentioned dedicated segmentation model, text recognition model and speech recognition model are performed in parallel.

[0069] The multimodal content review method provided by the embodiment of the present invention has the following advantages:

[0070] Advantage 1: The method of conducting content review by integrating a large artificial intelligence model with a dedicated small model effectively makes up for the shortcomings of the dedicated small model in semantic understanding and cognitive ability, and greatly improves the completeness and precision of content review, especially the ability to review content that violates public order and good morals and is politically sensitive.

[0071] Advantage 2: It effectively solves the problem of insufficient understanding and recognition ability of general artificial intelligence large models for long videos. It can more comprehensively recognize all the content in long videos and generate accurate review results by combining thought chains and industry knowledge bases.

[0072] Advantage 3: In response to new ideological or social issues, review personnel can quickly customize and build new content review capabilities, thereby changing the learning model of traditional customized small models that require a large amount of relevant data samples to be collected for training. New content review thinking chains can be created through a visual interface with zero samples, and verified and tested online to ensure that they can be solidified and saved after achieving the expected review effect, and used for normalized content review and monitoring and supervision, so as to better adapt to the complex and changeable monitoring and supervision business requirements.

[0073] The following describes the multimodal content review device provided by the present invention. The multimodal content review device described below and the multimodal content review method described above can be referenced to each other. Referring to the figure, the multimodal content review device includes:

[0074] A dedicated model calling module is used to call a dedicated recognition model to recognize subtitles and speech in the target review content and obtain text recognition results and speech recognition results; the dedicated recognition model is a model learned based on content samples and their subtitle labels and speech labels;

[0075] The artificial intelligence large model calling module is used to call the artificial intelligence large model, and according to the violation review rules, perform fusion analysis and reasoning on the text recognition results, the voice recognition results and the video images of the target review content, and output the violation behavior recognition results.

[0076] The long video recognition module is used to determine whether the length of the video to be reviewed is greater than the set threshold. If the length of the video to be reviewed is greater than the set threshold, the module jumps to the dedicated model calling module. The dedicated model calling module is also used to call the dedicated segmentation model to segment the video to be reviewed into scenes to obtain multiple scene sub-videos. The artificial intelligence large model calling module is also used to call the artificial intelligence large model to describe and understand the images of each scene sub-video respectively, obtain the description and understanding results corresponding to each scene sub-video, and call the artificial intelligence large model to perform fusion analysis and reasoning on the text recognition results, the speech recognition results, and the fusion results (the fusion results obtained by fusion of the description and understanding results in chronological order) according to the violation review rules.

[0077] As for how each module in the device works specifically, you can refer to the multimodal content review method embodiment described above and will not go into details here.

[0078] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5As shown, the electronic device may include: a processor (processor) 510, a communication interface (Communications Interface) 520, a memory (memory) 530 and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logic instructions in the memory 530 to execute the multimodal content review method, which includes: calling a dedicated recognition model to identify the subtitles and voice in the target review content to obtain text recognition results and voice recognition results; the dedicated recognition model is a model learned based on content samples and the subtitle tags and voice tags of the content samples; calling an artificial intelligence large model, according to the violation review rules, the text recognition results, the voice recognition results and the video images of the target review content are fused, analyzed and reasoned, and the violation identification results are output. As for a more specific method, please refer to the multimodal content review method described above, which will not be repeated here.

[0079] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0080] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multimodal content review method provided by the above methods, which includes: calling a dedicated recognition model to identify subtitles and voices in the target review content to obtain text recognition results and voice recognition results; the dedicated recognition model is a model learned based on content samples and subtitle labels and voice labels of content samples; calling an artificial intelligence large model to perform fusion analysis and reasoning on the text recognition results, the voice recognition results and the video images of the target review content according to the violation review rules, and output the violation identification results. As for a more specific method, please refer to the multimodal content review method described above, which will not be repeated here.

[0081] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented by the processor to execute the multimodal content review method provided by the above methods, the method comprising: calling a dedicated recognition model to identify the subtitles and voice in the target review content, and obtaining text recognition results and voice recognition results; the dedicated recognition model is a model learned based on content samples and the subtitle labels and voice labels of the content samples; calling an artificial intelligence large model, and performing fusion analysis and reasoning on the text recognition results, the voice recognition results and the video images of the target review content according to the violation review rules, and outputting the violation identification results. As for a more specific method, please refer to the multimodal content review method described above, which will not be repeated here.

[0082] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0083] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multimodal content review method, characterized in that: include: Calling a dedicated recognition model to identify the subtitles and voice in the target review content to obtain text recognition results and voice recognition results; The dedicated recognition model is a model learned based on content samples and their subtitle labels and voice labels; The artificial intelligence large model is called to perform fusion analysis and reasoning on the text recognition results, the voice recognition results and the video footage of the target review content according to the violation review rules, and output the violation recognition results.

2. The multimodal content review method according to claim 1, characterized in that: The artificial intelligence model is called to perform fusion analysis and reasoning on the text recognition results, the speech recognition results, and the video footage of the target review content according to the violation review rules, including: The artificial intelligence big model is called, and the violation review rules are input into the artificial intelligence big model. The artificial intelligence big model performs fusion analysis and reasoning on the text recognition results, the voice recognition results and the video images of the target review content based on the violation review rules; the violation review rules include prompt words related to the violation, a knowledge base and a review thinking logic chain.

3. The multimodal content review method according to claim 1, characterized in that: Before calling the dedicated recognition model to recognize the subtitles and voice in the target review content, it also includes: Determine whether the length of the video to be reviewed is greater than the set threshold; If not, the video to be reviewed is determined as the target review content; If so, a dedicated segmentation model is called to segment the video to be reviewed into scenes to obtain multiple scene sub-videos, and the scene sub-videos are used as the target review content.

4. The multimodal content review method according to claim 3, characterized in that: The artificial intelligence model is called to perform fusion analysis and reasoning on the text recognition results, the speech recognition results, and the video footage of the target review content according to the violation review rules, including: Calling the artificial intelligence big model to describe and understand the images of each of the scene sub-videos respectively, and obtaining description and understanding results corresponding to each of the scene sub-videos; Fusing the description understanding results in chronological order to obtain a fusion result; The artificial intelligence large model is called to perform fusion analysis and reasoning on the text recognition results, the speech recognition results and the fusion results according to the violation review rules.

5. The multimodal content review method according to claim 1, characterized in that: Calling a dedicated recognition model to identify subtitles and voice in the target review content, including: Separating the target review content from video and audio to obtain a target review video and a target review audio; Performing frame extraction processing on the target review video to obtain a frame image; Using a text recognition model to recognize the subtitles in the frame image to obtain the text recognition result; A speech recognition model is used to recognize the speech information in the target review audio to obtain the speech recognition result.

6. The multimodal content review method according to claim 5, characterized in that: The text recognition model is an OCR text recognition model.

7. The multimodal content review method according to claim 5, characterized in that: The speech recognition model is an ASR speech recognition model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, it implements the multimodal content review method as described in any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the multimodal content review method as described in any one of claims 1 to 7.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the multimodal content review method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Short video auditing method based on multiple modes

    CN115512259A

  • Media multi-mode content auditing method and system based on artificial intelligence technology

    CN119068399A

Cited By

  • Method and device for realizing operation specification review based on multi-modal large model, processor and computer readable storage medium thereof

    CN121233764A

  • Method and device for implementing operation specification review based on multi-modal large model, processor and computer readable storage medium thereof

    CN121233764B