Live broadcast violation information identification method and device, equipment, storage medium and computer program product

By converting live voice into speech text and extracting key video frames, combined with multimodal large model, the problem of low accuracy and recall of live broadcast violation information in the prior art is solved, and more efficient identification of violation information is achieved.

CN120034668APending Publication Date: 2025-05-23BEIJING QIHOOD TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510168917.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing methods of identifying live broadcast violation information rely on manual monitoring, resulting in low accuracy and recall.

Method used

By converting live voice into speech text and segmenting, the corresponding key video frames are extracted for character recognition, and the violation recognition results are generated in combination with the multimodal big model.

Benefits of technology

It improves the accuracy and recall rate of live broadcast violation information identification and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034668A_ABST
    Figure CN120034668A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a live broadcast violation information identification method and device, equipment, a storage medium and a computer program product, and the method comprises the steps: converting a live broadcast voice of to-be-identified live broadcast information into a voice text, segmenting the voice text, obtaining segmented voice texts, and storing the segmented voice texts in a server; key video frames corresponding to the segmented voice texts are extracted from a video picture of the to-be-recognized live broadcast information, character recognition is performed on the key video frames to obtain video frame texts, and violation recognition results of the to-be-recognized live broadcast information are generated through a multi-mode large model based on the key video frames, the segmented voice texts and the video frame texts; according to the method, the live broadcast violation information is identified by the multi-modal large model, and the image information, the voice information and the text information are considered, so that the accuracy and the recall rate of live broadcast violation information identification can be improved, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, storage medium and computer program product for identifying live broadcast illegal information. Background Art

[0002] At present, with the rise of live broadcasting, the supervision of live broadcasting information has become a difficult problem. The existing methods for identifying illegal live broadcasting information usually rely on manual monitoring of illegal information in live broadcasting, which has the defects of low accuracy and low recall rate. Summary of the invention

[0003] The main purpose of this application is to provide a method, device, equipment, storage medium and computer program product for identifying live broadcast violation information, aiming to solve the technical problem that related live broadcast violation information identification methods rely on manual monitoring of violation information in live broadcasts, resulting in low accuracy and low recall rate.

[0004] To achieve the above purpose, the present application provides a method for identifying live broadcast violation information, the method comprising:

[0005] Converting the live speech of the live broadcast information to be recognized into speech text, and segmenting the speech text to obtain segmented speech text;

[0006] Extracting key video frames corresponding to the segmented speech text from the video screen of the live broadcast information to be identified, and performing character recognition on the key video frames to obtain video frame text;

[0007] Based on the key video frames, the segmented voice texts and the video frame texts, a violation identification result of the live broadcast information to be identified is generated through a multimodal large model.

[0008] Optionally, extracting key video frames corresponding to the segmented voice text from the video screen of the live broadcast information to be identified, and performing character recognition on the key video frames to obtain video frame texts includes:

[0009] Dividing the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice texts;

[0010] Performing frame processing on the video segments to obtain key video frames corresponding to the segmented speech texts;

[0011] Character recognition is performed on the key video frame through an image recognition model to obtain video frame text.

[0012] Optionally, performing frame processing on the video segment to obtain key video frames corresponding to the segmented speech text includes:

[0013] Extracting frames from the video segment to obtain a set of video frames corresponding to the video segment;

[0014] The video frame set is deduplicated to obtain key video frames corresponding to the segmented speech text.

[0015] Optionally, deduplicating the video frame set to obtain key video frames corresponding to the segmented speech text includes:

[0016] Identify the image features of each video frame in the video frame set by using a preset image semantic model;

[0017] The video frame set is deduplicated based on the image features to obtain key video frames corresponding to the segmented speech text.

[0018] Optionally, dividing the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice texts includes:

[0019] Obtain the start time and end time of the segmented speech text;

[0020] The video screen of the live broadcast information to be identified is divided into video segments corresponding to the segmented voice texts according to the start time and the end time.

[0021] Optionally, the generating the violation identification result of the live broadcast information to be identified by using a multimodal large model based on the key video frame, the segmented voice text and the video frame text includes:

[0022] Generate prompt information of a multimodal large model according to the key video frame, the segmented voice text and the video frame text;

[0023] The prompt information is input into the multimodal large model, and a violation identification result of the live broadcast information to be identified is generated through the multimodal large model.

[0024] Optionally, the generating prompt information of the multimodal large model according to the key video frame, the segmented voice text and the video frame text includes:

[0025] Generate a violation identification task description according to the live broadcast information to be identified;

[0026] The key video frame, the segmented speech text, the video frame text and the violation identification task description are spliced ​​to obtain prompt information of a multimodal large model.

[0027] Optionally, the step of splicing the key video frame, the segmented speech text, the video frame text, and the violation identification task description to obtain prompt information of a multimodal large model includes:

[0028] Matching the corresponding target regulation library based on the live broadcast information to be identified;

[0029] The key video frames, the segmented speech texts, the video frame texts, the violation identification task descriptions and the target regulations library are spliced ​​to obtain prompt information of a multimodal large model.

[0030] Optionally, the step of inputting the prompt information into the multimodal large model and generating a violation identification result of the live broadcast information to be identified by using the multimodal large model includes:

[0031] Inputting the prompt information into the multimodal large model to obtain the penalty result and reasoning reason output by the multimodal large model;

[0032] The penalty result and the reasoning reason are analyzed to obtain a violation identification result of the live broadcast information to be identified.

[0033] Optionally, before converting the live voice of the live information to be identified into voice text and segmenting the voice text to obtain the segmented voice text, the method further includes:

[0034] Obtain historical penalty cases, and construct a supervised fine-tuning dataset based on the historical penalty cases;

[0035] The preset base model is fine-tuned based on the supervised fine-tuning dataset to obtain a multimodal large model.

[0036] Optionally, obtaining historical penalty cases and constructing a supervised fine-tuning dataset based on the historical penalty cases includes:

[0037] Obtaining historical penalty cases, and cleaning the historical penalty cases to obtain cleaned penalty cases;

[0038] The cleaned penalty cases are disambiguated to obtain disambiguated penalty cases, and a supervised fine-tuning dataset is constructed based on the disambiguated penalty cases.

[0039] In addition, to achieve the above purpose, the present application also proposes a live broadcast violation information identification device, the live broadcast violation information identification device comprising:

[0040] A voice processing module, used for converting the live voice of the live information to be recognized into voice text, and segmenting the voice text to obtain segmented voice text;

[0041] A video frame processing module is used to extract key video frames corresponding to the segmented voice text from the video screen of the live broadcast information to be identified, and perform character recognition on the key video frames to obtain video frame text;

[0042] A violation identification module is used to generate a violation identification result of the live broadcast information to be identified based on the key video frame, the segmented voice text and the video frame text through a multimodal large model.

[0043] Optionally, the video frame processing module is also used to divide the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice text; perform frame processing on the video segments to obtain key video frames corresponding to the segmented voice text; perform character recognition on the key video frames through an image recognition model to obtain video frame text.

[0044] Optionally, the video frame processing module is further used to extract frames from the video segment to obtain a set of video frames corresponding to the video segment; and to deduplicate the set of video frames to obtain key video frames corresponding to the segmented speech text.

[0045] Optionally, the video frame processing module is also used to identify image features of each video frame in the video frame set through a preset image semantic model; deduplicate the video frame set based on the image features to obtain key video frames corresponding to the segmented speech text.

[0046] Optionally, the video frame processing module is also used to obtain the start time and end time of the segmented voice text; and divide the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice text according to the start time and the end time.

[0047] Optionally, the violation identification module is also used to generate prompt information of the multimodal large model based on the key video frame, the segmented voice text and the video frame text; input the prompt information into the multimodal large model, and generate a violation identification result of the live broadcast information to be identified through the multimodal large model.

[0048] In addition, to achieve the above-mentioned purpose, the present application also proposes a live broadcast violation information identification device, which includes a memory, a processor, and a live broadcast violation information identification program stored on the memory and executable on the processor, and the live broadcast violation information identification program is configured to implement the live broadcast violation information identification method as described above.

[0049] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, on which a live broadcast violation information identification program is stored. When the live broadcast violation information identification program is executed by a processor, the live broadcast violation information identification method described above is implemented.

[0050] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a live broadcast violation information identification program, and when the live broadcast violation information identification program is executed by a processor, it implements the live broadcast violation information identification method described above.

[0051] One or more technical solutions proposed in this application have at least the following technical effects:

[0052] In the present application, it is disclosed that the live speech of the live information to be identified is converted into speech text, and the speech text is segmented to obtain segmented speech text, key video frames corresponding to the segmented speech text are extracted from the video screen of the live information to be identified, and character recognition is performed on the key video frames to obtain video frame text, and based on the key video frames, the segmented speech text and the video frame text, a violation identification result of the live information to be identified is generated through a multimodal large model; since the present application identifies the live violation information with a multimodal large model, and simultaneously considers the image information, speech information and text information, it can improve the accuracy and recall rate of the live violation information identification, and thus can improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0054] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0055] Figure 1 This is a flow chart of the first embodiment of the live broadcast violation information identification method of this application;

[0056] Figure 2 This is a flow chart of the second embodiment of the live broadcast violation information identification method of this application;

[0057] Figure 3 This is a flow chart of the third embodiment of the live broadcast violation information identification method of this application;

[0058] Figure 4 This is a specific flow chart of an embodiment of a method for identifying live broadcast violation information of this application;

[0059] Figure 5 This is a schematic diagram of the module structure of the device for identifying live broadcast violation information according to an embodiment of the present application;

[0060] Figure 6This is a schematic diagram of the device structure of the hardware operating environment involved in the live broadcast violation information identification method in the embodiment of the present application.

[0061] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0062] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0063] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0064] At present, with the rise of live broadcasting, the supervision of live broadcasting information has become a difficult problem. The supervision of live broadcasting information currently faces the following difficulties: (1) Due to the long time nature of live broadcasting information, manual supervision is costly and inefficient. (2) Rule-based supervision and traditional model-based methods only use text and voice information, but lack image information, resulting in low recognition accuracy and recall rate, which cannot meet the supervision needs.

[0065] Therefore, in order to overcome the above-mentioned defects, the present application provides a solution, which includes: converting the live voice of the live information to be identified into voice text, and segmenting the voice text to obtain segmented voice text, extracting key video frames corresponding to the segmented voice text from the video screen of the live information to be identified, and performing character recognition on the key video frames to obtain video frame text, and generating a violation identification result of the live information to be identified through a multimodal large model based on the key video frames, the segmented voice text and the video frame text; since the present application identifies the live violation information with a multimodal large model, and takes into account the image information, voice information and text information at the same time, it can improve the accuracy and recall rate of the live violation information identification, and thus improve the user experience.

[0066] It should be noted that the executor of this embodiment can be a live broadcast violation information identification device with data processing, network communication and program running functions, such as a server, etc., or other electronic devices that can achieve the same or similar functions, and this embodiment does not impose any restrictions on this.

[0067] Based on this, the present application embodiment provides a method for identifying live broadcast violation information, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the live broadcast violation information identification method of the present application.

[0068] In a first embodiment, the live broadcast violation information identification method includes:

[0069] Step S10: converting the live speech of the live broadcast information to be recognized into speech text, and segmenting the speech text to obtain segmented speech text.

[0070] It should be understood that the live broadcast information to be identified may refer to live broadcast content that needs to be identified as illegal or unlawful content, including video images, live broadcast voice, live broadcast title and other information. In a specific implementation, the live broadcast information to be identified may be live broadcast advertisements, live video broadcasts and other information. Live broadcast voice may refer to the audio part of the live broadcast content, including the host's explanation, background music, etc. Speech text may refer to the text content converted from live broadcast voice through Automatic Speech Recognition (ASR) technology. Segmented speech text may refer to multiple text segments that are divided into speech text according to specific strategies (such as semantic integrity, time interval, etc.).

[0071] In the specific implementation, the ASR model is used to convert the live speech into text content, that is, speech text. The speech text is segmented according to the preset strategy (such as based on punctuation, pause time, semantic integrity, etc.) to form multiple text segments, each of which contains complete semantic information or corresponds to a specific time period. For example, the ASR model is used to convert the live speech into speech text, and the speech text is divided into sentences according to the preset strategy. The data format is (start_time, end_time, text), where text is the segmented speech text, start_time is the start time of the segmented speech text, and end_time is the end time of the segmented speech text.

[0072] Furthermore, in order to enable the multimodal large model to more accurately identify illegal content in live broadcast information, before step S10, it also includes: obtaining historical penalty cases, and constructing a supervised fine-tuning data set based on the historical penalty cases; fine-tuning the preset base model based on the supervised fine-tuning data set to obtain a multimodal large model. Among them, historical penalty cases may refer to live broadcast information cases (such as live broadcast advertising cases) that have occurred in the past and were judged to be illegal or illegal, and the historical penalty cases contain the specific content of the live broadcast information, the penalty results, and possible reasons for the penalty. The supervised fine-tuning data set may refer to a data set constructed based on historical penalty cases, which is used to fine-tune the preset base model. The supervised fine-tuning data set contains positive examples (such as legal advertisements) and negative examples (such as illegal or illegal advertisements), as well as corresponding labels (legal / illegal). The preset base model may refer to a pre-trained multimodal model that can process image and text data. The preset base model is the basis for building a multimodal large model, but it needs to be fine-tuned through supervision to adapt to specific illegal identification tasks. The multimodal large model can refer to a preset base model after supervised fine-tuning, which can more accurately identify illegal or illegal content in live broadcast advertisements. The model can comprehensively utilize information from multiple modalities such as images, voice and text.

[0073] In the specific implementation, we collect cases of live broadcast information that have been judged as illegal or in violation of regulations in history, as well as the corresponding penalty results and reasons. We build a supervised fine-tuning dataset based on live broadcast information cases, including positive and negative examples, and label each case. We use the supervised fine-tuning dataset to train the preset base model and adjust the parameters of the model to adapt to specific illegal identification tasks. Through multiple iterative training, the model gradually learns the characteristics of live broadcast illegal information and improves its recognition accuracy.

[0074] Furthermore, in order to improve the quality of the supervised fine-tuning dataset, the acquisition of historical penalty cases and the construction of a supervised fine-tuning dataset based on the historical penalty cases include: acquiring historical penalty cases, and cleaning the historical penalty cases to obtain cleaned penalty cases; disambiguating the cleaned penalty cases to obtain disambiguated penalty cases, and constructing a supervised fine-tuning dataset based on the disambiguated penalty cases. Among them, the cleaned penalty cases may refer to cases obtained by preprocessing historical penalty cases to remove duplications, errors or irrelevant cases. The disambiguated penalty cases may refer to cases obtained by further processing the cleaned penalty cases to resolve ambiguous or fuzzy information in the cases.

[0075] In the specific implementation, historical penalty cases are collected from various channels, including regulatory agencies, live broadcast platforms, etc. The collected cases are pre-processed to remove duplicate cases, erroneous information, incomplete cases, etc., to ensure the accuracy and consistency of case data. The cleaned penalty cases are further analyzed to resolve ambiguous or vague information in the cases (such as clarification of case descriptions, clarification of violation types, etc.), and obtain the penalty cases after disambiguation.

[0076] Step S20: extracting key video frames corresponding to the segmented voice text from the video screen of the live broadcast information to be identified, and performing character recognition on the key video frames to obtain video frame text.

[0077] It is understood that a key video frame may refer to a main frame extracted from a video image that can represent the video content in a corresponding time period. Optical Character Recognition (OCR) may refer to the process of extracting text information from an image using optical character recognition technology. Video frame text may refer to text content extracted from a key video frame using OCR technology.

[0078] In a specific implementation, according to the timestamp information of the segmented speech text, the key video frames of the corresponding time period are intercepted from the video screen. The selection of key video frames can be based on strategies such as screen changes, face detection, object recognition, etc., which are not limited in this embodiment. The OCR technology is used to extract text information from the extracted key video frames to obtain the video frame text.

[0079] Step S30: Generate a violation identification result of the live broadcast information to be identified based on the key video frame, the segmented voice text and the video frame text through a multimodal large model.

[0080] It should be understood that a multimodal large model may refer to a deep learning model that can process multiple types of data (such as images, text, etc.) to generate violation identification results.

[0081] In the specific implementation, key video frames, segmented speech texts and video frame texts are used as inputs and sent to the multimodal large model for processing. The multimodal large model can comprehensively consider multiple types of information such as images and texts to generate violation identification results.

[0082] This embodiment uses a multimodal large model to identify live broadcast violation information, while taking into account image information, voice information, and text information, thereby improving the accuracy and recall rate of live broadcast violation information identification, and further improving user experience.

[0083] Reference Figure 2 , Figure 2This is a flow chart of the second embodiment of the live broadcast violation information identification method of this application, based on the above Figure 1 The first embodiment shown proposes a second embodiment of the live broadcast violation information identification method of the present application.

[0084] In the second embodiment, the step S20 includes:

[0085] Step S201: Divide the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice texts.

[0086] It should be understood that in order to facilitate the subsequent detailed analysis of the video content, in this embodiment, the video screen of the live broadcast information to be identified is first divided into video segments corresponding to the segmented voice text, and then the video segments are frame processed to obtain the key video frames corresponding to the segmented voice text, and then the key video frames are character recognized by the image recognition model to obtain the video frame text. Among them, the video screen can refer to the video part of the live broadcast information, which is used to display the live broadcast scene, the image of the anchor, the product display and other content. The video segment can refer to the video segment corresponding to the voice text segment cut out from the video screen according to the timestamp information of the segmented voice text.

[0087] In a specific implementation, dividing the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice text may be to obtain the start time and end time of the segmented voice text; dividing the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice text according to the start time and the end time. For example, the video is intercepted into segments according to the start_time and end_time of the segmented voice text text, and the segmented voice text corresponding to the video segment is text.

[0088] Step S202: performing frame processing on the video segment to obtain key video frames corresponding to the segmented speech text.

[0089] It can be understood that in order to obtain key video frames that can represent the main content of the video segment for subsequent character recognition or image processing, in this embodiment, frame processing is performed on the video segment to obtain key video frames corresponding to the segmented speech text.

[0090] Furthermore, in order to improve the quality of key video frames, the step S202 includes: extracting frames from the video segment to obtain a video frame set corresponding to the video segment; and removing duplicates from the video frame set to obtain key video frames corresponding to the segmented speech text. The video frame set may refer to a set of multiple video frames extracted from the video segment for subsequent processing and analysis.

[0091] In the specific implementation, multiple video frames are extracted from the video segment according to certain strategies (such as fixed time interval, key frame extraction, etc.) to form a video frame set. The video frame set is processed using an image semantic model or other deduplication strategies to select video frames that can represent the main content of the video segment as key video frames.

[0092] Furthermore, in order to improve the deduplication effect of video frames, the video frame set is deduplicated to obtain key video frames corresponding to the segmented speech text, including: identifying image features of each video frame in the video frame set through a preset image semantic model; deduplicating the video frame set based on the image features to obtain key video frames corresponding to the segmented speech text.

[0093] In the specific implementation, the pre-trained image semantic model is used to perform image feature recognition on each video frame in the video frame set to extract the image features of the video frame. The image features of the video frame are deduplicated and the video frames that can represent the main content of the video segment are selected as key video frames.

[0094] Step S203: Perform character recognition on the key video frame through an image recognition model to obtain video frame text.

[0095] It should be noted that the image recognition model may refer to a machine learning model that can recognize text, objects and other information in an image, for example, an OCR model for character recognition.

[0096] In a specific implementation, an OCR model is used to perform character recognition on key video frames, extract text content in the video frames, and obtain video frame text.

[0097] This embodiment first divides the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice texts, then performs frame processing on the video segments to obtain key video frames corresponding to the segmented voice texts, and then performs character recognition on the key video frames through an image recognition model to obtain video frame texts, thereby improving the accuracy of the key video frames, and further improving the accuracy and recall rate of identifying live broadcast violation information.

[0098] Reference Figure 3 , Figure 3 This is a flow chart of the third embodiment of the live broadcast violation information identification method of the present application. Based on the above embodiments, the third embodiment of the live broadcast violation information identification method of the present application is proposed.

[0099] In the third embodiment, the step S30 includes:

[0100] Step S301: Generate prompt information of a multimodal large model according to the key video frame, the segmented speech text and the video frame text.

[0101] It should be understood that in order to enable the multimodal large model to process multiple types of data such as text and images at the same time, the accuracy and robustness of violation identification are improved. This embodiment first generates prompt information for the multimodal large model based on key video frames, segmented voice texts, and video frame texts, and then inputs the prompt information into the multimodal large model, and generates violation identification results for the live broadcast information to be identified through the multimodal large model. Among them, the prompt information may refer to input information designed to guide the multimodal large model to perform specific tasks, which may include data of multiple modes (such as text, images, etc.) and be organized in a certain format.

[0102] In the specific implementation, key video frames, segmented voice texts and video frame texts are organized according to a certain format to generate prompt information of a multimodal large model.

[0103] Furthermore, in order to improve the accuracy of the prompt information, the step S301 includes: generating a violation identification task description based on the live broadcast information to be identified; splicing the key video frame, the segmented voice text, the video frame text and the violation identification task description to obtain the prompt information of the multimodal large model. Among them, the violation identification task description may refer to a general description of the live broadcast information to be identified, clarifying the goals and requirements of violation identification. For example, the violation identification task description may include descriptions such as background description, workflow, and output requirements, which are not limited in this embodiment.

[0104] In the specific implementation, the live broadcast information to be identified is analyzed to generate a description of the violation identification task. The key video frames, segmented voice texts, video frame texts, and violation identification task descriptions are spliced ​​to obtain the prompt information of the multimodal large model.

[0105] Furthermore, in order to make the violation identification work more reliable and improve the accuracy and authority of identification, the key video frames, the segmented voice texts, the video frame texts and the violation identification task descriptions are spliced ​​to obtain prompt information of a multimodal large model, including: matching the corresponding target regulation library based on the live broadcast information to be identified; splicing the key video frames, the segmented voice texts, the video frame texts, the violation identification task descriptions and the target regulation library to obtain prompt information of a multimodal large model.

[0106] In the specific implementation, according to the type, content and other characteristics of the live broadcast information to be identified, relevant laws, regulations, policies and other information are screened out from the regulatory database as the basis for violation identification. The key video frames, segmented voice texts, video frame texts, violation identification task descriptions and target regulatory databases are spliced ​​according to certain rules to generate prompt information of a multimodal large model.

[0107] For ease of understanding, the following examples are given, but are not intended to limit the present application. As an example, the format of the prompt information may be: prompt: {background description}, regulatory library: {regulatory library list}, workflow: {workflow}, key video frame: {image}, segmented speech text: {text}, video frame text: {text}, output requirement: {output requirement}.

[0108] Step S302: input the prompt information into the multimodal large model, and generate a violation identification result of the live broadcast information to be identified through the multimodal large model.

[0109] It can be understood that the generated prompt information is input into the multimodal large model for processing, and the multimodal large model performs comprehensive analysis and judgment based on the multiple modal data in the prompt information to generate a violation identification result for the live broadcast information to be identified.

[0110] Furthermore, in order to make the violation identification result more credible, the step S302 includes: inputting the prompt information into the multimodal large model to obtain the penalty result and reasoning reason output by the multimodal large model; parsing the penalty result and the reasoning reason to obtain the violation identification result of the live broadcast information to be identified. Among them, the penalty result may refer to the result output by the multimodal large model after the input prompt information is identified as a violation, indicating whether the live broadcast information is in violation. The reasoning reason may refer to the basis or logic for explaining the penalty result given by the multimodal large model when outputting the penalty result.

[0111] In the specific implementation, the generated prompt information is passed as input to the multimodal large model. The multimodal large model conducts comprehensive analysis and processing based on the multiple modal data (such as text, images, etc.) in the prompt information, and finally outputs the penalty results and reasoning reasons. The penalty results and reasoning reasons output by the multimodal large model are parsed and processed, and converted into specific violation identification results according to business needs or rule settings. This violation identification result can be a simple binary classification (such as violation / no violation) or more detailed information such as the type or degree of violation.

[0112] This embodiment first generates prompt information of the multimodal large model according to key video frames, segmented voice texts and video frame texts, then inputs the prompt information into the multimodal large model, and generates violation identification results of the live broadcast information to be identified through the multimodal large model, thereby enabling the multimodal large model to process multiple types of data such as text and image at the same time, thereby improving the accuracy and robustness of violation identification.

[0113] For ease of understanding, refer to Figure 4 This invention is provided for illustration, but is not intended to limit the present application. Figure 4This is a specific flow chart of an embodiment of a method for identifying live broadcast violation information in this application. Figure 4 In the example, it is assumed that the live broadcast information to be identified is a live broadcast advertisement. Live broadcast advertisements have three modal data, namely, images, voice, and text (text in video images). In this embodiment, the three modal data of live broadcast advertisements are converted into a data format that can be identified by the image-text model. The data format is converted in the following three steps.

[0114] 1. Convert the voice data. Use the ASR model to convert the voice of the video into text, and divide the voice text into sentences according to the strategy. The data format is (start_time, end_time, ad_text), where ad_text is the advertisement text of the sentence, start_time is the start time of the advertisement text, and end_time is the end time of the advertisement text.

[0115] 2. Process the video based on the sentence information of the speech text. According to the start_time and end_time of the advertisement text ad_text, the video is cut into paragraphs. The advertisement text corresponding to the video paragraph is ad_text. The strategy is used to extract frames from the video paragraph to obtain the video frame set corresponding to the video paragraph. Then, the image semantic model is used to deduplicate the video frame set to obtain N key video frames.

[0116] 3. Based on the deduplicated video frames, use the OCR model to extract the text in the video frames and obtain the video frame text. Finally, splice the video frames, ASR voice text, and OCR text to generate prompt information in the specified format, and use the image-text multimodal large model for illegal identification.

[0117] For ease of understanding, the following examples are given, but are not intended to limit the present application. As an example, assuming that the live broadcast information to be identified is a live broadcast advertisement, the steps for identifying illegal live broadcast advertisements are as follows:

[0118] 1. For the live advertisement video A to be identified, use the tool to separate the video image and voice to generate video image B and voice C.

[0119] 2. Process the speech C and use the ASR (automatic speech recognition) model to convert the speech ASR_C into text ASR_C. Segment the sentence according to the strategy to generate a text sentence set (start_time, end_time, text)*n, where (start_time, end_time, text) is a segment of the ASR speech, including n segments in total, and each segment information is recorded as ASR_C_SEG.

[0120] 3. According to the ASR speech segmentation information in step 2, first segment the video screen B to generate a video screen segmentation set (start_time, end_time, image_list)*n, where (start_time, end_time, image_list) is a segment of the video screen B, including a total of n segments, each segment is recorded as B_SEG. Next, deduplicate each B_SEG to generate key frame segmentation information, recorded as B_SEG_DUP.

[0121] 4. Perform OCR to extract text from image_list in each B_SEG_DUP in step 3. The extracted OCR text is recorded as IMG_OCR.

[0122] 5. Finally, ASR_C_SEG and IMG_OCR are stitched together for each image in image_list in B_SEG_DUP and input into the multimodal large model for illegal identification.

[0123] 6. Output the results of illegal identification.

[0124] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the live broadcast violation information identification method of the present application. More simple transformations based on this technical concept are all within the protection scope of the present application.

[0125] This application also provides a live broadcast violation information identification device, please refer to Figure 5 , the live broadcast violation information identification device comprises:

[0126] The speech processing module 10 is used to convert the live speech of the live broadcast information to be recognized into speech text, and segment the speech text to obtain segmented speech text;

[0127] The video frame processing module 20 is used to extract the key video frame corresponding to the segmented voice text from the video screen of the live broadcast information to be identified, and perform character recognition on the key video frame to obtain the video frame text;

[0128] The violation identification module 30 is used to generate a violation identification result of the live broadcast information to be identified based on the key video frame, the segmented voice text and the video frame text through a multimodal large model.

[0129] The live broadcast violation information identification device provided by the present application adopts the live broadcast violation information identification method in the above-mentioned embodiment, which can solve the technical problem that the related live broadcast violation information identification method manually monitors the violation information in the live broadcast, thus having the defects of low accuracy and low recall rate. Compared with the prior art, the beneficial effects of the live broadcast violation information identification device provided by the present application are the same as the beneficial effects of the live broadcast violation information identification method provided by the above-mentioned embodiment, and the other technical features in the live broadcast violation information identification device are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0130] The present application provides a live broadcast violation information identification device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the live broadcast violation information identification method in the above-mentioned embodiment one.

[0131] Reference below Figure 6 , which shows a schematic diagram of the structure of a live broadcast violation information identification device suitable for implementing the embodiment of the present application. The live broadcast violation information identification device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions, tablet computers), PMPs (Portable Media Players, portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The live broadcast violation information identification device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0132] like Figure 6As shown, the live broadcast violation information identification device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 to the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the live broadcast violation information identification device are also stored. The processing device 1001, the ROM 1002 and the RAM 1004 are connected to each other through the bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the live broadcast violation information identification device to communicate with other devices wirelessly or by wire to exchange data. Although the live broadcast violation information identification device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0133] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0134] The live broadcast violation information identification device provided by the present application adopts the live broadcast violation information identification method in the above embodiment, which can solve the technical problem that the related live broadcast violation information identification method manually monitors the violation information in the live broadcast, thus having the defects of low accuracy and low recall rate. Compared with the prior art, the beneficial effects of the live broadcast violation information identification device provided by the present application are the same as the beneficial effects of the live broadcast violation information identification method provided by the above embodiment, and the other technical features in the live broadcast violation information identification device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0135] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0136] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0137] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the live broadcast violation information identification method in the above-mentioned embodiment.

[0138] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory) or flash memory, optical fiber, CD-ROM (CD-Read Only Memory, portable compact disk read-only memory), optical storage device, magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0139] The above-mentioned computer-readable storage medium may be included in the live broadcast violation information identification device; or it may exist independently without being assembled into the live broadcast violation information identification device.

[0140] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the live broadcast violation information identification device, the live broadcast violation information identification device executes the above-mentioned live broadcast violation information identification method.

[0141] The computer program code for performing the operation of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on the remote computer, or completely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or it can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).

[0142] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0143] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0144] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned live broadcast violation information identification method, and can solve the technical problem that the related live broadcast violation information identification method manually monitors the violation information in the live broadcast, thus having the defects of low accuracy and low recall rate. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the live broadcast violation information identification method provided by the above-mentioned embodiment, and will not be repeated here.

[0145] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the live broadcast violation information identification method as described above.

[0146] The computer program product provided by the present application can solve the technical problem that the related live broadcast violation information identification method manually monitors the violation information in the live broadcast, which has the defects of low accuracy and low recall rate. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the live broadcast violation information identification method provided by the above embodiment, which will not be repeated here.

[0147] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

[0148] The present application discloses A1, a method for identifying live broadcast violation information, the method for identifying live broadcast violation information comprising:

[0149] Converting the live speech of the live broadcast information to be recognized into speech text, and segmenting the speech text to obtain segmented speech text;

[0150] Extracting key video frames corresponding to the segmented speech text from the video screen of the live broadcast information to be identified, and performing character recognition on the key video frames to obtain video frame text;

[0151] Based on the key video frames, the segmented voice texts and the video frame texts, a violation identification result of the live broadcast information to be identified is generated through a multimodal large model.

[0152] A2. The method for identifying live broadcast violation information as described in A1, wherein the key video frame corresponding to the segmented voice text is extracted from the video screen of the live broadcast information to be identified, and character recognition is performed on the key video frame to obtain the video frame text, including:

[0153] Dividing the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice texts;

[0154] Performing frame processing on the video segments to obtain key video frames corresponding to the segmented speech texts;

[0155] Character recognition is performed on the key video frame through an image recognition model to obtain video frame text.

[0156] A3. The method for identifying live broadcast violation information as described in A2, wherein the frame processing of the video segment is performed to obtain the key video frame corresponding to the segmented speech text, including:

[0157] Extracting frames from the video segment to obtain a set of video frames corresponding to the video segment;

[0158] The video frame set is deduplicated to obtain key video frames corresponding to the segmented speech text.

[0159] A4. The method for identifying live broadcast violation information as described in A3, wherein the video frame set is deduplicated to obtain key video frames corresponding to the segmented speech text, including:

[0160] Identify the image features of each video frame in the video frame set by using a preset image semantic model;

[0161] The video frame set is deduplicated based on the image features to obtain key video frames corresponding to the segmented speech text.

[0162] A5. The method for identifying live broadcast violation information as described in A2, wherein the video screen of the live broadcast information to be identified is divided into video segments corresponding to the segmented voice texts, including:

[0163] Obtain the start time and end time of the segmented speech text;

[0164] The video screen of the live broadcast information to be identified is divided into video segments corresponding to the segmented voice texts according to the start time and the end time.

[0165] A6. The live broadcast violation information identification method as described in any one of A1 to A5, wherein the violation identification result of the live broadcast information to be identified is generated by a multimodal large model based on the key video frame, the segmented voice text, and the video frame text, including:

[0166] Generate prompt information of a multimodal large model according to the key video frame, the segmented voice text and the video frame text;

[0167] The prompt information is input into the multimodal large model, and a violation identification result of the live broadcast information to be identified is generated through the multimodal large model.

[0168] A7. The method for identifying live broadcast violation information as described in A6, wherein the prompt information of the multimodal large model is generated according to the key video frame, the segmented voice text and the video frame text, including:

[0169] Generate a violation identification task description according to the live broadcast information to be identified;

[0170] The key video frame, the segmented speech text, the video frame text and the violation identification task description are spliced ​​to obtain prompt information of a multimodal large model.

[0171] A8. The live broadcast violation information identification method as described in A7, wherein the key video frame, the segmented voice text, the video frame text, and the violation identification task description are spliced ​​to obtain the prompt information of the multimodal large model, including:

[0172] Matching the corresponding target regulation library based on the live broadcast information to be identified;

[0173] The key video frames, the segmented speech texts, the video frame texts, the violation identification task descriptions and the target regulations library are spliced ​​to obtain prompt information of a multimodal large model.

[0174] A9. The method for identifying live broadcast violation information as described in A6, wherein the prompt information is input into the multimodal large model, and a violation identification result of the live broadcast information to be identified is generated by the multimodal large model, including:

[0175] Inputting the prompt information into the multimodal large model to obtain the penalty result and reasoning reason output by the multimodal large model;

[0176] The penalty result and the reasoning reason are analyzed to obtain a violation identification result of the live broadcast information to be identified.

[0177] A10. The method for identifying live broadcast violation information as described in any one of A1 to A5, wherein the live broadcast voice of the live broadcast information to be identified is converted into voice text, and the voice text is segmented, and before obtaining the segmented voice text, the method further comprises:

[0178] Obtain historical penalty cases, and construct a supervised fine-tuning dataset based on the historical penalty cases;

[0179] The preset base model is fine-tuned based on the supervised fine-tuning dataset to obtain a multimodal large model.

[0180] A11. The method for identifying live broadcast violation information as described in A10, wherein the step of obtaining historical penalty cases and constructing a supervised fine-tuning dataset based on the historical penalty cases includes:

[0181] Obtaining historical penalty cases, and cleaning the historical penalty cases to obtain cleaned penalty cases;

[0182] The cleaned penalty cases are disambiguated to obtain disambiguated penalty cases, and a supervised fine-tuning dataset is constructed based on the disambiguated penalty cases.

[0183] The present application also discloses B12, a live broadcast violation information identification device, the live broadcast violation information identification device comprising:

[0184] A voice processing module, used for converting the live voice of the live information to be recognized into voice text, and segmenting the voice text to obtain segmented voice text;

[0185] A video frame processing module is used to extract key video frames corresponding to the segmented voice text from the video screen of the live broadcast information to be identified, and perform character recognition on the key video frames to obtain video frame text;

[0186] A violation identification module is used to generate a violation identification result of the live broadcast information to be identified based on the key video frame, the segmented voice text and the video frame text through a multimodal large model.

[0187] B13. In the live broadcast violation information identification device as described in B12, the video frame processing module is further used to divide the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice text; perform frame processing on the video segments to obtain key video frames corresponding to the segmented voice text; and perform character recognition on the key video frames through an image recognition model to obtain video frame text.

[0188] B14. In the live broadcast violation information identification device as described in B13, the video frame processing module is further used to extract frames from the video segment to obtain a set of video frames corresponding to the video segment; and to deduplicate the video frame set to obtain key video frames corresponding to the segmented voice text.

[0189] B15. In the live broadcast violation information identification device as described in B14, the video frame processing module is further used to identify the image features of each video frame in the video frame set through a preset image semantic model; deduplicate the video frame set based on the image features to obtain key video frames corresponding to the segmented speech text.

[0190] B16. In the live broadcast violation information identification device as described in B13, the video frame processing module is also used to obtain the start time and end time of the segmented voice text; and divide the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice text according to the start time and the end time.

[0191] B17. In the live broadcast violation information identification device as described in any one of B12 to B16, the violation identification module is further used to generate prompt information of the multimodal large model according to the key video frame, the segmented voice text and the video frame text; input the prompt information into the multimodal large model, and generate a violation identification result of the live broadcast information to be identified through the multimodal large model.

[0192] The present application also discloses C18, a live broadcast violation information identification device, which includes: a memory, a processor, and a live broadcast violation information identification program stored in the memory and executable on the processor, and when the live broadcast violation information identification program is executed by the processor, the live broadcast violation information identification method as described above is implemented.

[0193] The present application also discloses D19, a storage medium, on which a live broadcast violation information identification program is stored, and when the live broadcast violation information identification program is executed by a processor, the live broadcast violation information identification method as described above is implemented.

[0194] The present application also discloses E20, a computer program product, which includes a live broadcast violation information identification program, and when the live broadcast violation information identification program is executed by a processor, it implements the live broadcast violation information identification method described above.

Claims

1. A method for identifying live broadcast violation information, characterized in that: The live broadcast violation information identification method comprises: Converting the live speech of the live broadcast information to be recognized into speech text, and segmenting the speech text to obtain segmented speech text; Extracting key video frames corresponding to the segmented speech text from the video screen of the live broadcast information to be identified, and performing character recognition on the key video frames to obtain video frame text; Based on the key video frames, the segmented voice texts and the video frame texts, a violation identification result of the live broadcast information to be identified is generated through a multimodal large model.

2. The method for identifying live broadcast violation information according to claim 1, characterized in that: The step of extracting key video frames corresponding to the segmented speech text from the video screen of the live broadcast information to be identified, and performing character recognition on the key video frames to obtain video frame texts includes: Dividing the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice texts; Performing frame processing on the video segments to obtain key video frames corresponding to the segmented speech texts; Character recognition is performed on the key video frame through an image recognition model to obtain video frame text.

3. The method for identifying live broadcast violation information according to claim 2, characterized in that: The performing frame processing on the video segment to obtain key video frames corresponding to the segmented speech text includes: Extracting frames from the video segment to obtain a set of video frames corresponding to the video segment; The video frame set is deduplicated to obtain key video frames corresponding to the segmented speech text.

4. The method for identifying live broadcast violation information according to claim 3, characterized in that: The step of removing duplicates from the video frame set to obtain key video frames corresponding to the segmented speech text includes: Identify the image features of each video frame in the video frame set by using a preset image semantic model; The video frame set is deduplicated based on the image features to obtain key video frames corresponding to the segmented speech text.

5. The method for identifying live broadcast violation information according to claim 2, characterized in that: The step of dividing the video screen of the live broadcast information to be identified into video segments corresponding to the segmented voice texts includes: Obtain the start time and end time of the segmented speech text; The video screen of the live broadcast information to be identified is divided into video segments corresponding to the segmented voice texts according to the start time and the end time.

6. The method for identifying live broadcast violation information according to any one of claims 1 to 5, characterized in that: The generating of the violation identification result of the live broadcast information to be identified through a multimodal large model based on the key video frame, the segmented voice text and the video frame text includes: Generate prompt information of a multimodal large model according to the key video frame, the segmented voice text and the video frame text; The prompt information is input into the multimodal large model, and a violation identification result of the live broadcast information to be identified is generated through the multimodal large model.

7. A device for identifying live broadcast violation information, characterized in that: The live broadcast violation information identification device comprises: A voice processing module, used for converting the live voice of the live information to be recognized into voice text, and segmenting the voice text to obtain segmented voice text; A video frame processing module is used to extract key video frames corresponding to the segmented voice text from the video screen of the live broadcast information to be identified, and perform character recognition on the key video frames to obtain video frame text; A violation identification module is used to generate a violation identification result of the live broadcast information to be identified based on the key video frame, the segmented voice text and the video frame text through a multimodal large model.

8. A device for identifying live broadcast violation information, characterized in that: The live broadcast violation information identification device includes: a memory, a processor, and a live broadcast violation information identification program stored in the memory and executable on the processor. When the live broadcast violation information identification program is executed by the processor, the live broadcast violation information identification method according to any one of claims 1 to 6 is implemented.

9. A storage medium, characterized in that: The storage medium stores a live broadcast violation information identification program, and when the live broadcast violation information identification program is executed by the processor, the live broadcast violation information identification method according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that The computer program product includes a live broadcast violation information identification program, and when the live broadcast violation information identification program is executed by a processor, it implements the live broadcast violation information identification method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Live broadcast video stream analysis method and system based on multi-modal large model

    CN120475195A

  • A live video stream analysis method and system based on a multi-modal large model

    CN120475195B

  • Massive network live broadcast batch data acquisition method and system

    CN120769077A