A method, device, medium and equipment for detecting video content risks

By dividing the multimodal detection data of real-time video stream data into different information density levels for processing, first detection and integration, the problem of insufficient risk identification of single modal data is solved, fast and accurate risk detection is achieved, and computing complexity and resource consumption are reduced.

CN119274119BActive Publication Date: 2025-07-04ANT ZHIXIN HANGZHOU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411814529.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-07-04
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

The existing video content risk detection technology relies on single modal data, making it difficult to accurately identify potential risk information in live videos, resulting in misjudgment or misjudgment. In addition, the multimodal data detection scheme has high computational complexity, large amount of repeated calculations, and low risk detection efficiency.

Method used

By extracting the multimodal detection data of real-time video stream data, it is divided into first information density data and second information density data, the first information density data is detected first and fused into the second information density data, content decisions are made, repeated calculations are reduced, and risk detection efficiency is improved.

Benefits of technology

It realizes rapid discovery of risk content in real-time video streaming data, reduces computing complexity, shortens detection time span, saves computing resources, and improves risk detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274119B_ABST
    Figure CN119274119B_ABST
Patent Text Reader

Abstract

The embodiments of this specification disclose a method, device, medium, and equipment for video content risk detection. First, real-time video stream data is obtained, and corresponding multi-modal detection data is extracted. The multi-modal detection data includes first information density data and second information density data whose information density is greater than that of the first information density data. Then, the first information density data is detected to determine the risk detection result, and the first information density data and the risk detection result are fused into the second information density data, and content decision-making is performed on the fused second information density data to obtain the content risk decision result. This technical solution processes multi-modal data hierarchically by information density, which can effectively improve the detection efficiency of risk content in real-time video stream data and shorten the risk detection time limit. At the same time, it can reduce the repeated calculation amount of high-information density data in the content decision-making process and improve the content decision-making efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of video risk detection, and particularly to a method, device, medium, and equipment for detecting risks in video content. Background Art

[0002] With the rapid development of live broadcast services and real-time video stream services, due to the complexity and real-time requirements of live videos and real-time video streams, higher challenges are posed in content review and risk control technologies. Traditional content detection systems generally identify risks based on single-modal data, such as identifying risks through frame images, audio slices, or keyword analysis. However, due to the diversity of information presentation forms and the complexity of risk factors in live content, it is difficult to accurately identify potential risk information relying solely on single-modal detection, and situations of missed or misjudged detections often occur due to insufficient information density.

[0003] Currently, in related technologies, a detection method using multi-modal data is adopted to improve the risk identification ability. However, since this technical solution uses multi-modal data for risk identification, the data complexity is relatively high, resulting in possible duplicate data among multi-modal data, leading to a large amount of duplicate calculations, wasting computing resources. At the same time, the time span from the occurrence of a risk to its detection is relatively long, and the risk detection efficiency is relatively low. Summary of the Invention

[0004] In the first aspect of the embodiments of the specification, a method for detecting risks in video content is provided. This method can reduce the data complexity of risk identification, reduce the amount of duplicate calculations, save computing resources, shorten the time span from the occurrence of a risk to its detection, and improve the risk detection efficiency. The method includes:

[0005] Obtain real-time video stream data, and extract multi-modal detection data corresponding to the real-time video stream data. The multi-modal detection data includes first information density data and second information density data, and the information density of the second information density data is greater than that of the first information density data;

[0006] Detect the first information density data to determine a risk detection result;

[0007] Fuse the first information density data and the risk detection result into the second information density data, and perform content decision-making on the fused second information density data to obtain a content risk decision result corresponding to the real-time video stream data.

[0008] Further, in some embodiments, before fusing the first information density data and the risk detection result into the second information density data, the method further includes: sequentially inputting the risk detection results corresponding to the first information density data into the content decision system according to the output order to obtain preliminary risk decision results corresponding to the first information density data; if it is determined that the content risk level of any one of the preliminary risk decision results is greater than or equal to the content risk level threshold, directly using the preliminary risk decision result as the content risk decision result corresponding to the real-time video stream data; if it is determined that the content risk levels of all the preliminary risk decision results are less than the content risk level threshold, fusing the first information density data and the risk detection result into the second information density data.

[0009] Further, in some embodiments, the making a content decision on the fused second information density data to obtain the content risk decision result corresponding to the real-time video stream data includes: inputting the fused second information density data into the content decision system, so that the content decision system makes a content decision on the second information density data based on the risk detection result corresponding to the first information density data to obtain an intermediate risk decision result of the fused second information density data; determining the content risk decision result corresponding to the real-time video stream data according to the preliminary risk decision results and the intermediate risk decision result.

[0010] Further, in some embodiments, the determining the content risk decision result corresponding to the real-time video stream data according to the preliminary risk decision results and the intermediate risk decision result includes: obtaining the multi-modal risk detection weights corresponding to the first information density data and the second information density data; fusing the preliminary risk decision results and the intermediate risk decision result according to the multi-modal risk detection weights to determine the content risk decision result corresponding to the real-time video stream data.

[0011] Further, in some embodiments, the obtaining the multi-modal risk detection weights corresponding to the first information density data and the second information density data includes: obtaining the initial risk detection weights corresponding to the first information density data and the second information density data; updating the initial risk detection weights according to the content risk decision result at the historical moment associated with the real-time video stream data to obtain the multi-modal risk detection weights, so that the updated multi-modal risk detection weights are adapted to the video content of the real-time video stream data.

[0012] Further, in some embodiments, the extraction of the multimodal detection data corresponding to the real-time video stream data includes: obtaining a preset detection extraction time, where the detection extraction time includes a first extraction duration and a second extraction duration, and the second extraction duration is greater than the first extraction duration; extracting the first information density data by the first extraction duration in combination with the timestamp information of the real-time video stream data; and extracting the second information density data by the second extraction duration in combination with the timestamp information of the real-time video stream data.

[0013] Further, in some embodiments, the first extraction duration includes a video frame capture interval and an audio slice extraction duration; the extraction of the first information density data by the first extraction duration in combination with the timestamp information of the real-time video stream data includes: performing frame extraction processing on the real-time video stream data based on the video frame capture interval to obtain frame-captured images, and determining the timestamp information corresponding to the frame-captured images according to the timestamp information of the real-time video stream data; performing audio extraction on the real-time video stream data based on the audio slice extraction duration to obtain audio slice data, and determining the timestamp information corresponding to the audio slice data according to the timestamp information of the real-time video stream data; performing speech recognition on the audio slice data to determine speech recognition text data, and determining the timestamp information corresponding to the speech recognition text data according to the timestamp information corresponding to the audio slice data; and using the frame-captured images, the audio slice data, the speech recognition text data, and the corresponding timestamp information as the first information density data.

[0014] Further, in some embodiments, the second extraction duration includes a video slice extraction duration; the extraction of the second information density data by the second extraction duration in combination with the timestamp information of the real-time video stream data includes: performing video extraction on the real-time video stream data based on the video slice extraction duration to obtain video slice data, and determining the absolute start timestamp and the absolute end timestamp corresponding to the video slice data according to the timestamp information of the real-time video stream data.

[0015] Further, in some embodiments, the first information density data at least includes frame - intercepted images, audio slice data, and speech recognition text data; the detecting the first information density data to determine a risk detection result includes: performing feature analysis on the frame - intercepted images to obtain frame - intercepted image features, and determining a frame - intercepted detection result through the frame - intercepted image features; performing audio analysis on the audio slice data to obtain pitch content features, and determining an audio detection result through the pitch content features; performing content analysis on the speech recognition text data to obtain keyword features, and determining a text detection result through the keyword features; storing the frame - intercepted detection result, the audio detection result, and the text detection result as the risk detection result corresponding to the first information density data.

[0016] Further, in some embodiments, the fusing the first information density data and the risk detection result into the second information density data includes: obtaining the absolute start timestamp and the absolute end timestamp corresponding to the second information density data; matching the first information density data within the corresponding timestamp range based on the absolute start timestamp and the absolute end timestamp; integrating the matched first information density data and the corresponding risk detection result into the detection information associated with the second information density data according to the timestamp information of the matched first information density data, to obtain the fused second information density data.

[0017] In the second aspect of the embodiments of the specification, a video content risk detection device is further proposed, including:

[0018] A multimodal detection data extraction module, configured to obtain real - time video stream data and extract the multimodal detection data corresponding to the real - time video stream data, where the multimodal detection data includes first information density data and second information density data, and the information density of the second information density data is greater than that of the first information density data;

[0019] A risk detection module, configured to detect the first information density data to determine a risk detection result;

[0020] A content risk decision - making module, configured to fuse the first information density data and the risk detection result into the second information density data, and perform content decision - making on the fused second information density data to obtain the content risk decision result corresponding to the real - time video stream data.

[0021] In the third aspect of the embodiments of the specification, a computer program product is further provided. The computer program product stores at least one instruction, and the at least one instruction is adapted to be loaded and executed by a processor to perform the method steps in the first aspect.

[0022] In the fourth aspect of the embodiments of this specification, a storage medium is further provided. The storage medium stores a computer program, and the computer program is adapted to be loaded and executed by a processor to perform the steps of the method in the first aspect.

[0023] In the fifth aspect of the embodiments of this specification, an electronic device is further provided, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the method in the first aspect.

[0024] In the embodiments of this specification, by extracting multi-modal detection data corresponding to real-time video stream data, and classifying the multi-modal detection data into first information density data and second information density data according to information density, and the information density of the second information density data is greater than that of the first information density data; detecting the first information density data to determine a risk detection result, and fusing the first information density data and the risk detection result into the second information density data, and performing content decision on the fused second information density data to obtain a content risk decision result corresponding to the real-time video stream data. On the one hand, by hierarchically processing the multi-modal detection data according to information density and performing prior risk decision processing on the first information density data, it is possible to quickly discover risk content in real-time video stream data, avoid the large computational pressure of high information density data, reduce the computational complexity in the overall detection process, shorten the time span from the occurrence of risk to being detected, improve the risk detection efficiency, thereby improving the response speed to risk content in real-time video stream data and ensuring network data security; on the other hand, after obtaining the risk detection result of the first information density data through prior risk decision processing, fusing the risk detection result of the first information density data with the second information density data. At this time, if there are duplicate modal data in the second information density data, the risk detection result of the first information density data can be directly used as the risk detection result of the duplicate data, avoiding a large amount of duplicate calculations caused by performing risk decisions on duplicate data, saving computational resources, and at the same time improving the risk decision efficiency of the second information density data, further improving the risk detection efficiency while ensuring the accuracy and stability of the risk detection result. Description of the Drawings

[0025] Figure 1 A schematic diagram of the system architecture showing an exemplary application environment to which a video content risk detection method and apparatus according to the embodiments of this specification can be applied.

[0026] Figure 2 A flowchart showing a video content risk detection method provided by the embodiments of this specification.

[0027] Figure 3 A flowchart showing a content risk detection based on multi-modal detection data provided by the embodiments of this specification.

[0028] Figure 4 A schematic flowchart for dynamically updating and adjusting the weights of multimodal risk detection provided by an embodiment of this specification.

[0029] Figure 5 A schematic flowchart for determining the risk detection result of the first information density data provided by an embodiment of this specification.

[0030] Figure 6 A schematic flowchart for integrating the first information density data and its risk detection result into the second information density data provided by an embodiment of this specification.

[0031] Figure 7 A schematic flowchart for realizing risk detection by hierarchically processing multimodal detection data according to information density provided by an embodiment of this specification.

[0032] Figure 8 A schematic structural diagram of a video content risk detection device provided by an embodiment of this specification.

[0033] Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of this specification. Detailed implementation manners

[0034] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0035] Figure 1 A schematic diagram showing the system architecture of an exemplary application environment of a document forgery detection method and device to which the embodiments of this specification can be applied.

[0036] As Figure 1 shown, the system architecture 100 may include one or more of the terminal devices 101, 102, and 103, the network 104, and the server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with artificial intelligence (AI) computing capabilities, including but not limited to desktop computers, portable computers, smartphones, and tablet computers, etc. It should be understood,Figure 1 The numbers of the terminal devices, networks, and servers in [it] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. For example, the server 105 can be a server cluster composed of multiple servers, etc.

[0037] The video content risk detection method provided by the embodiments of this specification is generally executed by the server 105. Correspondingly, the video content risk detection device is generally set in the server 105. However, those skilled in the art can easily understand that the video content risk detection method provided by the embodiments of this specification can also be executed by the terminal devices 101, 102, and 103. Correspondingly, the video content risk detection device can also be set in the terminal devices 101, 102, and 103. No special limitation is made in this exemplary embodiment.

[0038] Please refer to Figure 2 , which is a schematic flowchart of a video content risk detection method provided by the embodiments of this specification. In the embodiments of this specification, the video content risk detection method can be applied to terminal devices or servers. No special limitation is made in this exemplary embodiment. Taking the server executing this method as an example below, Figure 2 the following describes the process shown in detail. The video content risk detection method in the embodiments of this specification can specifically include the following steps:

[0039] Step S210: Obtain real-time video stream data and extract the multi-modal detection data corresponding to the real-time video stream data. The multi-modal detection data includes first information density data and second information density data, and the information density of the second information density data is greater than that of the first information density data;

[0040] Step S220: Detect the first information density data to determine the risk detection result;

[0041] Step S230: Integrate the first information density data and the risk detection result into the second information density data, and perform content decision on the integrated second information density data to obtain the content risk decision result corresponding to the real-time video stream data.

[0042] According to the video content risk detection method in the embodiments of the present specification, on the one hand, by hierarchically processing multi-modal detection data through information density and performing an antecedent risk decision process on the first information density data, it is possible to quickly discover risk content in real-time video stream data, avoid the large computational pressure of high information density data, reduce the computational complexity in the overall detection process, shorten the time span from the occurrence of a risk to its detection, improve the risk detection efficiency, thereby increasing the response speed to risk content in real-time video stream data and ensuring network data security; on the other hand, after obtaining the risk detection result of the first information density data through the antecedent risk decision process, the risk detection result of the first information density data is fused with the second information density data. At this time, if there are duplicate modal data in the second information density data, the risk detection result of the first information density data can be directly used as the risk detection result of the duplicate data, avoiding a large amount of duplicate calculations caused by making risk decisions on duplicate data, saving computational resources, and at the same time improving the risk decision efficiency of the second information density data. While ensuring the accuracy and stability of the risk detection result, the risk detection efficiency is further improved.

[0043] The video content risk detection method in the embodiments of the present specification will be described in detail below.

[0044] In step S210, real-time video stream data is obtained, and multi-modal detection data corresponding to the real-time video stream data is extracted. The multi-modal detection data includes first information density data and second information density data, and the information density of the second information density data is greater than that of the first information density data.

[0045] In an exemplary embodiment of the present specification, real-time video stream data refers to a continuous sequence of video frames transmitted through a network, which includes meta-information such as video signals, audio signals, and timestamps. For example, the real-time video stream data can be a live video stream or a real-time chat video in a chat tool. The type of real-time video stream data is not specifically limited in this exemplary embodiment.

[0046] Multi-modal detection data refers to detection data with different modalities extracted from real-time video stream data. For example, the multi-modal detection data can be frame images, comment information, bullet screen information, or audio slice data, speech recognition text data, video slice data. The types of multi-modal data that can be extracted from real-time video stream data are not specifically limited in this exemplary embodiment.

[0047] Multi-modal detection data can be extracted from real-time video stream data through a variety of related technologies, and this embodiment does not make special limitations on this. For example, a real-time video stream data can be frame-processed by a streaming media parser to extract interlaced frame images as intercepted frame images; an audio stream can be extracted by an audio decoder and short-time audio slices can be generated to obtain audio slice data; the audio slice data can be converted into speech recognition text data through ASR (Automatic Speech Recognition) technology, and video slice data can be extracted from the real-time video stream data by a video slice tool. It can be understood that the extracted multi-modal detection data can simultaneously carry its corresponding timestamp information, and the timestamp information can be used to mark the start and end moments of each modal data to ensure the temporal consistency of subsequent data matching.

[0048] In the embodiments of this specification, the multi-modal detection data can be divided into first information density data and second information density data according to the information density, and the information density of the second information density data is greater than that of the first information density data. For example, the multi-modal detection data can be data such as intercepted frame images, comment information, barrage information, audio slice data, speech recognition text data, video slice data, etc., and among them, the information density of the video slice data is greater than that of the intercepted frame images, comment information, barrage information, audio slice data, speech recognition text data. Therefore, data such as intercepted frame images, comment information, barrage information, audio slice data, speech recognition text data, etc. can be used as the first information density data, and the video slice data can be used as the second information density data; of course, the data belonging to the first information density data and the second information density data can also be pre-divided according to the data repetition situation. For example, the multi-modal detection data can include intercepted frame images, audio slice data, speech recognition text data, video slice data, and among them, the intercepted frame images are part of the video slice data, and the speech recognition text data is part of the audio slice data. Therefore, the intercepted frame images and the speech recognition text data can be used as the first information density data, and the audio slice data and the video slice data can be used as the second information density data. It can be understood that the division method and division content of the multi-modal detection data of the first information density data and the second information density data can be custom-set according to the actual application scenario, and this embodiment does not make special limitations on this. For the convenience of description, the embodiments of this specification will hereinafter take the first information density data as intercepted frame images, audio slice data, speech recognition text data, and the second information density data as video slice data as an example for description.

[0049] In step S220, the first information density data is detected to determine a risk detection result.

[0050] In an exemplary embodiment of this specification, the risk detection result refers to the detection result obtained by performing risk content detection by combining the content information included in the first information density data. For example, assume that the first information density data may include frame capture images, audio slice data, and speech recognition text data. Specific risk object image features can be extracted through a deep learning model (such as the Faster R-CNN network or the YOLO network model, etc.) or traditional image processing algorithms and matched with a preset risk feature library to determine the risk detection result of the frame capture image. The time domain and frequency domain features of the audio slice data (such as MFCC (Mel Frequency Cepstrum Coefficient), pitch, timbre, tone features, etc.) can be extracted through acoustic analysis techniques, and the extracted time domain and frequency domain features are compared and analyzed with sensitive audio templates to obtain the risk detection result of the audio slice data. Potential risk content in the speech recognition text data can be identified through natural language processing techniques (such as keyword matching or semantic analysis). It can be understood that only a possible detection method for performing risk content detection on the first information density data is given here. Other types of risk detection methods can also be used to implement the detection of the first information density data and determine the risk detection result. This exemplary embodiment is not limited thereto.

[0051] Risk content detection can be sequentially performed on each data in the first information density data. The specific detection order can be determined according to the extraction speed of each first information density data. For example, the frame capture image is extracted before the audio slice data and the speech recognition text data. After obtaining the frame capture image, risk detection is directly performed. That is, the first information density data is only a classification of multi-modal detection data, and the data in the first information density data are still independently detected.

[0052] It can be understood that in one implementation manner, after obtaining the risk detection result of any first information density data, the first information density data and the corresponding risk detection result can be directly input into the content decision system for content decision to determine the risk decision result of the current real-time video stream data, which can effectively improve the risk detection efficiency of the real-time video stream data and shorten the risk detection time limit. The risk detection time limit is the time taken to detect potential risk information in the content, that is, the time span from the occurrence of the risk to its detection. In another implementation manner, it is also possible to wait until the risk detection of each modal detection data included in the first information density data is completed, and then input all the first information density data and the corresponding risk detection results into the content decision system for content decision to determine the risk decision result of the current real-time video stream data, which can effectively improve the accuracy and stability of the content decision result of the real-time video stream data.

[0053] When the content decision result of the first information density data is risk-free content, the first information density data, its corresponding risk detection result, and timestamp information can be stored in a pre-set content database for subsequent calls and processing. Optionally, the content database can be configured with a timestamp index structure to accelerate data matching, or the structure of the content database is a query database (such as an SQL database).

[0054] In step S230, the first information density data and the risk detection result are fused into the second information density data, and content decision-making is performed on the fused second information density data to obtain the content risk decision result corresponding to the real-time video stream data.

[0055] In an exemplary embodiment of this specification, the core of data fusion lies in matching the timestamp ranges of the first information density data and the second information density data, and associating the risk detection result of the first information density data with the second information density data. Precise matching of timestamps is the key to data fusion, which can be achieved by establishing an index or querying a database for modal association within the time sequence range. For example, according to the timestamp range of the second information density data such as video slice data, the corresponding timestamps of the first information density data such as frame capture images, audio slice data, and speech recognition text data can be retrieved based on this timestamp range.

[0056] By hierarchically processing multi-modal detection data through information density and performing prior risk decision processing on the first information density data, it is possible to quickly discover risk content in real-time video stream data, avoid the large computational pressure of high information density data, reduce the computational complexity in the overall detection process, shorten the time span from the occurrence of a risk to its detection, improve the risk detection efficiency, thereby increasing the response speed to risk content in real-time video stream data and ensuring network data security; in addition, after obtaining the risk detection result of the first information density data through prior risk decision processing, the risk detection result of the first information density data is fused with the second information density data. At this time, if there are duplicate modal data in the second information density data, the risk detection result of the first information density data can be directly used as the risk detection result of the duplicate data, avoiding a large amount of duplicate calculations caused by performing risk decision-making on duplicate data, saving computational resources, and at the same time improving the risk decision efficiency of the second information density data, further improving the risk detection efficiency while ensuring the accuracy and stability of the risk detection result.

[0057] The content in steps S210 to S230 will be described in detail below.

[0058] In an exemplary embodiment of this specification, before fusing the first information density data and the risk detection result into the second information density data, content decision-making may be implemented based on the risk detection result of the first information density data. Refer to Figure 3 as shown, and it may be specifically implemented through the following steps:

[0059] Step S310: Input the risk detection results corresponding to each of the first information density data into the content decision-making system in the output order, and obtain the preliminary risk decision results corresponding to each of the first information density data;

[0060] Step S320: If it is determined that the content risk level of any one of the preliminary risk decision results is greater than or equal to the content risk level threshold, then directly use the preliminary risk decision result as the content risk decision result corresponding to the real-time video stream data;

[0061] Step S330: If it is determined that the content risk levels of all the preliminary risk decision results are less than the content risk level threshold, then fuse the first information density data and the risk detection result into the second information density data.

[0062] Among them, the output order refers to the order of extracting each first information density data from the real-time video stream data. For example, the first information density data may include frame images, audio slice data, and speech recognition text data. Among them, the frame image may be a video frame image extracted from the real-time video stream data at regular intervals, and the audio slice data is the continuous audio data within a certain period of time in the real-time video stream data. The speech recognition text data is the data obtained after recognizing the audio slice data. Therefore, there is a time difference in the output order of each first information density data.

[0063] The risk detection results corresponding to each first information density data and each first information density data may be input into the content decision-making system in the output order of each first information density data for risk content decision-making, and the preliminary risk decision results corresponding to each first information density data are obtained. For example, the preliminary risk decision result may be risk-free, low risk, medium risk, high risk, etc. Of course, the preliminary risk decision result may also be divided into levels 1 to 5. Level 1 corresponds to a lower content risk level, and level 5 corresponds to a higher content risk level; this exemplary embodiment does not make special limitations on the division method of the preliminary risk decision result.

[0064] When it is determined that the content risk level of any preliminary risk decision result is greater than or equal to the content risk level threshold. For example, the content risk level threshold can be medium risk or level 3. At this time, it can be considered that medium to high risk content appears in the real-time video stream data. Then, the preliminary risk decision result can be directly used as the content risk decision result corresponding to the real-time video stream data, realizing the rapid discovery of medium to high risk content through the first information density data, effectively shortening the time span from the occurrence of the risk to its detection, and improving the risk detection efficiency.

[0065] When it is determined that the content risk levels of all preliminary risk decision results are less than the content risk level threshold, it can be considered that the first information density data preliminarily detects risk content and further confirmation is required. Then, the first information density data and its corresponding risk detection results can be fused into the second information density data, so as to realize the further detection of the risk content of the real-time video stream data. Since the information density of the second information density data is relatively high, the accuracy and stability of the risk content decision result can be ensured. At the same time, the second information density data is fused with the first information density data and its corresponding risk detection results, so that when making a risk content decision on the second information density data, the detection and processing of duplicate data can be avoided, effectively reducing the amount of duplicate calculation, saving computing resources, and further improving the risk detection efficiency.

[0066] Optionally, after determining that the content risk levels of all preliminary risk decision results are less than the content risk level threshold and fusing the first information density data and its corresponding risk detection results into the second information density data, content decision can be made on the fused second information density data to obtain the content risk decision result corresponding to the real-time video stream data. Specifically, it can be achieved through the following steps:

[0067] The fused second information density data can be input into the content decision system, so that the content decision system makes content decision on the second information density data based on the risk detection results corresponding to the first information density data, obtaining the intermediate risk decision result of the fused second information density data. Then, based on the preliminary risk decision results and the intermediate risk decision result, the content risk decision result corresponding to the real-time video stream data can be determined. For example, the content risk decision result corresponding to the real-time video stream data can be obtained by weighted calculation through the weights corresponding to the preliminary risk decision results and the intermediate risk decision result; of course, the preliminary risk decision results and the intermediate risk decision result can also be input into a pre-set standardized risk scoring function to comprehensively determine the content risk decision result corresponding to the real-time video stream data. In this exemplary embodiment, there is no special limitation on the specific method of determining the content risk decision result corresponding to the real-time video stream data based on the preliminary risk decision results and the intermediate risk decision result.

[0068] Optionally, the content risk decision result corresponding to the real-time video stream data can be determined according to the preliminary risk decision results and the intermediate risk decision results through the following steps, including:

[0069] The multi-modal risk detection weights corresponding to the first information density data and the second information density data can be obtained, and the preliminary risk decision results and the intermediate risk decision results can be fused according to the multi-modal risk detection weights to determine the content risk decision result corresponding to the real-time video stream data.

[0070] For example, assuming that the first information density data can include frame capture images, audio fragment data, and speech recognition text data, and the second information density data can include video fragment data, the content risk decision result corresponding to the real-time video stream data can be determined according to the preliminary risk decision results and the intermediate risk decision results through the following relational expression:

[0071]

[0072] Among them, can represent the content risk decision result, , , and can respectively represent the multi-modal risk detection weights corresponding to the frame capture image, audio fragment data, speech recognition text data, and video fragment data. For example, the multi-modal risk detection weights can be 0.2 for the frame capture image, 0.2 for the audio fragment data, 0.2 for the speech recognition text data, and 0.4 for the video fragment data. It can be understood that this is only a schematic example here, and specific custom settings can be made according to the actual situation; , , can respectively represent the preliminary risk decision results corresponding to the first information density data, can represent the intermediate risk decision result of the second information density data.

[0073] Fusing the preliminary risk decision results and the intermediate risk decision results through the multi-modal risk detection weights to obtain the content risk decision result corresponding to the real-time video stream data can effectively improve the accuracy and stability of the content risk decision result, realize the effective detection of the real-time video stream data with risky content, and ensure network data security.

[0074] Optionally, the update of the multi-modal risk detection weights corresponding to the first information density data and the second information density data can be realized through the steps in Figure 4 , as shown in Figure 4 , and specifically can include:

[0075] Step S410, obtain the initial risk detection weights corresponding to the first information density data and the second information density data;

[0076] Step S420, update the initial risk detection weights according to the content risk decision results at the historical moments associated with the real-time video stream data to obtain multi-modal risk detection weights, so that the updated multi-modal risk detection weights are adapted to the video content of the real-time video stream data.

[0077] Among them, the initial risk detection weight refers to the multi-modal risk detection weight corresponding to the first information density data and the second information density data set in advance based on experience. For example, the initial risk detection weight can be 0.2 for frame capture images, 0.2 for audio slice data, 0.2 for speech recognition text data, and 0.4 for video slice data. Of course, the initial risk detection weight can be customized according to the actual situation, and this embodiment is not limited thereto.

[0078] After at least one round of content decisions for the first information density data and the second information density data, the content risk decision results at the historical moments associated with the real-time video stream data can be obtained. At this time, the content risk decision results can show the decisive factors in the first information density data and the second information density data at the historical moments. For example, if the content risk decision result at a historical moment is determined by the audio slice data in the first information density data, then it can be considered that the high-risk content in the current real-time video stream data is in the audio information. Therefore, the multi-modal risk detection weight corresponding to the audio slice data can be adjusted. For example, the multi-modal risk detection weight can be adjusted to 0.1 for frame capture images, 0.4 for audio slice data, 0.2 for speech recognition text data, and 0.3 for video slice data; of course, similar updates are performed on other first information density data and second information density data, and the specific update ratio or update weight can be customized, and this embodiment does not make special limitations on this.

[0079] After multiple rounds of content decisions, the dynamically updated and adjusted multi-modal risk detection weights will be more and more adapted to the video content of the real-time video stream data, so that the content risk decision results determined based on the multi-modal risk detection weights are more in line with the video content in the real-time video stream data, effectively improving the accuracy and effectiveness of the content risk decision results, ensuring the precision of the risk processing of the real-time video stream data, and avoiding false detection rates.

[0080] In an exemplary embodiment of this specification, the extraction of multi-modal detection data corresponding to real-time video data can be implemented through the following steps, which specifically include the following steps:

[0081] A preset detection and extraction time can be obtained. The detection and extraction time can include a first extraction duration and a second extraction duration, and the second extraction duration is greater than the first extraction duration. By using the first extraction duration and combining with the timestamp information of the real-time video stream data, first information density data is extracted. By using the second extraction duration and combining with the timestamp information of the real-time video stream data, second information density data is extracted.

[0082] Among them, the first extraction duration refers to the duration interval data preset for extracting the first information density data in the real-time video stream data. Its duration is usually short to meet the real-time requirement. For example, the first extraction duration can be the video frame capture interval for extracting frame images, or the audio slice extraction duration for extracting audio slice data. The specific interval or duration can be set customarily. In this embodiment, no special limitation is imposed on the value setting of the first extraction duration.

[0083] The second extraction duration refers to the duration interval data preset for extracting the second information density data in the real-time video stream data. Its duration is long to capture more complete video content. For example, the second extraction duration can be the video slice extraction duration of video slice data, or the duration between video segmentation nodes. The specific duration can be set customarily or according to the application situation. In this exemplary embodiment, no special limitation is imposed on the value setting of the second extraction duration.

[0084] The second extraction duration can be greater than the first extraction duration. For example, the initial values of the first extraction duration and the second extraction duration can be defined through a system configuration file. For example, the video frame capture interval of frame images can be set to 0.5 seconds, the audio slice extraction duration of audio slice data can be set to 1 second, and the video slice extraction duration of video slice data can be set to 3 seconds. Furthermore, the start time and end time of each modality data can be marked by using the timestamp information to ensure the data time sequence consistency.

[0085] It can be understood that the first extraction duration and the second extraction duration may not be fixed. A real-time dynamic adjustment mechanism can be used to dynamically adjust the first extraction duration and the second extraction duration according to the current network latency and processing performance. Of course, a machine learning model can also be used to predict the risk hot spots in the real-time video stream data, and increase the sampling density for the risk hot spots, so as to dynamically adjust the first extraction duration and the second extraction duration according to the sampling density. No special limitation is imposed on the specific real-time dynamic adjustment mechanism in this embodiment.

[0086] Optionally, the extraction of the first information density data and the second information density data can be implemented through the following steps, which specifically can include the following steps:

[0087] Frame extraction processing can be performed on real-time video stream data based on the video frame extraction interval to obtain frame-extracted images, and the timestamp information corresponding to the frame-extracted images can be determined according to the timestamp information of the real-time video stream data. For example, after decoding the real-time video stream data through a streaming media decoder, frame sampling can be performed on the decoded real-time video stream data in combination with the video frame extraction interval to obtain frame-extracted images, and the timestamp information in the real-time video stream data at the current sampling moment can be used as the timestamp information corresponding to the frame-extracted images, and the frame-extracted images can be marked with timestamps in combination with the timestamp information to ensure the consistency of subsequent matching.

[0088] Audio extraction can be performed on real-time video stream data based on the audio slice extraction duration to obtain audio slice data, and the timestamp information corresponding to the audio slice data can be determined according to the timestamp information of the real-time video stream data. For example, the original audio data in the real-time video stream data can be extracted in real time through an audio decoder, and the original audio data can be segmented into audio slice data in combination with the audio slice extraction duration, and the absolute start timestamp and absolute end timestamp corresponding to the audio slice data can be determined in combination with the timestamp information in the real-time video stream data at the current sampling moment and the audio slice extraction duration to ensure the consistency of subsequent matching and fusion.

[0089] Speech recognition can be performed on the audio slice data segmented based on the audio slice extraction duration to determine speech recognition text data, and the timestamp information corresponding to the speech recognition text data can be determined according to the timestamp information corresponding to the audio slice data, such as the absolute start timestamp and absolute end timestamp, that is, the absolute start timestamp and absolute end timestamp of the speech recognition text data are consistent with the audio slice data. Furthermore, the frame-extracted images, audio slice data, speech recognition text data, and the corresponding timestamp information can be used as first information density data.

[0090] Video extraction can be performed on real-time video stream data based on the video slice extraction duration to obtain video slice data, and the absolute start timestamp and absolute end timestamp corresponding to the video slice data can be determined according to the timestamp information of the real-time video stream data.

[0091] By detecting the extraction time, real-time video stream data with high complexity can be segmented into multi-modal detection data, and the extraction of first information density data and second information density data can be achieved in combination with the first extraction duration and the second extraction duration, and the marking of sequential first information density data and second information density data can be achieved in combination with the timestamp information corresponding to the real-time video stream data, providing a sequential basis for the fusion of subsequent multi-modal detection data, improving the accuracy and efficiency of data fusion, and improving the risk detection efficiency and the accuracy of risk content decision results.

[0092] In an exemplary embodiment of this specification, it can be achieved through Figure 5The steps in detect the first information density data to determine the risk detection result, with reference to Figure 5 As shown in, specifically, it may include:

[0093] Step S510, perform feature analysis on the intercepted frame image to obtain the intercepted frame image features, and determine the intercepted frame detection result through the intercepted frame image features;

[0094] Step S520, perform audio analysis on the audio slice data to obtain the pitch content features, and determine the audio detection result through the pitch content features;

[0095] Step S530, perform content analysis on the speech recognition text data to obtain keyword features, and determine the text detection result through the keyword features;

[0096] Step S540, store the intercepted frame detection result, the audio detection result, and the text detection result as the risk detection result corresponding to the first information density data.

[0097] Among them, the risk content features in the intercepted frame image can be extracted through an image feature extraction model based on a convolutional neural network to obtain the intercepted frame image features; or the risk content features in the intercepted frame image can be extracted based on traditional image processing technologies such as object recognition and edge segmentation to obtain the intercepted frame image features. In this exemplary embodiment, there is no special limitation on the feature analysis method of the intercepted frame image. After obtaining the intercepted frame image features, the intercepted frame image features can be compared with the risk image features in the preset risk content database to determine the intercepted frame detection result.

[0098] The extraction and classification of pitch features such as timbre and pitch contained in the audio slice data can be realized through an acoustic feature extraction model based on a Transformer network to obtain the pitch content features; of course, the extraction and classification of spectral features, high-frequency components, etc. contained in the audio slice data can also be realized through Mel frequency cepstral coefficients to obtain the pitch content features. In this exemplary embodiment, there is no special limitation on the audio analysis method of the audio slice data. After obtaining the pitch content features, the pitch content features can be compared with the risk pitch content features in the preset risk content database to determine the audio detection result.

[0099] The content analysis of the speech recognition text data can be realized through an emotion analysis model based on LSTM (Long Short Term Memory) combined with a keyword extraction model to obtain keyword features, and the keyword features can be compared with the risk keyword features in the preset risk content database to determine the text detection result.

[0100] Through comprehensive detection of the first information density data, due to the different output orders of the data in the first information density data, it is possible to quickly detect risk content in real-time video stream data, effectively improve the risk detection efficiency, shorten the risk detection time limit, and when the detected risk level is relatively low, integrate it into the second information density data to assist the risk detection and content decision-making of the second information density data, effectively avoiding the calculation of duplicate data, reducing the amount of duplicate calculation, saving computing resources, and further improving the risk detection efficiency while ensuring the accuracy of the risk content decision result.

[0101] In an exemplary embodiment of this specification, the integration of the first information density data and the risk detection result into the second information density data can be achieved through the steps in Figure 6 , with reference to Figure 6 shown, and specifically may include:

[0102] Step S610, obtain the absolute start timestamp and the absolute end timestamp corresponding to the second information density data;

[0103] Step S620, match the first information density data within the corresponding timestamp range based on the absolute start timestamp and the absolute end timestamp;

[0104] Step S630, according to the timestamp information corresponding to the matched first information density data, integrate the matched first information density data and the corresponding risk detection result into the detection information associated with the second information density data to obtain the fused second information density data.

[0105] Among them, the absolute start timestamp is a time identifier used to mark the start time point in video slice data. This time identifier is based on the playback time axis of real-time video stream data and is consistent with the overall time sequence of real-time video stream data, ensuring that different modality data can be synchronized and associated through timestamps.

[0106] The absolute end timestamp is a time identifier used to mark the end time point of video slice data. This time identifier marks the end time point of the video slice and is also based on the playback time axis of the real-time video stream, ensuring the consistency of the end points of different modality data.

[0107] The absolute start timestamp and the absolute end timestamp are exactly consistent with the playback time axis of the real-time video stream in marking the start time point and the end time point of video slice data, which is different from the timestamp marking using the system clock or based on frame numbers.

[0108] By matching the corresponding timestamp range through the absolute start timestamp and the absolute end timestamp, and matching the first information density data and the corresponding risk detection results to achieve the association of modal data, it can ensure the complete alignment of the part of the second information density data that has duplicate data with the first information density data. Thus, the first information density data and its detection results can be directly used as the associated information for the same part of the data in the second information density data, thereby avoiding duplicate calculations for this part of the duplicate data, effectively reducing the amount of duplicate calculations, saving computing resources, and improving the risk detection efficiency.

[0109] Figure 7 This is a schematic flowchart of a process for implementing risk detection by hierarchically processing multi-modal detection data according to information density provided by an embodiment of this specification.

[0110] Reference Figure 7As shown, in step S710, obtain real-time video stream data transmitted in real time; in step S720, the preprocessing module can perform data processing on the real-time video stream data to obtain multi-modal detection data corresponding to the real-time video stream data; for example, frame extraction processing can be performed on the real-time video stream data based on the video frame capture interval to obtain frame-captured images, and the time-stamp information corresponding to the frame-captured images can be determined according to the time-stamp information of the real-time video stream data; audio extraction can be performed on the real-time video stream data based on the audio slice extraction duration to obtain audio slice data, and the time-stamp information corresponding to the audio slice data can be determined according to the time-stamp information of the real-time video stream data; speech recognition can be performed on the audio slice data to determine speech recognition text data, and the time-stamp information corresponding to the speech recognition text data can be determined according to the time-stamp information corresponding to the audio slice data; the frame-captured images, audio slice data, speech recognition text data, and the corresponding time-stamp information are used as the first information density data; video extraction can be performed on the real-time video stream data based on the video slice extraction duration to obtain video slice data, and the absolute start time-stamp and absolute end time-stamp corresponding to the video slice data can be determined according to the time-stamp information of the real-time video stream data, and the video slice data and the corresponding absolute start time-stamp and absolute end time-stamp are used as the second information density data; in step S730, the risk detection results corresponding to each first information density data are sequentially input into the risk detection system according to the output order for risk detection to obtain the preliminary risk decision results corresponding to each first information density data, and the obtained preliminary risk decision results are input into the content decision system for content decision; for example, feature analysis can be performed on the frame-captured images to obtain frame-captured image features, and the frame-captured detection results can be determined through the frame-captured image features; audio analysis can be performed on the audio slice data to obtain pitch content features, and the audio detection results can be determined through the pitch content features; content analysis can be performed on the speech recognition text data to obtain keyword features, and the text detection results can be determined through the keyword features; in step S740, the frame-captured detection results, audio detection results, and text detection results are stored as the risk detection results corresponding to the first information density data and stored in the database, and the video slice data used as the second information density data is stored in the database; if it is determined that the content risk level of any preliminary risk decision result is greater than or equal to the content risk level threshold, the preliminary risk decision result is directly used as the content risk decision result corresponding to the real-time video stream data, and the current process ends; if it is determined that the content risk levels of all preliminary risk decision results are less than the content risk level threshold, the first information density data and the risk detection results are fused into the second information density data, and step S750 is executed; in step S750, the absolute start time-stamp and absolute end time-stamp corresponding to the second information density data can be obtained; match the first information density data within the corresponding time-stamp range based on the absolute start time-stamp and absolute end time-stamp;According to the timestamp information corresponding to the first information density data that is matched, integrate the matched first information density data and the corresponding risk detection results into the detection information associated with the second information density data to obtain the fused second information density data; in step S760, the fused second information density data can be input into the content decision-making system, so that the content decision-making system makes content decisions on the second information density data based on the risk detection results corresponding to the first information density data to obtain the intermediate risk decision result of the fused second information density data; determine the content risk decision result corresponding to the real-time video stream data according to each preliminary risk decision result and the intermediate risk decision result, and end the current process.

[0111] In summary, in the video content risk detection method provided in the embodiments of this specification, by extracting multi-modal detection data corresponding to real-time video stream data, and dividing the multi-modal detection data into first information density data and second information density data according to the information density, and the information density of the second information density data is greater than that of the first information density data; detect the first information density data to determine the risk detection result, fuse the first information density data and the risk detection result into the second information density data, and make content decisions on the fused second information density data to obtain the content risk decision result corresponding to the real-time video stream data. On the one hand, by hierarchically processing the multi-modal detection data according to the information density and performing the prior risk decision processing on the first information density data, it is possible to quickly discover the risk content in the real-time video stream data, avoid the large computational pressure of the high-information density data, reduce the computational complexity in the overall detection process, shorten the time span from the occurrence of the risk to being detected, improve the risk detection efficiency, thereby improving the response speed to the risk content in the real-time video stream data and ensuring the network data security; on the other hand, after obtaining the risk detection result of the first information density data through the prior risk decision processing, fuse the risk detection result of the first information density data with the second information density data. At this time, if there are duplicate modal data in the second information density data, the risk detection result of the first information density data can be directly used as the risk detection result of the duplicate data, avoiding a large amount of duplicate calculations caused by making risk decisions on the duplicate data, saving computational resources, and at the same time improving the risk decision efficiency of the second information density data, further improving the risk detection efficiency while ensuring the accuracy and stability of the risk detection result.

[0112] Please refer to Figure 8 , which is a schematic structural diagram of a video content risk detection device provided in the embodiments of this specification. As Figure 8As shown, the video content risk detection device 800 can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, the video content risk detection device 800 includes a multi-modal detection data extraction module 810, a risk detection module 820, and a content risk decision-making module 830, specifically including:

[0113] The multi-modal detection data extraction module 810 can be used to obtain real-time video stream data and extract the multi-modal detection data corresponding to the real-time video stream data. The multi-modal detection data includes first information density data and second information density data, and the information density of the second information density data is greater than that of the first information density data.

[0114] The risk detection module 820 can be used to detect the first information density data and determine the risk detection result.

[0115] The content risk decision-making module 830 can be used to fuse the first information density data and the risk detection result into the second information density data, and perform content decision-making on the fused second information density data to obtain the content risk decision result corresponding to the real-time video stream data.

[0116] Optionally, the video content risk detection device 800 further includes a first information density data sequential processing module, which is configured to:

[0117] Input the risk detection results corresponding to each of the first information density data into the content decision-making system in the output order to obtain the preliminary risk decision results corresponding to each of the first information density data.

[0118] If it is determined that the content risk level of any of the preliminary risk decision results is greater than or equal to the content risk level threshold, then directly use the preliminary risk decision result as the content risk decision result corresponding to the real-time video stream data.

[0119] If it is determined that the content risk levels of all the preliminary risk decision results are less than the content risk level threshold, then fuse the first information density data and the risk detection result into the second information density data.

[0120] Optionally, the content risk decision-making module 830 is configured to:

[0121] Input the fused second information density data into the content decision-making system, so that the content decision-making system performs content decision-making on the second information density data based on the risk detection result corresponding to the first information density data to obtain the intermediate risk decision result of the fused second information density data.

[0122] Based on each of the preliminary risk decision results and the intermediate risk decision results, determine the content risk decision result corresponding to the real-time video stream data.

[0123] Optionally, the content risk decision module 830 is configured to:

[0124] Obtain the multi-modal risk detection weights corresponding to the first information density data and the second information density data;

[0125] Fuse the preliminary risk decision result and the intermediate risk decision result according to the multi-modal risk detection weights to determine the content risk decision result corresponding to the real-time video stream data.

[0126] Optionally, the content risk decision module 830 is configured to:

[0127] Obtain the initial risk detection weights corresponding to the first information density data and the second information density data;

[0128] Update the initial risk detection weights according to the content risk decision results at historical moments associated with the real-time video stream data to obtain multi-modal risk detection weights, so that the updated multi-modal risk detection weights are adapted to the video content of the real-time video stream data.

[0129] Optionally, the multi-modal detection data extraction module 810 is configured to:

[0130] Obtain a preset detection extraction time, where the detection extraction time includes a first extraction duration and a second extraction duration, and the second extraction duration is greater than the first extraction duration;

[0131] Extract the first information density data through the first extraction duration in combination with the timestamp information of the real-time video stream data;

[0132] Extract the second information density data through the second extraction duration in combination with the timestamp information of the real-time video stream data.

[0133] Optionally, the first extraction duration includes a video frame capture interval and an audio slice extraction duration; the multi-modal detection data extraction module 810 is configured to:

[0134] Perform frame extraction processing on the real-time video stream data based on the video frame capture interval to obtain frame-captured images, and determine the timestamp information corresponding to the frame-captured images according to the timestamp information of the real-time video stream data;

[0135] Extract audio from the real-time video stream data based on the duration of the audio slice extraction to obtain audio slice data, and determine the time stamp information corresponding to the audio slice data according to the time stamp information of the real-time video stream data;

[0136] Perform speech recognition on the audio slice data to determine speech recognition text data, and determine the time stamp information corresponding to the speech recognition text data according to the time stamp information corresponding to the audio slice data;

[0137] Use the frame captured image, the audio slice data, the speech recognition text data, and the corresponding time stamp information as the first information density data.

[0138] Optionally, the second extraction duration includes a video slice extraction duration; the multimodal detection data extraction module 810 is configured to:

[0139] Extract video from the real-time video stream data based on the video slice extraction duration to obtain video slice data, and determine the absolute start time stamp and the absolute end time stamp corresponding to the video slice data according to the time stamp information of the real-time video stream data.

[0140] Optionally, the first information density data at least includes a frame captured image, audio slice data, and speech recognition text data; the risk detection module 820 is configured to:

[0141] Perform feature analysis on the frame captured image to obtain frame captured image features, and determine a frame captured detection result through the frame captured image features;

[0142] Perform audio analysis on the audio slice data to obtain pitch content features, and determine an audio detection result through the pitch content features;

[0143] Perform content analysis on the speech recognition text data to obtain keyword features, and determine a text detection result through the keyword features;

[0144] Store the frame captured detection result, the audio detection result, and the text detection result as the risk detection result corresponding to the first information density data.

[0145] Optionally, the content risk decision module 830 is configured to:

[0146] Obtain the absolute start time stamp and the absolute end time stamp corresponding to the second information density data;

[0147] Match the first information density data within the corresponding time stamp range based on the absolute start time stamp and the absolute end time stamp;

[0148] According to the timestamp information corresponding to the first information density data that is matched, integrate the matched first information density data and the corresponding risk detection result into the detection information associated with the second information density data, so as to obtain the fused second information density data.

[0149] The above device embodiments correspond to the method embodiments. For specific descriptions, reference can be made to the descriptions in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference can be made to the corresponding method embodiments.

[0150] The embodiments of the present specification also provide a computer storage medium. The computer storage medium can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the method as described in the embodiments of the present specification. The specific execution process can refer to the specific descriptions in the embodiments of the present specification and will not be elaborated here.

[0151] The present specification also provides a computer program product. The computer program product stores at least one instruction, and the at least one instruction is loaded and executed by the processor to perform the method as described in the embodiments of the present specification. The specific execution process can refer to the specific descriptions in the embodiments of the present specification and will not be elaborated here.

[0152] The embodiments of the present specification also provide Figure 9 the structural schematic diagram of the electronic device shown. As Figure 9 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above voice activity detection method.

[0153] Of course, in addition to the software implementation manner, the present specification does not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit and can also be hardware or a logic device.

[0154] For an improvement in a technology, it can be clearly distinguished whether it is an improvement in hardware (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement in method processes). However, with the development of technology, many improvements in method processes today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method process into the hardware circuit. Therefore, it cannot be said that an improvement in a method process cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. The designer can program by himself to "integrate" a digital system on a piece of PLD without asking the chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that only by slightly logically programming the method process with the above-mentioned several hardware description languages and programming it into the integrated circuit can the hardware circuit implementing the logical method process be easily obtained.

[0155] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0156] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0157] For the convenience of description, the above devices are described by dividing them into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0158] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0159] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0162] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0163] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0164] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0165] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0166] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0167] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0168] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0169] The above is only the embodiment of this specification and is not used to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A video content risk detection method, the method comprising: Obtaining real-time video stream data, and extracting multi-modal detection data corresponding to the real-time video stream data, the multi-modal detection data including first information density data and second information density data, and the information density of the second information density data being greater than that of the first information density data; Detecting the first information density data to determine a risk detection result; Sequentially inputting the risk detection results corresponding to the first information density data into a content decision-making system according to the output order to obtain preliminary risk decision results corresponding to the first information density data; If it is determined that the content risk level of any one of the preliminary risk decision results is greater than or equal to a content risk level threshold, then directly use the preliminary risk decision result as the content risk decision result corresponding to the real-time video stream data; If it is determined that the content risk levels of all the preliminary risk decision results are less than the content risk level threshold, then fuse the first information density data and the risk detection results into the second information density data, and perform content decision-making on the fused second information density data to obtain the content risk decision result corresponding to the real-time video stream data.

2. The video content risk detection method according to claim 1, wherein performing content decision-making on the fused second information density data to obtain the content risk decision result corresponding to the real-time video stream data comprises: Inputting the fused second information density data into the content decision-making system, so that the content decision-making system performs content decision-making on the second information density data based on the risk detection result corresponding to the first information density data to obtain an intermediate risk decision result of the fused second information density data; Determining the content risk decision result corresponding to the real-time video stream data according to the preliminary risk decision results and the intermediate risk decision result.

3. The video content risk detection method according to claim 2, wherein determining the content risk decision result corresponding to the real-time video stream data according to the preliminary risk decision results and the intermediate risk decision result comprises: Obtaining multi-modal risk detection weights corresponding to the first information density data and the second information density data; Fusing the preliminary risk decision results and the intermediate risk decision result according to the multi-modal risk detection weights to determine the content risk decision result corresponding to the real-time video stream data.

4. The video content risk detection method according to claim 3, wherein obtaining the multi-modal risk detection weights corresponding to the first information density data and the second information density data comprises: Obtaining initial risk detection weights corresponding to the first information density data and the second information density data; Updating the initial risk detection weights according to the content risk decision result at the historical moment associated with the real-time video stream data to obtain multi-modal risk detection weights, so that the updated multi-modal risk detection weights are adapted to the video content of the real-time video stream data.

5. The video content risk detection method according to claim 1, wherein the extraction of the multi-modal detection data corresponding to the real-time video stream data includes: Obtaining a preset detection extraction time, where the detection extraction time includes a first extraction duration and a second extraction duration, and the second extraction duration is greater than the first extraction duration; Extracting the first information density data by means of the first extraction duration in combination with the timestamp information of the real-time video stream data; Extracting the second information density data by means of the second extraction duration in combination with the timestamp information of the real-time video stream data.

6. The video content risk detection method according to claim 5, wherein the first extraction duration includes a video frame capture interval and an audio slice extraction duration; and the extraction of the first information density data by means of the first extraction duration in combination with the timestamp information of the real-time video stream data includes: Performing frame extraction processing on the real-time video stream data based on the video frame capture interval to obtain frame-captured images, and determining the timestamp information corresponding to the frame-captured images according to the timestamp information of the real-time video stream data; Performing audio extraction on the real-time video stream data based on the audio slice extraction duration to obtain audio slice data, and determining the timestamp information corresponding to the audio slice data according to the timestamp information of the real-time video stream data; Performing speech recognition on the audio slice data to determine speech recognition text data, and determining the timestamp information corresponding to the speech recognition text data according to the timestamp information corresponding to the audio slice data; Taking the frame-captured images, the audio slice data, the speech recognition text data, and the corresponding timestamp information as the first information density data.

7. The video content risk detection method according to claim 5, wherein the second extraction duration includes a video slice extraction duration; and the extraction of the second information density data by means of the second extraction duration in combination with the timestamp information of the real-time video stream data includes: Performing video extraction on the real-time video stream data based on the video slice extraction duration to obtain video slice data, and determining the absolute start timestamp and the absolute end timestamp corresponding to the video slice data according to the timestamp information of the real-time video stream data.

8. The video content risk detection method according to claim 1, wherein the first information density data at least includes frame-captured images, audio slice data, and speech recognition text data; The detection of the first information density data to determine the risk detection result includes: Performing feature analysis on the frame-captured images to obtain frame-captured image features, and determining the frame-captured detection result through the frame-captured image features; Performing audio analysis on the audio slice data to obtain pitch content features, and determining the audio detection result through the pitch content features; Performing content analysis on the speech recognition text data to obtain keyword features, and determining the text detection result through the keyword features; Store the frame truncation detection result, the audio detection result, and the text detection result as the risk detection result corresponding to the first information density data.

9. The video content risk detection method according to claim 1, wherein the fusing the first information density data and the risk detection result into the second information density data includes: Obtain the absolute start timestamp and the absolute end timestamp corresponding to the second information density data; Match the first information density data within the corresponding timestamp range based on the absolute start timestamp and the absolute end timestamp; According to the timestamp information corresponding to the matched first information density data, integrate the matched first information density data and the corresponding risk detection result into the detection information associated with the second information density data to obtain the fused second information density data.

10. A video content risk detection device, the device includes: A multimodal detection data extraction module, configured to obtain real-time video stream data and extract multimodal detection data corresponding to the real-time video stream data, the multimodal detection data includes first information density data and second information density data, and the information density of the second information density data is greater than that of the first information density data; A risk detection module, configured to detect the first information density data to determine a risk detection result; A content risk decision module, configured to sequentially input the risk detection results corresponding to each of the first information density data into a content decision system according to the output order to obtain preliminary risk decision results corresponding to each of the first information density data; If it is determined that the content risk level of any one of the preliminary risk decision results is greater than or equal to the content risk level threshold, then directly use the preliminary risk decision result as the content risk decision result corresponding to the real-time video stream data; if it is determined that the content risk levels of all the preliminary risk decision results are less than the content risk level threshold, then fuse the first information density data and the risk detection result into the second information density data, and perform content decision on the fused second information density data to obtain the content risk decision result corresponding to the real-time video stream data.

11. A storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.

12. An electronic device, characterized in that, Including: A processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the steps of the method according to any one of claims 1 to 9.

13. A computer program product having at least one instruction stored thereon, characterized in that, When the at least one instruction is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Content carrier risk detection method, device, apparatus and medium

    CN109492401A

  • Living body detection method, device and equipment

    CN115546908A