An audio and video auditing method, device and equipment and readable storage medium
By using a three-level review network model to filter audio and video data into broad and specific categories, and combining blacklist and whitelist feature libraries, the problem of high cost and low efficiency of manual review is solved, and efficient automatic review of audio and video data is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU QUYAN NETWORK TECH CO LTD
- Filing Date
- 2023-03-27
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, the review of audio and video data relies on manual monitoring, which results in high labor costs and low review efficiency, and cannot adapt to the rapid growth of large amounts of audio and video data.
The audio and video review model adopts a three-level review network, including a first-level analysis network, a second-level analysis network, and a third-level analysis network. It automatically reviews audio and video data by filtering through broad and detailed category analysis and combining blacklist and whitelist feature libraries for similarity calculation.
It reduces the pressure and cost of manual review, improves review efficiency, and achieves efficient automatic review of audio and video data.
Smart Images

Figure CN116561619B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of content review and identification, and more specifically, to an audio and video content review method, apparatus, device, and readable storage medium. Background Technology
[0002] Currently, with the rapid development of pan-entertainment social networking, video publishing, and live streaming, the number of users and the amount of audio and video uploaded to platforms have increased significantly, generating tens of thousands to hundreds of thousands of hours of audio and video content daily. The governance of the online environment is very strict, and reviewing uploaded audio and video data is a crucial task for all major platforms. Platforms have an obligation to screen and review user-uploaded audio and video data to determine whether it constitutes illegal, vulgar, or prohibited audio and video works, thus ensuring the healthy content uploaded by users.
[0003] To comply with regulatory requirements, platform management companies typically need to set up an audit department. Videos uploaded by users to the backend every day are sent to the audit department for cross-review. Only after the audit is approved can the videos be displayed on the platform and shared with other users. However, using manual monitoring is costly in terms of manpower and has low audit efficiency. It cannot adapt to situations with a large amount of audio and video data, and therefore cannot meet the platform's future development requirements. Summary of the Invention
[0004] In view of this, this application provides an audio and video review method, apparatus, device, and readable storage medium. The audio and video review model, consisting of a three-level review network, reviews the audio and video data to be reviewed. It is more cost-effective and efficient in terms of review machine costs. It innovatively uses features from blacklist and whitelist feature libraries corresponding to sub-category violation tags to enable automatic machine review of audio and video data, reducing the pressure and cost of manual review and improving review efficiency.
[0005] An audio / video review method includes:
[0006] Obtain audio and video data pending review;
[0007] An audio-visual review model is defined, comprising a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network performs broad-category analysis and filtering on the input audio-visual data to be reviewed according to major violation tags, obtaining first-level analysis results and generating each violation audio-visual dataset corresponding to each major violation tag. The second-level analysis network performs detailed-category analysis and filtering on each violation audio-visual data in each violation audio-visual dataset, obtaining second-level analysis results and determining the sub-category violation tags corresponding to each violation audio-visual data. The third-level analysis network extracts the implicit features of each violation audio-visual data and performs similarity calculation with features in the blacklist and whitelist feature libraries corresponding to the sub-category violation tags, obtaining similarity comparison results. The result processing network determines the review result of the audio-visual data to be reviewed based on the first-level analysis results, the second-level analysis results, and the similarity comparison results.
[0008] The audio and video data to be reviewed is input into the audio and video review model to obtain the review result of the audio and video review model on the audio and video data to be reviewed.
[0009] Preferably, the review result for the audio and video data to be reviewed is determined based on the primary analysis result, the secondary analysis result, and the similarity comparison result, including:
[0010] The review results of each audio and video data item in the audio and video data to be reviewed that has a non-violation result in the first-level analysis result and / or the second-level analysis result are determined as approved;
[0011] The first-level analysis result and the second-level analysis result are deemed to be violations, and the similarity comparison result is the review result of each audio and video data whose feature similarity with the feature in the blacklist feature library exceeds a preset threshold, which is determined to be the review failure;
[0012] If the first-level analysis result and the second-level analysis result are deemed as violations, and the similarity comparison result is the review result of each audio and video data whose feature similarity with the whitelist feature library exceeds a preset threshold, then the review is deemed as passed.
[0013] Preferred options also include:
[0014] The audio and video data whose first-level analysis results and second-level analysis results are deemed to be in violation, and whose similarity comparison results are all within the preset thresholds for feature similarity with the features in the blacklist feature library and the whitelist feature library, are sent to the manual review module, and the feedback result of the manual review module is determined as the review result.
[0015] Preferably, the first-level analysis network consists of a depthwise convolutional network with an inverted residual structure and pointwise convolutional layers;
[0016] The process by which the primary analysis network performs category analysis and filtering on the input audio and video data to be reviewed according to major violation tags, obtains primary analysis results, and generates each violation audio and video dataset corresponding to the major violation tags includes:
[0017] The depthwise convolutional network of the first-level analysis network performs depthwise convolution on the input audio and video data to be reviewed according to the parameters corresponding to the major categories of violation tags, and uses lightweight filtering to obtain the information values of each channel.
[0018] The pointwise convolutional layers of the first-level analysis network are constructed by linearly combining the information values of each channel, and the violation audio and video datasets corresponding to each major category of violation label are determined according to the filtering range corresponding to each major category of violation label.
[0019] Preferably, the secondary analysis network consists of a residual convolutional network, a pooling layer, and a result output layer;
[0020] The secondary analysis network performs detailed category analysis and filtering on each piece of illegal audio and video data in each illegal audio and video dataset, obtains secondary analysis results, and determines the sub-category violation label corresponding to each piece of illegal audio and video data. This process includes:
[0021] The residual convolutional network of the secondary analysis network extracts features from each illegal audio and video data in each illegal audio and video dataset to obtain the feature information of each illegal audio and video data in each illegal audio and video dataset.
[0022] The pooling layer of the secondary analysis network performs dimensionality reduction pooling on the feature information of each illegal audio and video data in each illegal audio and video dataset to generate integrated features of each illegal audio and video data in each illegal audio and video dataset.
[0023] The output layer of the secondary analysis network performs secondary violation screening based on the integrated features of each illegal audio and video data in each illegal audio and video dataset, obtains the secondary analysis results, and determines the sub-class violation label corresponding to each illegal audio and video data.
[0024] Preferably, the three-level analysis network consists of an image feature extraction layer, a speech feature extraction layer, and a similarity comparison layer;
[0025] The process of extracting implicit features from each piece of illegal audio and video data by the three-level analysis network, and calculating similarity between these features and the features in the blacklist and whitelist feature libraries corresponding to the sub-class illegal tags, to obtain the similarity comparison results, includes:
[0026] The image feature extraction layer of the three-level analysis network extracts the implicit features of the video frames in each illegal audio and video data to obtain the general features of the frames;
[0027] The speech feature extraction layer of the three-level analysis network extracts the implicit features of the speech signal in each illegal audio and video data to obtain general speech features;
[0028] The similarity calculation layer of the three-level analysis network calculates cosine similarity between the common features of the image and the common features of the voice and the features in the blacklist feature library and whitelist feature library corresponding to the subclass violation label, and obtains the similarity comparison result.
[0029] Preferably, the process of training the audio and video review model includes:
[0030] Acquire training audio and video data, which are labeled with corresponding review result information;
[0031] The training audio and video data is input into a preset initial audio and video review model to obtain the review result of the training audio and video data output by the initial audio and video review model.
[0032] The initial audio and video review model is trained with the goal of ensuring that the review results of the training audio and video data are consistent with the corresponding review result information labeled in the training audio and video data.
[0033] When the initial audio and video review model meets the preset training conditions, the trained initial audio and video review model is used as the audio and video review model.
[0034] An audio and video review device, comprising:
[0035] The audio and video acquisition module is used to acquire audio and video data to be reviewed.
[0036] The model determination module is used to determine the audio and video review model. The audio and video review model includes a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network is used to perform category analysis and filtering on the input audio and video data to be reviewed according to major category violation tags, to obtain first-level analysis results and generate each violation audio and video dataset corresponding to each major category violation tag. The second-level analysis network is used to perform detailed category analysis and filtering on each violation audio and video data in each violation audio and video dataset, to obtain second-level analysis results and determine the sub-category violation tags corresponding to each violation audio and video data. The third-level analysis network is used to extract the implicit features of each violation audio and video data and perform similarity calculation with the features in the blacklist feature library and whitelist feature library corresponding to the sub-category violation tags to obtain similarity comparison results. The result processing network is used to determine the review result of the audio and video data to be reviewed based on the first-level analysis results, the second-level analysis results, and the similarity comparison results.
[0037] The review result module is used to input the audio and video data to be reviewed into the audio and video review model, and obtain the review result of the audio and video review model on the audio and video data to be reviewed.
[0038] An audio and video review device, comprising a memory and a processor;
[0039] The memory is used to store programs;
[0040] The processor is used to execute the program to implement the various steps of the audio and video review method described above.
[0041] A readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the various steps of the audio and video review method described above.
[0042] As can be seen from the above technical solutions, the audio and video review method, apparatus, device and readable storage medium provided in this application embodiment obtains the audio and video data to be reviewed, determines the audio and video review model, inputs the audio and video data to be reviewed into the audio and video review model, and obtains the review result of the audio and video review model on the audio and video data to be reviewed. The audio and video review model includes a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network is used to perform category analysis and filtering on the input audio and video data to be reviewed according to major violation tags, to obtain first-level analysis results and generate each violation audio and video dataset corresponding to each major violation tag. The second-level analysis network is used to perform detailed category analysis and filtering on each violation audio and video data in each violation audio and video dataset, to obtain second-level analysis results and determine the sub-category violation tags corresponding to each violation audio and video data. The third-level analysis network is used to extract the implicit features of each violation audio and video data and perform similarity calculation with the features in the blacklist feature library and whitelist feature library corresponding to the sub-category violation tags to obtain similarity comparison results. The result processing network is used to determine the review result of the audio and video data to be reviewed based on the first-level analysis results, the second-level analysis results, and the similarity comparison results.
[0043] This application employs a cascaded review process, utilizing a three-tiered review network to review audio and video data. The first-tier analysis network is a small model with high recall and low precision, filtering out most normal audio and video data through broad category analysis, retaining potentially inappropriate data for the second-tier analysis network. The second-tier analysis network is a large model with high recall and high precision, aiming to finely classify violation types. It further filters potentially inappropriate audio and video data through fine category analysis, resulting in a more detailed classification of the inappropriate audio and video data. The third-tier analysis network is an implicit feature extraction model, aiming to implicitly encode images and audio signals using an unsupervised model, and then calculate similarity comparison results by comparing them with features in blacklist and whitelist feature libraries.
[0044] This application designs a cascaded model structure, which is more cost-effective and efficient in terms of machine review. It innovatively uses features from the blacklist and whitelist feature libraries corresponding to the sub-category violation tags to enable automatic machine review of audio and video data, reducing the pressure and cost of manual review and improving review efficiency. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0046] Figure 1 This is a flowchart of an audio / video review method disclosed in this application;
[0047] Figure 2 This is a schematic diagram of the structure of an audio and video review model disclosed in this application;
[0048] Figure 3 A flowchart illustrating one method for determining the audit result disclosed in this application;
[0049] Figure 4 This is a schematic diagram of the structure of a first-order analysis network disclosed in this application;
[0050] Figure 5 This is a schematic diagram of the structure of a two-level analysis network disclosed in this application;
[0051] Figure 6 This is a schematic diagram of the structure of a three-level analysis network disclosed in this application;
[0052] Figure 7 This is a structural block diagram of an audio and video review device disclosed in this application;
[0053] Figure 8 This is a hardware structure block diagram of the audio and video review equipment disclosed in this application. Detailed Implementation
[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0055] The following section introduces the solution proposed in this application. The technical solution is as follows, and details are provided below.
[0056] Figure 1 This is a flowchart of an audio / video review method disclosed in an embodiment of this application. This audio / video review method can be applied to audio / video review devices, such as... Figure 1 As shown, the method may include:
[0057] Step S1: Obtain the audio and video data to be reviewed.
[0058] Specifically, the audio and video review method provided in this application can assist reviewers in reviewing audio and video data. For example, in various video websites, in the scenario of reviewing audio and video data, the method can automatically review the audio and video data uploaded to the website. Only audio and video data that passes the review can be displayed to the public, while audio and video data that fails the review cannot be displayed to the public.
[0059] Step S2: Determine the audio and video review model.
[0060] Specifically, the audio and video review model includes a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network performs broad-category analysis and filtering of the input audio and video data to be reviewed according to major violation tags, obtaining first-level analysis results and generating datasets of violating audio and video data corresponding to each major violation tag. The second-level analysis network performs detailed-category analysis and filtering of each violating audio and video data in each dataset, obtaining second-level analysis results and determining the sub-category violation tags corresponding to each violating audio and video data. The third-level analysis network extracts the implicit features of each violating audio and video data and calculates similarity with features in the blacklist and whitelist feature libraries corresponding to the sub-category violation tags, obtaining similarity comparison results. The result processing network determines the review result of the audio and video data to be reviewed based on the first-level analysis results, second-level analysis results, and similarity comparison results.
[0061] After obtaining the audio and video data to be reviewed, the data is input into the audio and video review model. This model can analyze and identify the audio and video data to be reviewed and output the corresponding audio and video review results.
[0062] Figure 2 This is a schematic diagram of the audio / video review model disclosed in this application. The audio / video review model includes a first-level analysis network A, a second-level analysis network B, a third-level analysis network C, and a result processing network D. The cascaded review structure of the audio / video review model, composed of three review networks, reviews the audio / video data to be reviewed. The first-level analysis network is a small model with high recall and low precision, which filters out most normal audio / video data through broad category analysis, retaining potentially illegal audio / video data for the second-level analysis network. The second-level analysis network is a large model with high recall and high precision, aiming to finely classify violation types. It performs a second screening of potentially illegal audio / video data through fine category analysis, obtaining a more detailed classification of the illegal audio / video data. The third-level analysis network is an implicit feature extraction model, aiming to implicitly encode images and audio signals using an unsupervised model, and obtain similarity comparison results after calculating similarity with features in blacklist and whitelist feature libraries.
[0063] Understandably, in order to further shorten detection time and improve detection efficiency, this application can pre-train a preset initial audio-visual review model based on training audio-visual data labeled with corresponding review result information, thereby obtaining a trained audio-visual review model. When automated review of audio-visual data is required, the audio-visual review model can be directly obtained and used without wasting a lot of training time.
[0064] like Figure 3 As shown, the process of determining the review result of audio and video data to be reviewed based on the results of primary analysis, secondary analysis, and similarity comparison can specifically include the following four situations:
[0065] Scenario 1: The review results of each audio and video data item in the pending review whose first-level analysis result and / or second-level analysis result is not in violation are determined as approved.
[0066] Scenario 2: The review results of audio and video data that are deemed to be in violation by both the first-level and second-level analysis results and whose similarity comparison results exceed the preset threshold with the features in the blacklist feature library will be determined as failing the review.
[0067] Scenario 3: The review results of audio and video data that are deemed to be in violation by both the first-level and second-level analysis results, and whose similarity comparison results exceed the preset threshold with the features in the whitelist feature library, are determined to be approved.
[0068] Scenario 4: Send all audio and video data that are deemed to be in violation by both the first-level and second-level analysis results and whose similarity comparison results show that their similarity to features in both the blacklist and whitelist feature libraries does not exceed the preset threshold to the manual review module, and determine the feedback result from the manual review module as the review result.
[0069] Step S3: Input the audio and video data to be reviewed into the audio and video review model to obtain the review results of the audio and video data to be reviewed output by the audio and video review model.
[0070] Specifically, the audio and video data to be reviewed is input into the audio and video review model. The model reviews the data by performing broad category analysis and filtering using a first-level analysis network, obtaining first-level analysis results; a second-level analysis network performs detailed category analysis and filtering, obtaining second-level analysis results; and a third-level analysis network performs implicit feature and similarity calculations, obtaining similarity comparison results. Finally, the result processing network determines the review result for the audio and video data based on the first-level analysis results, second-level analysis results, and similarity comparison results, and outputs the review result.
[0071] As can be seen from the above technical solutions, the audio and video review method, apparatus, device and readable storage medium provided in this application embodiment obtains the audio and video data to be reviewed, determines the audio and video review model, inputs the audio and video data to be reviewed into the audio and video review model, and obtains the review result of the audio and video data to be reviewed output by the audio and video review model. The audio and video review model includes a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network is used to perform category analysis and filtering on the input audio and video data to be reviewed according to major violation tags, obtain the first-level analysis results, and generate each violation audio and video dataset corresponding to each major violation tag. The second-level analysis network is used to perform detailed category analysis and filtering on each violation audio and video data in each violation audio and video dataset, obtain the second-level analysis results, and determine the sub-category violation tags corresponding to each violation audio and video data. The third-level analysis network is used to extract the implicit features of each violation audio and video data and calculate the similarity with the features in the blacklist feature library and whitelist feature library corresponding to the sub-category violation tags to obtain the similarity comparison results. The result processing network is used to determine the review result of the audio and video data to be reviewed based on the first-level analysis results, the second-level analysis results, and the similarity comparison results.
[0072] This application employs a cascaded review process, utilizing a three-tiered review network to review audio and video data. The first-tier analysis network is a small model with high recall and low precision, filtering out most normal audio and video data through broad category analysis, retaining potentially inappropriate data for the second-tier analysis network. The second-tier analysis network is a large model with high recall and high precision, aiming to finely classify violation types. It further filters potentially inappropriate audio and video data through fine category analysis, resulting in a more detailed classification of the inappropriate audio and video data. The third-tier analysis network is an implicit feature extraction model, aiming to implicitly encode images and audio signals using an unsupervised model, and then calculate similarity comparison results by comparing them with features in blacklist and whitelist feature libraries.
[0073] This application designs a cascaded model structure, which is more cost-effective and efficient in terms of machine review. It innovatively uses features from the blacklist and whitelist feature libraries corresponding to the sub-category violation tags to enable automatic machine review of audio and video data, reducing the pressure and cost of manual review and improving review efficiency.
[0074] In some embodiments of this application, the structure and function of the audio and video review model of this application are described in detail. The audio and video review model includes a first-level analysis network A, a second-level analysis network B, a third-level analysis network C, and a result processing network D. The following describes the process in conjunction with... Figures 4 to 6 The first-level analysis network A, the second-level analysis network B, the third-level analysis network C, and the result processing network D will be introduced in detail in turn.
[0075] First-level analysis network A:
[0076] like Figure 4 As shown, the first-level analysis network A consists of a depthwise convolutional network with an inverted residual structure and pointwise convolutional layers.
[0077] The first-level analysis network performs category analysis and filtering on the input audio and video data to be reviewed according to major violation tags, obtains the first-level analysis results, and generates each violation audio and video dataset corresponding to the major violation tags. Specifically, it may include:
[0078] ① The depthwise convolutional network of the first-level analysis network performs depthwise convolution on the input audio and video data to be reviewed according to the parameters corresponding to the major categories of violation tags, and uses lightweight filtering to obtain the information values of each channel.
[0079] ② The pointwise convolutional layers of the first-level analysis network are constructed by linearly combining the information values of each channel, and the illegal audio and video datasets corresponding to the major categories of illegal labels are determined according to the filtering range corresponding to the major categories of illegal labels.
[0080] Specifically, the first-level analysis network employs an inverted residual structure, consisting of two parts: a depthwise convolutional network and a pointwise convolutional layer. This depthwise separable convolution is used to extract features. It first performs depthwise convolution, then pointwise convolution, significantly reducing the number of parameters while achieving the effect of ordinary convolution. The final features also contain all the feature information from the previous layer. Compared to conventional convolution operations, its parameter count and computational cost are relatively low. The ReLU6 activation function can be used in this application.
[0081] The characteristic of depthwise convolution is that it does not change the number of input channels, and one input channel corresponds to one output channel. Each output channel only contains information from its corresponding input channel, not information from all input channels. The number of parameters is calculated as follows:
[0082] para(num) = m * w * h
[0083] Where m is the number of input channels, and w and h are the width and height of the convolution kernel, respectively. Taking 4 input channels and a 3*3 kernel size as an example, the number of parameters is:
[0084] para(num) = 4 * 3 * 3 = 36
[0085] Pointwise convolution uses a 1x1 kernel to mix information from the channels output by depthwise convolution. Its parameters are calculated as follows:
[0086] para(num) = m * n * w * h
[0087] Where m, n, w, and h represent the number of input channels, the number of output channels, and the width and height of the convolution kernel, respectively.
[0088] The Level 1 analysis network uses a depthwise convolutional network to perform depthwise convolutions on the input audio and video data to be reviewed, based on the parameters corresponding to major violation labels, and employs lightweight filtering to obtain information values for each channel. Pointwise convolutional layers then linearly combine these channel information values to construct the corresponding violation audio and video datasets based on the filtering ranges for each major violation label. The Level 1 analysis network takes the entire volume of live audio and video data as input, which is massive in scale, and does not perform further detailed classification of each violation label. The Level 1 analysis network uses an inverted residual structure composed of depthwise convolutional networks and pointwise convolutional layers, which reduces model computation and cost, improves inference speed, and lowers deployment costs.
[0089] This application can also adjust the ratio of positive and negative samples and the category weights to achieve the goal of high recall of violation samples in the primary analysis network, thereby minimizing the omission of violations.
[0090] Second-level analysis network B:
[0091] like Figure 5 As shown, the secondary analysis network consists of a residual convolutional network, a pooling layer, and a result output layer.
[0092] The secondary analysis network performs detailed classification analysis and filtering on each piece of illegal audio and video data in each illegal audio and video dataset, obtains the secondary analysis results, and determines the sub-class violation label corresponding to each piece of illegal audio and video data. Specifically, it may include:
[0093] ① The residual convolutional network of the second-level analysis network extracts features from each illegal audio and video data in each illegal audio and video dataset to obtain the feature information of each illegal audio and video data in each illegal audio and video dataset.
[0094] ② The pooling processing layer of the secondary analysis network performs dimensionality reduction pooling processing on the feature information of each illegal audio and video data in each illegal audio and video dataset, generating integrated features of each illegal audio and video data in each illegal audio and video dataset.
[0095] ③ The output layer of the secondary analysis network performs secondary violation screening based on the integrated features of each illegal audio and video data in each illegal audio and video dataset, obtains the secondary analysis results, and determines the sub-class violation label corresponding to each illegal audio and video data.
[0096] Specifically, the secondary analysis network performs a second round of filtering on the potentially illegal audio and video data obtained from the primary analysis network. It subcategorizes the data by violation type, eliminating normal samples that were mistakenly identified as violations by the primary model, and determining the corresponding sub-class violation label for each illegal audio and video data entry, resulting in a more detailed classification of the illegal audio and video data. Sub-classes of violation types are manually defined, and based on these definitions, each illegal audio and video data entry obtained from the secondary filtering is labeled with its corresponding sub-class violation label. Because of the pre-screening by the primary analysis network, the secondary analysis network actually uses fewer samples, consuming minimal resources, and the larger model improves overall recall and precision.
[0097] Three-level analysis network C:
[0098] like Figure 6 As shown, the three-level analysis network consists of an image feature extraction layer, a speech feature extraction layer, and a similarity comparison layer.
[0099] The process of extracting implicit features from each piece of illegal audio and video data using a three-level analysis network, and calculating similarity between these features and the features in the blacklist and whitelist feature libraries corresponding to the subclass illegal tags, to obtain the similarity comparison results, can specifically include:
[0100] ① The image feature extraction layer of the three-level analysis network extracts the implicit features of video frames in each illegal audio and video data to obtain the general features of the frames;
[0101] ② The speech feature extraction layer of the three-level analysis network extracts the implicit features of the speech signal in each illegal audio and video data to obtain the general speech features;
[0102] ③ The similarity calculation layer of the three-level analysis network calculates cosine similarity between the features in the blacklist and whitelist feature libraries corresponding to the subclass violation tags based on the common features of the image and the common features of the voice, and obtains the similarity comparison results.
[0103] Specifically, the image feature extraction layer can use techniques such as CLIP (Contrastive Language-Image Pre-training) to extract implicit features from video frames in each piece of illegal audio and video data, obtaining general features for the frames. The speech feature extraction layer can use techniques such as Hidden-Unit BERT (Bi-direction Encoder Representations from Transformer) to extract implicit features from speech signals in each piece of illegal audio and video data, obtaining general speech features.
[0104] The blacklist feature library records common and general features of the collected illegal audio and video data, while the whitelist feature library records common and general features of the collected non-illegal audio and video data. Common and easily identifiable images / audio are extracted using the above method. The similarity between the general visual and audio features and the features in the blacklist and whitelist feature libraries corresponding to the sub-category of illegal tags is calculated using cosine similarity comparison. If the similarity to features in the blacklist feature library is high (exceeding a preset threshold), then all audio and video data similar to the blacklist are identified as illegal audio and video data, i.e., they fail the review. If the similarity to features in the whitelist feature library is high (exceeding a preset threshold), then all audio and video data similar to the whitelist are identified as non-illegal audio and video data, i.e., they pass the review. In this case, no manual review is required, reducing review costs. If the similarity comparison result calculated does not exceed the preset threshold with each feature in the blacklist feature library and each feature in the whitelist feature library, that is, the similarity between the features in the blacklist feature library and the whitelist feature library corresponding to the sub-category violation tag is low, then each audio and video data can be sent to the manual review module, and the feedback result of the manual review module can be determined as the review result.
[0105] It is foreseeable that in practical applications, the features in the blacklist and whitelist feature libraries will be updated and adjusted at any time, allowing for manual additions and deletions. When a new type of violation is discovered, a feature can be manually added to the blacklist feature library, or when a new type of compliance is discovered, a feature can be manually added to the whitelist feature library. This eliminates the need to retrain the review model with deep learning after each new type of violation; only adjustments to the features in the blacklist and whitelist feature libraries are required. Therefore, as the application time increases and the amount of audio and video data reviewed grows, the features in the blacklist and whitelist feature libraries will gradually become richer, and the amount of audio and video data requiring manual review will gradually decrease.
[0106] In some embodiments of this application, the process of training an audio and video review model may include:
[0107] The first step is to obtain training audio and video data, which is labeled with corresponding review result information.
[0108] The second step is to input the training audio and video data into the preset initial audio and video review model to obtain the review results of the training audio and video data output by the initial audio and video review model.
[0109] The third step is to train an initial audio and video review model with the goal of ensuring that the review results of the training audio and video data are consistent with the corresponding review result information labeled in the training audio and video data.
[0110] Step 4: When the initial audio and video review model meets the preset training conditions, the trained initial audio and video review model is used as the audio and video review model.
[0111] Specifically, after acquiring a large amount of training audio and video data, this data can be input into a pre-defined initial audio and video review model for training. The goal is to ensure that the review results of the training audio and video data match the corresponding review result information labeled in the training data. During training, the parameters of each network structure are continuously adjusted and corrected.
[0112] When the initial audio and video review model meets the preset training conditions, the trained initial audio and video review model can be used as the audio and video review model. This audio and video review model can quickly and accurately determine the violation situation and violation type of the audio and video data, and review the audio and video data.
[0113] The audio and video review device provided in the embodiments of this application is described below. The audio and video review device described below can be referred to in correspondence with the audio and video review method described above.
[0114] See Figure 7 , Figure 7 This is a structural block diagram of an audio and video review device disclosed in an embodiment of this application.
[0115] like Figure 7 As shown, the audio and video review device may include:
[0116] The audio and video acquisition module 110 is used to acquire audio and video data to be reviewed;
[0117] The model determination module 120 is used to determine the audio and video review model. The audio and video review model includes a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network is used to perform category analysis and filtering on the input audio and video data to be reviewed according to major category violation tags, obtain the first-level analysis results, and generate each violation audio and video dataset corresponding to the major category violation tags. The second-level analysis network is used to perform detailed category analysis and filtering on each violation audio and video data in each violation audio and video dataset, obtain the second-level analysis results, and determine the sub-category violation tags corresponding to each violation audio and video data. The third-level analysis network is used to extract the implicit features of each violation audio and video data and perform similarity calculation with the features in the blacklist feature library and whitelist feature library corresponding to the sub-category violation tags to obtain the similarity comparison results. The result processing network is used to determine the review result of the audio and video data to be reviewed based on the first-level analysis results, the second-level analysis results, and the similarity comparison results.
[0118] The review result module 130 is used to input the audio and video data to be reviewed into the audio and video review model and obtain the review result of the audio and video data to be reviewed output by the audio and video review model.
[0119] As can be seen from the above technical solutions, the audio and video review method, apparatus, device and readable storage medium provided in this application embodiment obtains the audio and video data to be reviewed, determines the audio and video review model, inputs the audio and video data to be reviewed into the audio and video review model, and obtains the review result of the audio and video data to be reviewed output by the audio and video review model. The audio and video review model includes a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network is used to perform category analysis and filtering on the input audio and video data to be reviewed according to major violation tags, obtain the first-level analysis results, and generate each violation audio and video dataset corresponding to each major violation tag. The second-level analysis network is used to perform detailed category analysis and filtering on each violation audio and video data in each violation audio and video dataset, obtain the second-level analysis results, and determine the sub-category violation tags corresponding to each violation audio and video data. The third-level analysis network is used to extract the implicit features of each violation audio and video data and calculate the similarity with the features in the blacklist feature library and whitelist feature library corresponding to the sub-category violation tags to obtain the similarity comparison results. The result processing network is used to determine the review result of the audio and video data to be reviewed based on the first-level analysis results, the second-level analysis results, and the similarity comparison results.
[0120] This application employs a cascaded review process, utilizing a three-tiered review network to review audio and video data. The first-tier analysis network is a small model with high recall and low precision, filtering out most normal audio and video data through broad category analysis, retaining potentially inappropriate data for the second-tier analysis network. The second-tier analysis network is a large model with high recall and high precision, aiming to finely classify violation types. It further filters potentially inappropriate audio and video data through fine category analysis, resulting in a more detailed classification of the inappropriate audio and video data. The third-tier analysis network is an implicit feature extraction model, aiming to implicitly encode images and audio signals using an unsupervised model, and then calculate similarity comparison results by comparing them with features in blacklist and whitelist feature libraries.
[0121] This application designs a cascaded model structure, which is more cost-effective and efficient in terms of machine review. It innovatively uses features from the blacklist and whitelist feature libraries corresponding to the sub-category violation tags to enable automatic machine review of audio and video data, reducing the pressure and cost of manual review and improving review efficiency.
[0122] In some embodiments of this application, the first-level analysis network of the audio and video review model consists of a depthwise convolutional network with an inverted residual structure and a pointwise convolutional layer;
[0123] The depthwise convolutional network is used to perform depthwise convolution on the input audio and video data to be reviewed according to the parameters corresponding to the major categories of violation tags, and uses lightweight filtering to obtain the information values of each channel.
[0124] Pointwise convolutional layers are used to construct linear combinations based on the information values of each channel, and determine the illegal audio and video datasets corresponding to each major category of illegal labels according to the filtering range corresponding to each major category of illegal labels.
[0125] In some embodiments of this application, the secondary analysis network of the audio and video review model consists of a residual convolutional network, a pooling processing layer, and a result output layer;
[0126] Residual convolutional networks are used to extract features from each illegal audio and video data in each illegal audio and video dataset to obtain the feature information of each illegal audio and video data in each illegal audio and video dataset;
[0127] The pooling layer is used to perform dimensionality reduction pooling on the feature information of each illegal audio and video data in each illegal audio and video dataset, and generate integrated features of each illegal audio and video data in each illegal audio and video dataset;
[0128] The output layer is used to perform secondary violation screening based on the integrated features of each illegal audio and video data in each illegal audio and video dataset, to obtain secondary analysis results and determine the subclass violation label corresponding to each illegal audio and video data.
[0129] In some embodiments of this application, the three-level analysis network of the audio and video review model consists of an image feature extraction layer, a speech feature extraction layer, and a similarity comparison layer;
[0130] The image feature extraction layer is used to extract the implicit features of video frames in each illegal audio and video data to obtain general features of the frames;
[0131] The speech feature extraction layer is used to extract the implicit features of the speech signal in each illegal audio and video data to obtain general speech features;
[0132] The similarity calculation layer is used to calculate the cosine similarity between features in the blacklist and whitelist feature libraries corresponding to subclass violation tags, based on general features of the image and general features of the voice, to obtain the similarity comparison results.
[0133] In some embodiments of this application, the audio and video review device may further include a model training module;
[0134] The model training module trains the audio and video review model, including:
[0135] Acquire training audio and video data, which are labeled with corresponding review result information;
[0136] Input the training audio and video data into the preset initial audio and video review model to obtain the review results of the training audio and video data output by the initial audio and video review model;
[0137] The goal is to train an initial audio and video review model that matches the review results of the training audio and video data with the corresponding review result information labeled in the training audio and video data.
[0138] When the initial audio and video review model meets the preset training conditions, the trained initial audio and video review model will be used as the audio and video review model.
[0139] The audio and video review device provided in this application embodiment can be applied to audio and video review equipment. Figure 8 The hardware structure block diagram of the audio and video review equipment is shown. Figure 8 The hardware structure of the audio and video review equipment may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0140] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;
[0141] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0142] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;
[0143] The memory stores a program, which the processor can call. The program is used for:
[0144] Obtain audio and video data pending review;
[0145] An audio-visual review model is defined, comprising a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network performs broad-category analysis and filtering on the input audio-visual data to be reviewed according to major violation tags, obtaining first-level analysis results and generating each violation audio-visual dataset corresponding to each major violation tag. The second-level analysis network performs detailed-category analysis and filtering on each violation audio-visual data in each violation audio-visual dataset, obtaining second-level analysis results and determining the sub-category violation tags corresponding to each violation audio-visual data. The third-level analysis network extracts the implicit features of each violation audio-visual data and performs similarity calculation with features in the blacklist and whitelist feature libraries corresponding to the sub-category violation tags, obtaining similarity comparison results. The result processing network determines the review result of the audio-visual data to be reviewed based on the first-level analysis results, the second-level analysis results, and the similarity comparison results.
[0146] The audio and video data to be reviewed is input into the audio and video review model to obtain the review result of the audio and video review model on the audio and video data to be reviewed.
[0147] Optionally, the refined and extended functions of the program can be referred to the above description.
[0148] This application embodiment also provides a readable storage medium that can store a program suitable for execution by a processor, the program being used for:
[0149] Obtain audio and video data pending review;
[0150] An audio-visual review model is defined, comprising a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network performs broad-category analysis and filtering on the input audio-visual data to be reviewed according to major violation tags, obtaining first-level analysis results and generating each violation audio-visual dataset corresponding to each major violation tag. The second-level analysis network performs detailed-category analysis and filtering on each violation audio-visual data in each violation audio-visual dataset, obtaining second-level analysis results and determining the sub-category violation tags corresponding to each violation audio-visual data. The third-level analysis network extracts the implicit features of each violation audio-visual data and performs similarity calculation with features in the blacklist and whitelist feature libraries corresponding to the sub-category violation tags, obtaining similarity comparison results. The result processing network determines the review result of the audio-visual data to be reviewed based on the first-level analysis results, the second-level analysis results, and the similarity comparison results.
[0151] The audio and video data to be reviewed is input into the audio and video review model to obtain the review result of the audio and video review model on the audio and video data to be reviewed.
[0152] Optionally, the refined and extended functions of the program can be referred to the above description.
[0153] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0154] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0155] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio and video review method, characterized in that, include: Obtain audio and video data pending review; An audio / video review model is defined, comprising a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network performs broad-category analysis and filtering on the input audio / video data to be reviewed according to major violation tags, obtaining first-level analysis results and generating datasets of violating audio / video data corresponding to each major violation tag. The second-level analysis network performs detailed-category analysis and filtering on each piece of violating audio / video data in each dataset, obtaining second-level analysis results and determining the sub-category violation tags corresponding to each piece of violating audio / video data. The third-level analysis network extracts the violation tags from each piece of data. The implicit features of the video data are compared with the features in the blacklist and whitelist feature libraries corresponding to the subclass violation tags to obtain similarity comparison results. The result processing network is used to determine the review result of the audio and video data to be reviewed based on the first-level analysis result, the second-level analysis result, and the similarity comparison result. The first-level analysis network consists of a depthwise convolutional network with an inverted residual structure and a pointwise convolutional layer. The second-level analysis network consists of a residual convolutional network, a pooling layer, and a result output layer. The third-level analysis network consists of an image feature extraction layer, a speech feature extraction layer, and a similarity comparison layer. The audio and video data to be reviewed is input into the audio and video review model to obtain the review result of the audio and video review model on the audio and video data to be reviewed.
2. The method according to claim 1, characterized in that, The review result for the audio and video data to be reviewed is determined based on the primary analysis result, the secondary analysis result, and the similarity comparison result, including: The review results of each audio and video data item in the audio and video data to be reviewed that has a non-violation result in the first-level analysis result and / or the second-level analysis result are determined as approved; The first-level analysis result and the second-level analysis result are deemed to be violations, and the similarity comparison result is the review result of each audio and video data whose feature similarity with the feature in the blacklist feature library exceeds a preset threshold, which is determined to be the review failure; If the first-level analysis result and the second-level analysis result are deemed as violations, and the similarity comparison result is the review result of each audio and video data whose feature similarity with the whitelist feature library exceeds a preset threshold, then the review is deemed as passed.
3. The method according to claim 2, characterized in that, Also includes: The audio and video data whose first-level analysis results and second-level analysis results are deemed to be in violation, and whose similarity comparison results are all within the preset thresholds for feature similarity with the features in the blacklist feature library and the whitelist feature library, are sent to the manual review module, and the feedback result of the manual review module is determined as the review result.
4. The method according to claim 1, characterized in that, The process by which the primary analysis network performs category analysis and filtering on the input audio and video data to be reviewed according to major violation tags, obtains primary analysis results, and generates each violation audio and video dataset corresponding to the major violation tags includes: The depthwise convolutional network of the first-level analysis network performs depthwise convolution on the input audio and video data to be reviewed according to the parameters corresponding to the major categories of violation tags, and uses lightweight filtering to obtain the information values of each channel. The pointwise convolutional layers of the first-level analysis network are constructed by linearly combining the information values of each channel, and the violation audio and video datasets corresponding to each major category of violation label are determined according to the filtering range corresponding to each major category of violation label.
5. The method according to claim 1, characterized in that, The secondary analysis network performs detailed category analysis and filtering on each piece of illegal audio and video data in each illegal audio and video dataset, obtains secondary analysis results, and determines the sub-category violation label corresponding to each piece of illegal audio and video data. This process includes: The residual convolutional network of the secondary analysis network extracts features from each illegal audio and video data in each illegal audio and video dataset to obtain the feature information of each illegal audio and video data in each illegal audio and video dataset. The pooling layer of the secondary analysis network performs dimensionality reduction pooling on the feature information of each illegal audio and video data in each illegal audio and video dataset to generate integrated features of each illegal audio and video data in each illegal audio and video dataset. The output layer of the secondary analysis network performs secondary violation screening based on the integrated features of each illegal audio and video data in each illegal audio and video dataset, obtains the secondary analysis results, and determines the sub-class violation label corresponding to each illegal audio and video data.
6. The method according to claim 1, characterized in that, The process of extracting implicit features from each piece of illegal audio and video data by the three-level analysis network, and calculating similarity between these features and the features in the blacklist and whitelist feature libraries corresponding to the sub-class illegal tags, to obtain the similarity comparison results, includes: The image feature extraction layer of the three-level analysis network extracts the implicit features of the video frames in each illegal audio and video data to obtain the general features of the frames; The speech feature extraction layer of the three-level analysis network extracts the implicit features of the speech signal in each illegal audio and video data to obtain general speech features; The similarity calculation layer of the three-level analysis network calculates cosine similarity between the common features of the image and the common features of the voice and the features in the blacklist feature library and whitelist feature library corresponding to the subclass violation label, and obtains the similarity comparison result.
7. The method according to claim 1, characterized in that, The process of training the audio and video review model includes: Acquire training audio and video data, which are labeled with corresponding review result information; The training audio and video data is input into a preset initial audio and video review model to obtain the review result of the training audio and video data output by the initial audio and video review model. The initial audio and video review model is trained with the goal of ensuring that the review results of the training audio and video data are consistent with the corresponding review result information labeled in the training audio and video data. When the initial audio and video review model meets the preset training conditions, the trained initial audio and video review model is used as the audio and video review model.
8. An audio and video review device, characterized in that, include: The audio and video acquisition module is used to acquire audio and video data to be reviewed. The model determination module is used to determine the audio and video review model. The audio and video review model includes a first-level analysis network, a second-level analysis network, a third-level analysis network, and a result processing network. The first-level analysis network is used to perform category analysis and filtering on the input audio and video data to be reviewed according to major violation tags, obtaining first-level analysis results and generating each violation audio and video dataset corresponding to each major violation tag. The second-level analysis network is used to perform detailed category analysis and filtering on each violation audio and video data in each violation audio and video dataset, obtaining second-level analysis results and determining the sub-category violation tags corresponding to each violation audio and video data. The third-level analysis network is used to extract... The implicit features of each piece of illegal audio and video data are compared with the features in the blacklist and whitelist feature libraries corresponding to the sub-category of illegal tags to obtain similarity comparison results. The result processing network is used to determine the review result of the audio and video data to be reviewed based on the first-level analysis result, the second-level analysis result, and the similarity comparison result. The first-level analysis network consists of a depthwise convolutional network with an inverted residual structure and a pointwise convolutional layer. The second-level analysis network consists of a residual convolutional network, a pooling layer, and a result output layer. The third-level analysis network consists of an image feature extraction layer, a speech feature extraction layer, and a similarity comparison layer. The review result module is used to input the audio and video data to be reviewed into the audio and video review model, and obtain the review result of the audio and video review model on the audio and video data to be reviewed.
9. An audio and video review device, characterized in that, Including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the audio and video review method as described in any one of claims 1-7.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the audio and video review method as described in any one of claims 1-7.
Citation Information
Patent Citations
Teaching video auditing method and device, equipment and medium
CN112860943A
Audio audit processing method and device, equipment and storage medium
CN114168788A