A content review method and device, equipment, storage medium
By combining a dual AI review model with an MD5 database, video content is quickly screened and processed by a deep convolutional neural network, solving the problems of slow video content review speed and resource waste in cloud storage, and achieving efficient and accurate video content recognition and adaptive capabilities.
Patent Information
- Application Number
- CN202210002751.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-04
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-01-04
AI Technical Summary
In existing technologies, cloud-based video content review suffers from slow review speed, low efficiency, and waste of human and material resources. Furthermore, AI algorithm models cannot learn and adapt to new prohibited videos, and manual review consumes a large amount of resources.
A dual AI review model is used to review video and text data separately. The initial screening is combined with the MD5 database. The number of neural network layers is adjusted using deep convolutional neural networks and self-organizing map algorithms to achieve efficient extraction and recognition of video frames.
It improves the speed and accuracy of video content review, saves human and material resources, dynamically adapts to the identification of prohibited videos, and enhances the adaptability and efficiency of the review system.
Smart Images

Figure CN116467487B_ABST
Abstract
Description
Technical Field
[0001] This application pertains to the application of artificial intelligence technology, and is used in areas such as improving the efficiency and accuracy of video content review in cloud storage applications. Background Technology
[0002] Cloud storage bridges mobile and PC platforms, enabling file sharing but also facilitating the spread of prohibited information. Cloud storage providers have an obligation to verify the legality and security of stored files. Currently, a combination of AI and human review remains the mainstream approach for reviewing video content. However, human review suffers from slow speed, low efficiency, and wasted human and material resources. Summary of the Invention
[0003] In view of this, embodiments of this application provide a content review method, apparatus, device, and storage medium.
[0004] The technical solution of this application embodiment is implemented as follows:
[0005] In a first aspect, embodiments of this application provide a content review method applied to cloud storage. The method includes: acquiring review data of a video to be reviewed; wherein the review data includes text data and video data of the video to be reviewed; reviewing the text data using a first AI review model to obtain a first review result; reviewing the video data using a second AI review model to obtain a second review result; and determining the review result of the video to be reviewed based on the first review result and the second review result.
[0006] Secondly, embodiments of this application provide a content review apparatus, the apparatus comprising: an acquisition module, configured to acquire review data of a video to be reviewed; wherein the review data includes text data and video data of the video to be reviewed; a first review module, configured to review the text data using a first AI review model to obtain a first review result; a second review module, configured to review the video data using a second AI review model to obtain a second review result; and a first determination module, configured to determine the review result of the video to be reviewed based on the first review result and the second review result.
[0007] Thirdly, this application provides a content review device, the device comprising: a memory for storing executable instructions; and a processor for executing the executable instructions stored in the memory to implement the content review method provided in this application.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium for storing a computer program thereon, the computer program being used to implement the content review method provided in embodiments of this application when executed by a processor.
[0009] In this embodiment, firstly, text data and video data of the video to be reviewed are acquired as review data; secondly, the text data is reviewed using a first AI review model to obtain a first review result; then, the video data is reviewed using a second AI review model to obtain a second review result; finally, based on the first review result and the second review result, the review result of the video to be reviewed is determined. Thus, firstly, by dividing the review data of the video to be reviewed into text data and video data for separate review and obtaining review scores for each, and then weighting and fusing them, the accuracy of the review is improved; secondly, by reviewing the text data using the first AI review model and the video data using the second AI review model, the manual review step is eliminated, which not only significantly improves the speed and efficiency of the review but also saves human and material resources. Attached Figure Description
[0010] Figure 1 A schematic diagram illustrating the implementation process of a content review method provided in this application embodiment;
[0011] Figure 2 A schematic diagram illustrating the implementation process of a content review method provided in this application embodiment;
[0012] Figure 3 A schematic diagram illustrating the implementation process of a content review method provided in this application embodiment;
[0013] Figure 4A A schematic diagram illustrating the implementation process of a content review method provided in this application embodiment;
[0014] Figure 4B This is a schematic diagram illustrating the implementation process of a method for adjusting the number of convolutional kernels and the number of layers in a neural network, as provided in an embodiment of this application.
[0015] Figure 4C This is a schematic diagram of the convolutional kernel connections in the current convolutional layer provided in an embodiment of this application;
[0016] Figure 4D A schematic diagram illustrating the implementation process of a method for adjusting the number of convolutional kernels and the number of layers in a neural network, provided in an embodiment of this application;
[0017] Figure 4E This is a schematic diagram illustrating the data processing flow of the AI reviewer provided in an embodiment of this application.
[0018] Figure 5 A schematic diagram illustrating the implementation process of a content review method provided in this application embodiment;
[0019] Figure 6 A schematic diagram of the composition structure of a content review device provided in this application embodiment;
[0020] Figure 7 This is a schematic diagram of the hardware entity of a content review device provided in an embodiment of this disclosure. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0024] As a digital fingerprint of a file, each video has a unique MD5 hash. For cloud storage applications or video websites, establishing a continuously updated MD5 database of inappropriate videos and automatically reviewing the MD5 hashes of uploaded videos is crucial. Videos falling into the category of inappropriate content can be blocked or deleted, and the uploader can be blacklisted as a dangerous user. Currently, the review standards and establishment process for inappropriate videos primarily rely on continuously enriching the databases of prohibited words, audio, and videos, monitoring dangerous users, and establishing reward mechanisms for reporting violations. In its "Clean Internet Campaign," Baidu Cloud directly replaced inappropriate videos with 8-second promotional videos. Baidu Cloud reviews video content by combining AI algorithms with human review after examining the MD5 hash. From an economic and technical perspective, establishing a continuously updated MD5 database and using a review scheme that combines AI and human review remains the mainstream choice for reviewing video content.
[0025] If artificial intelligence is used appropriately for review, eliminating the need for manual review, it will not only significantly improve the speed and efficiency of the review process, but also save considerable human resources costs.
[0026] In view of this, embodiments of this application provide a content review method applied to cloud storage, such as... Figure 1 As shown, the method includes:
[0027] Step S101: Obtain the review data of the video to be reviewed; wherein, the review data includes the text data and video data of the video to be reviewed;
[0028] Here, the video to be reviewed includes the text information contained in the video, audio files, and each frame of the video. The audio files in the video need to be converted into text information.
[0029] Step S102: Review the text data using the first AI review model to obtain the first review result;
[0030] Here, the first AI review model includes a text AI reviewer. The first review result is represented by a review score (Score 1). In some embodiments, the algorithm in the text AI reviewer is a deep convolutional neural network algorithm, which includes convolutional layers, pooling layers, fully connected layers, and dropout layers to prevent overfitting. In some embodiments, text data enters the text AI reviewer for review. The text AI reviewer calculates a text-audio review score (Score 1) by performing a global comparison with text data in the prohibited video database. Videos that fail the review are rejected and uploaded to backup storage, replaced with warning images, and the videos are updated in the prohibited video database. In some embodiments, a text-audio review score (score 1) greater than or equal to 0.75 indicates that the review has failed.
[0031] Step S103: Review the video data using the second AI review model to obtain a second review result;
[0032] Here, the second review result is represented by the review score Score2. The second AI review model includes N AI reviewers, where N is an integer greater than or equal to 1. When the second AI review model includes two or more AI reviewers, the two or more AI reviewers can be the same or different.
[0033] In some embodiments, the second AI review model includes a first AI reviewer, a second AI reviewer, and a third AI reviewer, wherein the review level of the third AI reviewer is higher than that of the first and second AI reviewers.
[0034] Step S104: Based on the first review result and the second review result, determine the review result of the video to be reviewed.
[0035] Here, the review score of the first review result (score1) and the review score of the second review result (score2) are weighted and calculated to obtain the violation degree O using formula (1), which is shown below:
[0036]
[0037] Where O represents the degree of violation, ∧ represents the intersection, ∨ represents the union, score1≤0.30∧score2≤0.30 means that score1 is less than or equal to 0.3 and score2 is less than or equal to 0.3, score1≥0.75∨score2≥0.75 means that score1 is greater than or equal to 0.75 or score2 is greater than or equal to 0.75, and 0.5×(score1+1.2score2) means that the degree of violation O is obtained by weighted calculation of score1 and score2.
[0038] In some embodiments, a violation score of 0.5 or higher is considered a failure to pass review, while a score below 0.5 is considered a pass. Passed content is added to a graylist video library, and the file is also backed up and stored on storage media. Failed videos are refused upload to the backup storage, replaced with warning images, and the videos are updated in the prohibited video library.
[0039] In some embodiments, the system is configured to periodically (every 7 days) re-examine and verify the content in the graylist video library, and re-execute the review process from steps S101 to S104. Video content that passes the review process twice consecutively is added to the whitelist video library, and the system is configured to periodically (every 30 days) re-examine and verify the content in the whitelist video library.
[0040] In this embodiment, firstly, text data and video data of the video to be reviewed are acquired as review data; secondly, the text data is reviewed using a first AI review model to obtain a first review result; then, the video data is reviewed using a second AI review model to obtain a second review result; finally, based on the first review result and the second review result, the review result of the video to be reviewed is determined. Thus, firstly, by dividing the review data of the video to be reviewed into text data and video data for separate review and obtaining review scores for each, and then weighting and fusing them, the accuracy of the review is improved; secondly, by reviewing the text data using the first AI review model and the video data using the second AI review model, the manual review step is eliminated, which not only significantly improves the speed and efficiency of the review but also saves human and material resources.
[0041] This application provides a content review method, such as... Figure 2 As shown, the method includes:
[0042] Step S201: Divide the video to be uploaded into segments to obtain at least one video data block;
[0043] Here, video splitting can be achieved using a video splitter (Boilsoft Video Splitter, BVS). BVS can directly cut video files without encoding, and it is fast and easy to operate.
[0044] Step S202: Determine the MD5 data value of each video data block;
[0045] Here, the MD5 data value is obtained by the MD5 Message Digest Algorithm (MD5MDA), which is a cryptographic hash function that produces a 128-bit (16-byte) hash value to ensure the integrity and consistency of transmitted information.
[0046] Step S203: Based on the MD5 data value of each video data block and the MD5 data value in the preset MD5 database, determine the first review result of each video data block; wherein, the first review result includes pass or fail;
[0047] Here, when the first review result is "not passed," the video is deemed to be prohibited and added to the prohibited video library.
[0048] Step S204: If the first review result of all video data blocks in the video to be uploaded is passed, the video to be uploaded is determined as the video to be reviewed; wherein, the data to be reviewed includes the text data and video data of the video to be reviewed;
[0049] Here, if any video data block in the video to be uploaded fails the first review, the video to be uploaded is determined to be prohibited and added to the prohibited video library.
[0050] Step S205: Review the text data using the first AI review model to obtain the first review result;
[0051] Step S206: Review the video data using the second AI review model to obtain a second review result;
[0052] Step S207: Based on the first review result and the second review result, determine the review result of the video to be reviewed.
[0053] Here, steps S205 to S207 can be understood by referring to steps S102 to S104 above.
[0054] In this embodiment, the video to be uploaded is segmented to obtain at least one video data block. The MD5 hash of each video data block is determined. Based on the MD5 hash of each video data block and the MD5 hash of a pre-set MD5 database, a first review result is determined for each video data block. The video to be uploaded where the first review result of all video data blocks is passed is identified as the video to be reviewed. In this way, preliminary video review through simple MD5 value comparison reduces the workload of the subsequent AI review model, thereby improving review speed and efficiency.
[0055] This application provides a content review method, such as... Figure 3 As shown, the method includes:
[0056] Step S301: Based on text information extraction technology, determine the first text data of the video to be reviewed;
[0057] Here, text information extraction technology is a technique for extracting specific information from text. Text information includes characters, words, phrases, sentences, paragraphs, or combinations of these specific units forming a whole. Text information extraction includes extracting noun phrases, names of people, place names, events, actions, etc., from text information. Text information extraction models include Hidden Markov Models, Maximum Entropy Markov Models, Conditional Random Fields, and Voting Perceptron Models. In some embodiments, text information extraction technology employs Optical Character Recognition (OCR).
[0058] Step S302: Based on speech recognition technology, determine the second text data of the video to be reviewed;
[0059] Here, speech recognition technology is also called Automatic Speech Recognition (ASR), which aims to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. In some embodiments, the recognition methods in speech recognition technology include, but are not limited to, the following: (1) Linguistic and acoustic-based methods: This is the earliest method applied to speech recognition. (2) Stochastic model method: This method uses steps such as feature extraction, template training, template classification, and template judgment to recognize speech. The techniques involved in the stochastic model method include: Dynamic Time Warping (DTW), Hidden Markov Model (HMM) theory, and Vector Quantization (VQ) technology. (3) Neural network: This is a method that simulates human neural activity and has some human characteristics, such as automatic adaptation and autonomous learning. In addition, neural networks have strong classification and mapping capabilities. Using neural networks in speech recognition significantly improves the efficiency of speech recognition. (4) Probabilistic syntax analysis method: a technique that can identify long passages and use knowledge of the corresponding level to identify different levels of knowledge.
[0060] Step S303: Based on the first text data and the second text data, determine the text data of the video to be reviewed;
[0061] Step S304: Based on the image recognition algorithm, determine the video data of the video to be reviewed;
[0062] Here, image recognition algorithms include extracting important features of an image and eliminating redundant information to recognize the image. The process of image recognition technology includes: (1) Information acquisition: Information acquisition includes converting information such as light or sound into electrical information through sensors, that is, acquiring the basic information of the research object and converting it into information that the machine can recognize through some method. (2) Image preprocessing: Image preprocessing includes operations such as denoising, smoothing, and transformation in image processing, which are used to enhance the important features of the image. (3) Feature extraction and selection: Feature extraction and selection includes extracting the image's own features and selecting useful features in the image. (4) Classifier design: Classifier design refers to obtaining a recognition rule through training, and a feature classification can be obtained through the recognition rule, so that image recognition technology can achieve a high recognition rate. (5) Classification decision: Classification decision refers to classifying the object to be recognized in the feature space, so as to better identify which category the research object belongs to.
[0063] In some embodiments, the image feature extraction algorithm includes: Histogram of Oriented Gradient (HOG), Local Binary Pattern (LBP) algorithm, or Haar-like features algorithm.
[0064] In this embodiment, the LBP algorithm is used to extract image features from the video to be reviewed. Here, the LBP algorithm is an operator used to describe the local texture features of an image, and the extracted features are the local texture features of the image.
[0065] In some embodiments, the implementation of step S304 includes:
[0066] Step S3041: Divide each video frame in the video to be reviewed into N×N sub-regions; here, N is a number greater than 1, for example, the sub-region can be a 3×3 region or an 8×8 region, etc.
[0067] Step S3042: Determine the LBP value of each pixel in each sub-region;
[0068] Here, for example, each video frame in the video to be reviewed is divided into a 3×3 sub-region. Within a 3×3 sub-region, the grayscale values of the eight adjacent pixels are compared with the center pixel as a threshold. If the values of the surrounding pixels are greater than or equal to the center pixel value, the position of that pixel is marked as 1; otherwise, it is marked as 0. In this way, the comparison of the eight points in the 3×3 neighborhood generates an 8-bit binary number (usually converted to a decimal number, i.e., LBP code, of which there are 256 possible values), thus obtaining the LBP value of the center pixel within the sub-region.
[0069] Step S3043: Determine the statistical histogram of each sub-region based on the LBP value of each pixel in each sub-region;
[0070] In some embodiments, after obtaining the statistical histogram of each sub-region, the histogram needs to be normalized.
[0071] Step S3044: Connect the histograms of each sub-region to obtain the feature vector of each video frame in the video to be reviewed;
[0072] Here, the feature vector of each video frame can be a 1×M matrix.
[0073] Step S3045: Based on the timestamp of the feature vector of each video frame in the video to be reviewed, all video frames are concatenated sequentially to obtain the video data of the video to be reviewed.
[0074] Step S305: Review the text data using the first AI review model to obtain the first review result;
[0075] Step S306: Review the video data using the second AI review model to obtain a second review result;
[0076] Step S307: Based on the first review result and the second review result, determine the review result of the video to be reviewed.
[0077] Here, steps S305 to S307 can be understood by referring to steps S102 to S104 above.
[0078] In this embodiment, the first text data of the video to be reviewed is determined based on text information extraction technology; the second text data of the video to be reviewed is determined based on speech recognition technology; the text data of the video to be reviewed is determined based on the first and second text data; and the video data of the video to be reviewed is determined based on an image recognition algorithm. This improves review speed and efficiency while saving human and material resources.
[0079] This application provides a content review method, such as... Figure 4A As shown, the method includes:
[0080] Step S401: Obtain the review data of the video to be reviewed; wherein, the review data includes the text data and video data of the video to be reviewed;
[0081] Step S402: Review the text data using the first AI review model to obtain the first review result; here, steps S401 to S402 can be understood by referring to steps S101 to S102.
[0082] Step S403: Extract the video data according to a preset first extraction ratio to obtain the first video data;
[0083] Here, a Cartesian coordinate system is established with the timeline of the video frames as the x-axis and the probability of a video frame being extracted as the y-axis. The AI reviewer randomly extracts video frames according to a normal distribution, meaning that the probability of a video frame being extracted follows a normal distribution, with video frames in the middle of the timeline having a higher probability of being extracted and those at the ends of the timeline having a lower probability. In some embodiments, the preset first extraction ratio is 0.236, and the dynamic adjustment range of the first extraction ratio is 0.20-0.50.
[0084] Step S404: Review the first video data through the first AI reviewer to obtain a preliminary review result; wherein, the preliminary review result includes pass, suspected, or fail;
[0085] Here, the first AI reviewer calculates the degree of violation O. tmp To obtain the preliminary review results, the violation rate is O. tmp The relationship between the results and the preliminary review is shown in formula (2):
[0086]
[0087] Among them, O tmp Indicates the degree of prohibition.
[0088] Step S405: Extract the first video data whose pre-review result is passed according to the preset second extraction ratio to obtain the second video data; wherein, the second extraction ratio is greater than the first extraction ratio;
[0089] In some embodiments, the preset second extraction ratio is 0.286, the dynamic adjustment range of the second extraction ratio is 0.20-0.50, and the second extraction ratio is always greater than the first extraction ratio.
[0090] Step S406: Review the second video data using the second AI reviewer to obtain a second review result.
[0091] Here, the second review result is represented by Score2.
[0092] In some embodiments, if the preliminary review result in step S404 is suspected, the review steps for the first video data are as follows:
[0093] Step S405a: Extract the first video data whose pre-review result is suspected according to a preset third extraction ratio to obtain the third video data; wherein, the third extraction ratio is greater than the second extraction ratio.
[0094] In some embodiments, the preset third extraction ratio is 0.512, and the dynamic adjustment range of the third extraction ratio is 0.50-0.85.
[0095] Step S405b: The third video data is reviewed by the third AI reviewer to obtain a second review result.
[0096] Here, the second review result is represented by Score2.
[0097] In some embodiments, if the pre-screening result in step S404 is unsuccessful, the first video data will be rejected from being uploaded to backup storage, and a warning image will be used as a replacement while the video is updated to the prohibited video library.
[0098] In some embodiments, the review level of the third AI reviewer is higher than that of the second AI reviewer, and the review level of the second AI reviewer is higher than that of the first AI reviewer. The review level is determined based on a preset sampling ratio, and the higher the preset sampling ratio, the higher the review level.
[0099] Step S407: Based on the first review result and the second review result, determine the review result of the video to be reviewed.
[0100] In some embodiments, the review score (score1) of the first review result and the review score (score2) of the second review result are weighted to obtain a violation score (O). A violation score (O) greater than or equal to 0.5 is considered a failure to pass the review, while a score less than 0.5 is considered a pass. Content that passes the review is added to a graylist video library, and the file is also backed up and stored on storage media. Videos that fail the review are refused upload to the backup storage, replaced with warning images, and the videos are updated in the prohibited video library.
[0101] In some embodiments, because the prohibited video library is constantly being updated dynamically, the system will periodically (once every 7 days) review and verify the content in the gray list video library again, re-execute the system review process of steps S401 to S407, and add video content that passes the review process twice in a row to the white list video library. The system will periodically (once every 30 days) review and verify the content in the white list video library again.
[0102] In this embodiment, firstly, a first video data set is obtained by extracting videos to be reviewed at a preset first extraction ratio. This first video data is then reviewed by a first AI reviewer to obtain a pre-review result. Secondly, the first video data set that passed the pre-review is extracted at a preset second extraction ratio to obtain second video data. This second video data is then reviewed by a second AI reviewer to obtain a second review result. Simultaneously, the first video data set that was deemed suspicious in the pre-review is extracted at a preset third extraction ratio to obtain third video data. This third video data is then reviewed by a third AI reviewer to obtain a second review result. Finally, based on the first and second review results, the review result of the video to be reviewed is determined. This method, by using different AI reviewers in cooperation, firstly, improves the speed of video review; secondly, it improves the accuracy of video review; and thirdly, it avoids the waste of human and material resources associated with manual review.
[0103] In some embodiments, the algorithms of the first, second, and third AI reviewers are deep convolutional neural networks, wherein the number of convolutional neural network layers in the third AI reviewer is greater than the number of convolutional neural network layers in the first and second AI reviewers.
[0104] Here, the number of convolutional neural network layers in the first, second, and third AI reviewers is determined based on preset first, second, and third extraction ratios. Because the preset third extraction ratio is greater than the first and second extraction ratios, the number of convolutional neural network layers in the third AI reviewer is greater than the number of convolutional neural network layers in the first and second AI reviewers.
[0105] In some embodiments, the number of neural network layers in the first, second, and third AI reviewers is adjusted based on a self-organizing map algorithm.
[0106] Figure 4B This application provides a schematic diagram illustrating the implementation process of a method for adjusting the number of convolutional kernels and the number of layers in a neural network, as illustrated in the embodiments of this application. Figure 4B As shown, the method includes:
[0107] Step S40: Obtain the connection status of the convolutional kernels in the current convolutional layer and the connection iteration number of two interconnected convolutional kernels;
[0108] Here, the threshold for the number of convolutional kernels is set, for example, the threshold for the number of convolutional kernels is set to 10.
[0109] Here, examples are used to understand the connection status of the convolutional kernels in the current convolutional layer and the connection iteration number of two interconnected convolutional kernels.
[0110] For example, Figure 4CThis is a schematic diagram of the convolutional kernel connections in the current convolutional layer, as shown below. Figure 4C As shown, there are 4 convolutional kernels in the current convolutional layer. The two ends of convolutional kernel 1 are connected to convolutional kernel 2 and convolutional kernel 3 respectively. The two ends of convolutional kernel 2 are connected to convolutional kernel 1 and convolutional kernel 3 respectively. The two ends of convolutional kernel 3 are connected to convolutional kernel 1 and convolutional kernel 4 respectively. The two ends of convolutional kernel 4 are connected to convolutional kernel 2 and convolutional kernel 3 respectively. However, convolutional kernel 1 and convolutional kernel 4 are not connected.
[0111] Here, two matrices E and A are given below. Matrix E shows the connection status of the convolution kernels, and matrix A shows the iteration number in which the connection exists. From matrix A, we can see that the connection between convolution kernels 1 and 2 has undergone two iterations, while the connection between convolution kernels 2 and 4 has had 0 iterations. A value of 0 in matrix A can represent either unconnected convolution kernels or connections formed only in the most recent iterations. From matrix E, we can see that no connection was established between convolution kernels 1 and 4, while convolution kernels 2 and 4 were connected in the most recent iteration.
[0112]
[0113] Step S41: When the number of connection iterations of the two interconnected convolutional kernels is less than the number of iterations threshold, the weight vector of the convolutional kernels in the current convolutional layer is determined based on the self-organizing map algorithm after iteration.
[0114] In some embodiments, a self-organizing map algorithm is used to update the weight vector and position of the convolution kernel. After t+1 iterations, the neighborhood function of the winning convolution kernel can be expressed as Equation (3), and the weight vector of the winning convolution kernel can be expressed as Equation (4). Equations (3) and (4) are shown below:
[0115]
[0116] Among them, H ij (t) represents the neighborhood function; t represents the iteration number; r i (t),r j σ(t) represents the position vectors of convolution kernels i and j; σ(t) represents the learning rate.
[0117]
[0118] Among them, W i (t+1) represents the updated weight vector of the convolution kernel after t+1 iterations; t represents the iteration number; n j (t) represents the number of convolution kernels connected to convolution kernel j; H ij (t) represents the neighborhood function; σ(t) represents the average of all feature vectors of x assigned to the convolution kernel j; σ(t) represents the learning rate.
[0119] Based on this, after updating the convolution kernel weight vector in each iteration, it is necessary to calculate the distance between the convolution kernel vectors using formula (5) and update the convolution kernel position using formula (6). Formulas (5) and (6) are shown below:
[0120]
[0121] Where represents the update rate of the convolution kernel position at iteration number t; δ ij (t) is the neighborhood function, representing the distance between convolution kernel i and convolution kernel j; γ represents a parameter that controls the shrinkage of the neighborhood.
[0122]
[0123] Where, r i (t+1) represents the position vector of convolution kernel i after t+1 iterations; r i (t),r j (t) represents the position vectors of convolution kernels i and j after t iterations; δ ij (t) is the neighborhood function, representing the distance between convolution kernel i and convolution kernel j; n j α(t) represents the number of convolution kernels connected to convolution kernel j; α(t) represents the update rate of the convolution kernel position after t iterations.
[0124] Step S42: Determine the quantization error of the convolution kernel based on the weight vector of the convolution kernel in the current convolutional layer;
[0125] Here, the quantization error (QE) of the convolution kernel can be solved according to formula (7), which is shown below:
[0126] qe = w i (t+1)-x (7);
[0127] Where qe represents the quantization error, w i (t+1) represents the weight vector of the winning convolution kernel after t+1 iterations; x represents the feature vector of the image assigned to convolution kernel j.
[0128] Step S43: Determine the number of convolutional kernels in the current convolutional layer based on the quantization error of the convolutional kernels;
[0129] Here, when the quantization error of the convolutional kernel is greater than the growth threshold, the convolutional kernel with the largest quantization error is split into two convolutional kernels. The new convolutional kernels maintain the same connections as the original convolutional kernels, and the two new convolutional kernels are also interconnected; when the quantization error of the convolutional kernels is less than the growth threshold, the number of convolutional kernels in the current convolutional layer is not increased.
[0130] Step S44: Determine the number of layers in the neural network based on the average quantization error algorithm and the number of convolutional kernels in the current convolutional layer.
[0131] Here, the Mean Quadratuer Error (MQE) can be solved according to Equation (8), which is shown below:
[0132]
[0133] Where mq represents the average quantization error; w i (t+1) represents the weight vector of the winning convolutional kernel after t+1 iterations; x represents the feature vector of the image assigned to convolutional kernel j; n represents the number of convolutional kernels in the current convolutional layer.
[0134] In some embodiments, if the average quantization error between two consecutive time steps, meq(t+1) - mqe(t), is less than ξ1, the iteration ends, and the current convolutional layer is the layer number of the neural network. Here, ξ1 (e.g., 1E-06) is a very small value. During the algorithm iteration process, if the average quantization error between two consecutive time steps is still greater than ξ1, a new convolutional layer needs to be added.
[0135] Figure 4D This is a schematic diagram illustrating the implementation process of a method for adjusting the number of layers in a neural network, as provided in an embodiment of this application. Figure 4D As shown, the method includes:
[0136] Step S40a: Obtain the connection status of the convolutional kernels in the current convolutional layer and the connection iteration number of two interconnected convolutional kernels;
[0137] Here, step S40a can be understood by referring to step S40 above.
[0138] Step S41a: When the number of connection iterations of the two interconnected convolutional kernels is greater than the iteration threshold, disconnect the two interconnected convolutional kernels and delete the convolutional kernels that are not connected to any convolutional kernels to obtain the number of convolutional kernels in the current convolutional layer.
[0139] Here, the winning and second winning convolutional kernels are obtained by the self-organizing map algorithm. If they are connected to each other, the existing connection iteration count is reset to 0. After each iteration, if the connection iteration count between convolutional kernel i and convolutional kernel j is greater than the iteration threshold, they are disconnected. At this time, if there is a convolutional kernel i that is not connected to any other convolutional kernel, convolutional kernel i is deleted.
[0140] Step S42a: Determine the number of layers in the neural network based on the number of layers in the convolutional layer where the current convolutional kernel is located.
[0141] Here, because the number of convolutional kernels in the current convolutional layer is smaller, the number of convolutional kernels is less than the threshold. Therefore, the number of neural network layers in the current convolutional layer is the final number of neural network layers.
[0142] During video review, the high-dimensional video data mapped to a low-dimensional space sometimes results in overly scattered or dense clustering of the output, producing distorted variations that do not conform to the input pattern, making the reviewer's output difficult to identify. In this embodiment, the number of neural network layers in the first, second, and third AI reviewers is adjusted based on a self-organizing map algorithm. This allows the number of neural network layers in the first, second, and third AI reviewers to be automatically adjusted according to different videos to be reviewed. Using this online, variable-structure AI reviewer, the network structure can be dynamically changed during training to mitigate these distorted variations, improving the accuracy of prohibited video classification and identification.
[0143] This application pertains to the application of artificial intelligence technology, aiming to improve the efficiency and accuracy of video content review in cloud storage applications. Based on advanced artificial intelligence and big data analysis, it accurately identifies various scenarios involving politically sensitive, pornographic, violent, terrorist, spam, and logo watermark-related content, proactively preventing the risk of prohibited and illegal video content, purifying the online environment, and enhancing user experience.
[0144] With the development of integrated circuit technology and 5G communication technology, the computing performance of mobile terminals has been greatly enhanced. The rise of various social media platforms has ushered in an era of self-media, where all kinds of information can be published and disseminated online, and information interaction among netizens has significantly increased. However, the advent of the self-media era has also provided a convenient channel for the spread of illegal information. Illegal information can easily trigger adverse chain reactions, posing a serious threat to national security, social stability, and the health of the online environment, causing enormous negative impacts. Cloud storage connects mobile and PC terminals, enabling file sharing, but it also opens the door to the spread of prohibited information. However, cloud storage providers have an obligation to review the legality and security of stored files; therefore, each cloud storage application provider has its own content review system.
[0145] As a digital fingerprint of a file, each video has a unique MD5 hash. For cloud storage applications or video websites, establishing a continuously updated MD5 database of inappropriate videos and automatically reviewing the MD5 hashes of uploaded videos is crucial. Videos falling into the category of inappropriate content can be blocked or deleted, and the uploader can be blacklisted as a dangerous user. Currently, the standards and methods for establishing review criteria for inappropriate videos are mainly achieved through continuously enriching the databases of prohibited words, prohibited audio information, and prohibited videos, monitoring dangerous users, and establishing a reward mechanism for reporting violations. From an economic and technical perspective, establishing a continuously updated MD5 database and using a review scheme that combines AI and human review remains the mainstream choice for reviewing video content.
[0146] If artificial intelligence is used reasonably for review to eliminate the manual review step, it will not only significantly improve the speed and efficiency of review, but also save considerable human resources. The main problems in the relevant technologies are as follows: (1) When processing video frames, it mainly relies on extracting facial features of people as the data input for video frame image review, and the data processing speed is slow. (2) There are loopholes in the review process, which cannot completely prevent the upload of prohibited videos. A considerable number of video resources can still escape the review mechanism of cloud storage providers. Prohibited and illegal video resources are then uploaded to cloud storage and spread rampantly on the Internet, causing adverse effects on society. (3) The AI algorithm model in the review system is fixed and cannot learn on its own to adjust the model parameter structure, so the review system cannot keep up with the times to process new prohibited video files. (4) There is a manual review stage in the review mechanism, which consumes a lot of human resources and increases the human resource cost of cloud storage providers.
[0147] This application proposes a video content review method for cloud storage media, which can solve the technical difficulties that cannot be solved in the existing technical solutions. The technical solution provided by this application includes the following features: (1) Representing each pixel in all video frames in the video with 3D coordinates to quickly construct a 3D coordinate model of the outline of people and objects in the video. All pixels constituting the outline of people and objects are represented by vectors. The data of the people and objects model represented by vectors are used as the input of the algorithm model in the AI reviewer. (2) Randomly sampling the frames of the video to be reviewed in a normal distribution form. (3) Constructing an AI content reviewer using a deep convolutional neural network (CNN). By sampling video frames of different proportions in the video to be reviewed, constructing network models with different numbers of layers for the deep convolutional network, and building two reviewers with different functions, a fast AI reviewer and a strict AI reviewer. The two reviewers cooperate with each other to improve both the speed and accuracy of the video review process.
[0148] This application proposes a video content review method for cloud storage applications. The method uses an online variable-structure convolutional neural network model to classify and identify prohibited video data, compares the video content with data in the prohibited video database to obtain the degree of prohibitedness of the video to be uploaded, and determines whether to upload the video.
[0149] The process of this method includes: the review system provided in this application performs MD5 information processing on the uploaded video, compares the MD5 value with the data in the prohibited video library for initial screening, and uses OCR, ASR, image recognition technology, etc. to process the text information, audio information, and video frame information of the video into text information and image data, which are then input into the AI reviewer. The video content is then reviewed for the first time using a fast AI content reviewer. A second review is conducted using a strict AI content reviewer. Videos that fail both reviews are added to the prohibited video library to update the prohibited video library, and the uploaded prohibited videos are also processed.
[0150] The review system adopts a distributed architecture and relies on the computing power of mobile cloud computing. The entire review system is deployed on the physical machines of mobile cloud, with the AI reviewer service deployed and running on GPU cloud hosts. It relies on the super floating-point computing power of GPU cloud hosts to complete the training of AI models, thereby improving the review speed of the review system.
[0151] This application verifies the MD5 hash of video files and performs an initial comparison with the MD5 hashes of all files in the prohibited content library. It also customizes the number of video frames extracted, selecting a number proportional to the video's duration. The application analyzes the audio files contained in the video frames and the text information embedded within them. It uses Optical Character Recognition (OCR) to extract text information from corresponding images in all video frames, and Automatic Speech Recognition (ASR) to convert the audio into text. The text data is then merged with the processed audio data and reviewed by an AI reviewer. An image recognition algorithm identifies each frame of the video to construct the outlines of people and objects within it. These video frame images serve as input for both rapid and rigorous AI review. Finally, an AI content reviewer using a variable-depth convolutional neural network model categorizes the videos according to their review level.
[0152] Figure 4E This is a schematic diagram of the data processing flow of the AI reviewer provided in the embodiments of this application, such as... Figure 4E As shown, the AI reviewer includes the following steps when processing data:
[0153] Step S40b: Input data;
[0154] Here, the AI reviewer is divided into a fast AI reviewer and a strict AI reviewer based on the different proportions of video frames extracted when processing video frame data. Before inputting data, video frames need to be randomly extracted in a normal distribution (video frames are divided along the timeline, and the probability of extracting video frames follows a normal distribution, with video frames in the middle of the timeline having a higher probability of being extracted and those at the ends of the timeline having a lower probability of being extracted). The proportion of video frames extracted from data input into the fast AI reviewer is less than that from data input into the strict AI reviewer. Then, the extracted video frame data is input into either the fast AI reviewer or the strict AI reviewer.
[0155] Step S41b: Input data construction;
[0156] Here, the input data construction involves representing each pixel in all video frames using 3D coordinates to quickly build a 3D coordinate model of the outlines of people and objects in the video.
[0157] Step S42b: Process data based on the algorithm;
[0158] Here, the algorithm model for processing data is a deep convolutional neural network, which includes convolutional layers, pooling layers, fully connected layers, and overfitting prevention layers (Dropout). The number of convolutional and pooling layers is related to the extraction ratio of video frames. The number of convolutional and pooling layers varies depending on the extraction ratio of video frames.
[0159] Step S43b: Output the review results.
[0160] In implementation, the sampling ratios of the first and second fast AI reviewers are α1 and α2, respectively, with initial values of 0.236 and 0.286. The dynamic adjustment range of coefficients α1 and α2 is 0.20-0.50, and α1 is always < α2. The initial number of convolutional neural network layers is 36, with the number of network layers variable, ranging from 32 to 60 layers. The strict AI reviewer randomly samples video frames in a normal distribution, with a sampling ratio of β, an initial value of 0.512, and the coefficient β dynamically adjusted from 0.50 to 0.85. The initial number of convolutional neural network layers is 124, with the number of network layers variable, ranging from 80 to 160 layers. The output of the fast AI reviewer is a violation score of O. tmp The calculation formula is shown in formula (9), the violation degree O value range is (0,1), and the classification calculation formula for the review results is shown in formula (10). The review of text and audio information yields score1, and the review of video information yields score2.
[0161]
[0162]
[0163] Where O represents the degree of violation, ∧ represents the intersection, ∨ represents the union, score1≤0.30∧score2≤0.30 means score1 is less than or equal to 0.3 and score2 is less than or equal to 0.3, score1≥0.75∨score2≥0.75 means score1 is greater than or equal to 0.75 or score2 is greater than or equal to 0.75, and 0.5×(score1+1.2score2) represents the weighted average of score1 and score2 to obtain the degree of violation O. A degree of violation O greater than or equal to 0.5 is judged as failing the review, and a degree of violation O less than 0.5 is judged as passing the review.
[0164] Based on this, embodiments of this application provide a content review method, such as... Figure 5 As shown, the method includes:
[0165] Step S501: Divide the video file 51 to be uploaded into segments and upload the resulting segments in parallel.
[0166] For example, when splitting the uploaded video file 51, it can be divided into video data blocks of the same size according to a fixed size, which is called segmentation.
[0167] Step S502, MD5 information value retrieval and comparison steps:
[0168] This step includes: processing each uploaded video data block to obtain the MD5 information value of each video data block, comparing and verifying the obtained MD5 information value with the MD5 values of all videos in the prohibited content video library, and proceeding to the video review process if the preliminary verification is passed.
[0169] Here, videos that fail the initial verification are deemed prohibited and added to the prohibited video library, while the videos are also processed.
[0170] Step S503, Video Information Extraction Step:
[0171] This step includes: using OCR technology to process each frame of the video to obtain text data in the video, using ASR technology to process the audio contained in the video to obtain machine-translated text data, and using video information extraction technology to extract video information;
[0172] Step S504: Use the first strict AI reviewer 53 to review the text data;
[0173] Here, the first strict AI reviewer is used to globally compare the text data in the video with the text data in the prohibited database, and the text and audio review score is obtained after calculation.
[0174] In some embodiments, the video is determined to pass the review based on score1. If the review passes, the review judgment function 56 is entered. If the review fails, the video is rejected and uploaded to the backup storage. A warning image is used as a replacement and the video is updated to the prohibited video library.
[0175] Step S505: The first fast AI reviewer 52 will review the content of each video frame in the video content.
[0176] The review results of the first rapid AI reviewer include Pass (A), Suspected (B), and Fail (C);
[0177] Step S506: If the review result of the first fast AI reviewer is "passed", the video will enter the second fast AI reviewer for a second review, and the video frame review score will be score2. At the same time, if the review result of the second fast AI reviewer is "suspected", the video will enter the second strict AI reviewer for review, and the video frame review score will be score2.
[0178] Step S507: Use the review judgment function 56 to perform a weighted calculation on the text audio review score score1 and the video frame review score score2 to obtain the violation degree O; where the violation degree O is higher than or equal to 0.5, it is judged as failing the review, and lower than 0.5, it is judged as passing the review.
[0179] In some embodiments, approved content is added to a graylist video library, and the file is backed up and stored on cloud storage medium 57. Unapproved videos are rejected from being uploaded to the backup storage, and are replaced with warning images while the videos are updated to the prohibited video library.
[0180] It should be noted that, in Figure 5 In the process, text data enters the first strict AI reviewer for review. If the review result is "pass" (A), it enters the second fast AI reviewer for secondary review. If the review result is "suspected" (B), it enters the second strict AI reviewer for secondary review.
[0181] In some embodiments, because the prohibited video library is constantly and dynamically updated, the system is set to periodically (every 7 days) re-examine and verify the content in the graylist video library, and re-execute the system review process of steps one to five. Video content that passes the review process twice consecutively is added to the whitelist video library, and the system is set to periodically (every 30 days) re-examine and verify the content in the whitelist video library.
[0182] In this embodiment, a self-organizing map network is used to adjust the number of convolutional kernels in each layer of a convolutional neural network, thereby determining the number of layers in the neural network. Here, adjusting the number of convolutional kernels in each layer of the convolutional neural network includes increasing the number of convolutional kernels and deleting the number of convolutional kernels. The methods for determining the number of layers in the neural network by increasing the number of convolutional kernels and by deleting the number of convolutional kernels are described below.
[0183] Among them, (1) methods for increasing the number of convolutional kernels to determine the number of layers in a neural network include:
[0184] Step S50: If the current convolutional layer is the Nth convolutional layer, obtain the connection status between each convolutional kernel in the convolutional layer and the number of connection iterations between two interconnected convolutional kernels;
[0185] For example, such as Figure 4C As shown, the current convolutional layer has 4 convolutional kernels. Matrix E displays the kernel connections, and matrix A shows the iteration count of the connections. For example, matrix A shows that the connection between convolutional kernels 1 and 2 has undergone two iterations, while the connection between convolutional kernels 2 and 4 has 0 iterations. A value of 0 in matrix A can represent unconnected convolutional kernels or a recently formed connection. Matrix E shows that there is no connection between convolutional kernels 1 and 4, while a connection between convolutional kernels 2 and 4 was formed recently.
[0186]
[0187] Step S51: Calculate the weight vector of each convolution kernel after iteration;
[0188] Here, the maximum number of iterations is set to K. When the number of iterations between two interconnected convolutional kernels is less than K, the neighborhood function is first calculated according to formula (11), and then updated according to formula (12) based on the weight vector of the convolutional kernel. Formulas (11) and (12) are as follows:
[0189]
[0190] Among them, H ij (t) represents the neighborhood function; t represents the iteration number; r i (t),r j σ(t) represents the position vectors of convolution kernels i and j; σ(t) represents the learning rate.
[0191]
[0192] Among them, W i (t+1) represents the updated weight vector of the convolution kernel after t+1 iterations; t represents the iteration number; n j(t) represents the number of convolution kernels connected to convolution kernel j; H ij (t) represents the neighborhood function; σ(t) represents the average of all feature vectors of x assigned to the convolution kernel j; σ(t) represents the learning rate.
[0193] Step S52: Update the position of the convolution kernel according to the weight vector of the convolution kernel;
[0194] In some embodiments, after updating the convolution kernel weight vectors in each iteration, the distance between the convolution kernel vectors needs to be calculated according to formula (13), and the convolution kernel positions need to be updated according to formula (14). Formulas (13) and (14) are as follows:
[0195]
[0196] Where represents the update rate of the convolution kernel position at iteration number t; δ ij (t) is the neighborhood function, representing the distance between convolution kernel i and convolution kernel j; γ represents a parameter that controls the shrinkage of the neighborhood.
[0197]
[0198] Where, r i (t+1) represents the position vector of convolution kernel i after t+1 iterations; r i (t),r j (t) represents the position vectors of convolution kernels i and j after t iterations; δ ij (t) is the neighborhood function, representing the distance between convolution kernel i and convolution kernel j; n j α(t) represents the number of convolution kernels connected to convolution kernel j; α(t) represents the update rate of the convolution kernel position after t iterations.
[0199] Step S53: Calculate the quantization error of the convolution kernel;
[0200] In some embodiments, the quantization error of the convolution kernel can be solved according to formula (15), which is shown below:
[0201] qe = w i (t+1)-x (15);
[0202] Where qe represents the quantization error, w i (t+1) represents the weight vector of the winning convolution kernel after t+1 iterations; x represents the feature vector of the image assigned to convolution kernel j.
[0203] Step S54: Calculate the number of convolutional kernels in the current convolutional layer after iteration;
[0204] Here, a growth threshold of Q is set. When the quantization error qe of the convolutional kernel is greater than the growth threshold, the convolutional kernel with the largest quantization error is split into two convolutional kernels. The new convolutional kernels maintain the same connections as the original convolutional kernel, and the two new convolutional kernels are also interconnected. When the quantization error of the convolutional kernel is less than the growth threshold, the number of convolutional kernels in the current convolutional layer is not increased. Therefore, the number of convolutional kernels in the current convolutional layer after iteration can be calculated.
[0205] Step S55: Determine the number of layers in the neural network after iteration.
[0206] After updating the weight vector and position of the convolutional kernel, the growth of the convolutional layer network is controlled by the average quantization error (mqe), and its calculation formula (16) is as follows:
[0207]
[0208] Where mq represents the average quantization error; w i (t+1) represents the weight vector of the winning convolutional kernel after t+1 iterations; x represents the feature vector of the image assigned to convolutional kernel j; n represents the number of convolutional kernels in the current convolutional layer.
[0209] If the average quantization error between two consecutive time steps, meq(t) - mqe(t-1), is less than ξ1, the iteration terminates. At this point, the current convolutional layer is the same as the number of layers in the neural network. Here, ξ1 (e.g., 1E-06) is a very small value. During the algorithm iteration, if the average quantization error between two consecutive time steps is still greater than ξ1, a new convolutional layer needs to be added.
[0210] (2) Methods for determining the number of layers in a neural network by removing the number of convolutional kernels include:
[0211] Step S50a: When the number of connection iterations between two interconnected convolutional kernels is greater than the number of iterations K, disconnect the connection and delete the convolutional kernels that are not connected to any convolutional kernels to obtain the number of convolutional kernels in the current convolutional layer;
[0212] Here, the winning and second winning convolutional kernels are obtained by the self-organizing map algorithm. If they are connected to each other, the existing connection iteration count is reset to 0. After each iteration, if the connection iteration count between convolutional kernel i and convolutional kernel j is greater than the iteration threshold, they are disconnected. At this time, if there is a convolutional kernel i that is not connected to any other convolutional kernel, convolutional kernel i is deleted.
[0213] Step S50b: Determine the number of layers in the neural network after iteration.
[0214] Since the convolution kernel was deleted in step S50a, it means that the number of convolution kernels in the current convolutional layer is less than the number threshold. Therefore, the number of neural network layers in which the current convolutional layer is located is the final number of neural network layers.
[0215] In this embodiment, the audio files and text information contained in the video are processed into text information using Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) technologies. This text information serves as input to the AI reviewer. Image recognition algorithms construct the outlines of people and objects in the video frames. An algorithm model using video frame images as input and a variable-depth convolutional neural network model is used to classify the video's review level. While OCR extracts text from the video, ASR converts the audio information into text information. The text and audio information are compared with a prohibited word database to output a review result score1. The video frame image undergoes two quick AI review sessions or one quick AI review session followed by one strict AI review session, and is compared with video files in the video database to obtain a review result score2. The two scores are weighted to calculate the prohibition level of the video to be reviewed (score1 and score2 are probability values, ranging from 0 to 1). If either score is higher than 0.75, the video to be reviewed is determined to be prohibited; if both scores are lower than 0.3, the review is considered passed.
[0216] exist Figure 5 In this scenario, assuming the video file X to be reviewed passes the MD5 check, the text and audio data undergo content review by the first rigorous AI reviewer and are compared and verified against the text and audio data in the prohibited video database. The review result score 1 is 0.5. The video frame data undergoes content review by the first fast AI reviewer, and the review result is O after comparison and verification against the prohibited video database. tmp The score is 0.5, indicating that video file X is suspected of being prohibited and requires secondary review. After calculation by the second fast AI reviewer, its prohibition score2 is 0.5. After weighted calculation according to Formula 10, the prohibition score O of the video file X to be reviewed is 0.55. The video X to be reviewed is determined to be a prohibited video, the storage request for the video is rejected, the video is added to the prohibited video list, and the prohibited video database is updated.
[0217] Video violation review process dynamic adjustment strategy: a1 is the content size of the video file to be reviewed, in KB; a2 is the video frame extraction ratio (the proportion of extracted video frames to the total video frames (0,1)); a3 is the dataset acquisition time (the total time spent processing video frame information into matrix vectors and processing audio and text information into input matrices).
[0218] The dynamically structured convolutional neural network dynamically adjusts the convolutional neural network structure parameters of the AI review model according to the video violation review process. During the algorithm iteration process, it increases or decreases the number of convolutional and pooling layers and the number of convolutional kernels in each layer to determine whether the quantization error is less than a specific value. If it is less than the specific value, the iteration ends.
[0219] This application presents a system design scheme for applying an online variable neural network model to the field of video content review on cloud storage media. The system classifies prohibited videos by inputting text and audio information and video frames into the model, demonstrating excellent review efficiency. Furthermore, this review system uses a fast AI reviewer and a strict AI reviewer to construct the review process, eliminating the need for manual review during operation and saving significant human resource costs.
[0220] As can be seen from the above, (1) an AI reviewer was constructed based on an online variable depth convolutional neural network. By extracting video frames of different proportions from the video to be reviewed and by constructing and selecting different parameters of the deep convolutional network model, two different reviewers with different functions, namely a fast AI video content reviewer and a strict AI video content reviewer, were built. The two reviewers cooperated with each other to improve both the speed and accuracy of the video review process. (2) The review was conducted using artificial intelligence, and no human review was involved in the review process. This not only significantly improved the speed and efficiency of the review but also saved considerable human resources.
[0221] Through the above-mentioned technical solutions, it can be seen that the technical solutions provided in this application include the following features: (1) This system uses an online variable-structure convolutional neural network model to classify and identify prohibited video data. During the training process, the high-dimensional video data is mapped to a low-dimensional space. Sometimes the clustering of the output results is too scattered or dense, resulting in distorted changes that do not conform to the input pattern. However, it cannot be changed, making it difficult to identify the output results of the deep convolutional neural network. Therefore, the convolutional neural network that can change its network structure online can dynamically change its network structure during the training process to alleviate this distorted change and improve the accuracy of prohibited video classification and identification. (2) An AI reviewer is constructed using an online variable-structure deep convolutional neural network. By adjusting the model parameters, the AI reviewer is set to two types: a fast AI reviewer and a strict AI reviewer. The two fast AI reviewers and the strict AI reviewer are combined to form a complete video review system. (3) In this embodiment, the text and audio information in the video are separated from the video frame information and reviewed separately to obtain score1 and score2. Finally, the weighted score of the review results is used to dynamically adjust the parameters of the AI reviewer based on the violation level. Multiple fast AI reviewers and strict AI reviewers are combined to complete the video review process through the specific review process of this embodiment. (4) In this embodiment, the video review violation level is dynamically adjusted online using the video review violation level dynamic adjustment strategy. The AI reviewer is constructed using an online variable structure CNN network, which enables the AI reviewer to adaptively optimize the review model structure and improve the accuracy of the classification of prohibited videos.
[0222] Based on the foregoing embodiments, this application provides a content review device, which includes various modules and sub-modules included in each module. Each unit included in each sub-module can be implemented by a processor in a computer device; of course, it can also be implemented by logic circuits. In the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field programmable gate array (FPGA), etc. Figure 6 This is a schematic diagram of the composition structure of a content review device provided in an embodiment of this application, as shown below. Figure 6 As shown, the device includes:
[0223] The acquisition module 610 is used to acquire the data to be reviewed of the video to be reviewed; wherein, the data to be reviewed includes the text data and video data of the video to be reviewed; the first review module 620 is used to review the text data through a first AI review model to obtain a first review result; the second review module 630 is used to review the video data through a second AI review model to obtain a second review result;
[0224] The first determining module 640 is used to determine the review result of the video to be reviewed based on the first review result and the second review result.
[0225] In some embodiments, the apparatus further includes: a segmentation module for segmenting the video to be uploaded to obtain at least one video data block; a second determination module for determining the MD5 data value of each video data block; a third determination module for determining a first review result for each video data block based on the MD5 data value of each video data block and the MD5 data value in a preset MD5 database; wherein the first review result includes pass or fail; and the third review module for determining the video to be uploaded as the video to be reviewed if the first review result of all video data blocks in the video to be uploaded is pass.
[0226] In some embodiments, the acquisition module includes: a first determining submodule, configured to determine first text data of the video to be reviewed based on text information extraction technology; a second determining submodule, configured to determine second text data of the video to be reviewed based on speech recognition technology; a third determining submodule, configured to determine text data of the video to be reviewed based on the first text data and the second text data; and a fourth determining submodule, configured to determine video data of the video to be reviewed based on an image recognition algorithm.
[0227] In some embodiments, the fourth determining submodule includes: a partitioning unit, configured to divide each video frame in the video to be reviewed into N×N sub-regions; a first determining unit, configured to determine the LBP value of each pixel in each sub-region; a second determining unit, configured to determine a statistical histogram of each sub-region based on the LBP value of each pixel in each sub-region; a first connecting unit, configured to connect the histograms of each sub-region to obtain a feature vector of each video frame in the video to be reviewed; and a second connecting unit, configured to sequentially connect all video frames based on the timestamp of the feature vector of each video frame in the video to be reviewed to obtain video data of the video to be reviewed.
[0228] In some embodiments, the second AI review model includes a first AI reviewer and a second AI reviewer, and the second review module includes:
[0229] A first extraction submodule is used to extract the video data according to a preset first extraction ratio to obtain first video data; a first review submodule is used to review the first video data through the first AI reviewer to obtain a preliminary review result; wherein the preliminary review result includes pass, suspected, or fail; a second extraction submodule is used to extract the first video data whose preliminary review result is pass according to a preset second extraction ratio to obtain second video data; wherein the second extraction ratio is greater than the first extraction ratio; a second review submodule is used to review the second video data through the second AI reviewer to obtain a second review result.
[0230] In some embodiments, the second AI review model further includes a first AI reviewer and a third AI reviewer, and the second review module further includes: a third extraction submodule, used to extract the first video data whose pre-review result is suspected according to a preset third extraction ratio to obtain third video data; wherein the third extraction ratio is greater than the second extraction ratio; and a third review submodule, used to review the third video data through the third AI reviewer to obtain a second review result.
[0231] In some embodiments, the number of neural network layers in the third AI reviewer is greater than the number of neural network layers in the first and second AI reviewers.
[0232] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0233] It should be noted that, in the embodiments of this application, if the above-described content review method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0234] Figure 7 This is a schematic diagram of the hardware entity of a content review device provided in an embodiment of the present disclosure, such as... Figure 7As shown, the hardware entity of the content review device 700 (which may be a computer device in practice) includes a processor 710 and a memory 720, wherein the memory 720 stores a computer program that can run on the processor 710, and the processor 710 executes the program to implement the steps in the method of any of the above embodiments.
[0235] The memory 720 stores a computer program that can run on the processor. The memory 720 is configured to store instructions and applications executable by the processor 710, and may also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) for the processor 710 and various modules in the device 1600. This can be implemented using flash memory or random access memory (RAM). The steps of the method described above are implemented when the processor 710 executes the program.
[0236] This disclosure provides a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the content review method as described in any of the above embodiments.
[0237] It should be noted that the descriptions of the storage medium embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0238] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0239] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0240] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0241] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0242] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0243] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0244] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0245] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A content review method, characterized in that, The method includes: Obtain the review data of the video to be reviewed; wherein, the review data includes the text data and video data of the video to be reviewed; The text data is reviewed using a first AI review model to obtain a first review result; The video data is reviewed using a second AI review model to obtain a second review result; Based on the first review result and the second review result, the review result of the video to be reviewed is determined; The second AI review model includes a first AI reviewer and a second AI reviewer. The step of reviewing the video data using the second AI review model to obtain a second review result includes: The video data is extracted according to a preset first extraction ratio to obtain the first video data; The first AI reviewer reviews the first video data to obtain a preliminary review result; wherein the preliminary review result includes pass, suspected, or fail. The first video data that passed the preliminary review is extracted according to a preset second extraction ratio to obtain the second video data; wherein the second extraction ratio is greater than the first extraction ratio. The second AI reviewer reviews the second video data to obtain a second review result; The second AI review model further includes a first AI reviewer and a third AI reviewer. The method of reviewing the video data using the second AI review model to obtain a second review result further includes: The first video data whose preliminary review result is suspected is extracted according to a preset third extraction ratio to obtain third video data; wherein, the third extraction ratio is greater than the second extraction ratio; The third video data is reviewed by the third AI reviewer to obtain a second review result.
2. The method according to claim 1, characterized in that, The method further includes: The video to be uploaded is segmented to obtain at least one video data block; Determine the MD5 value of each video data block; Based on the MD5 data value of each video data block and the MD5 data value in a preset MD5 database, a third review result is determined for each video data block; wherein the third review result includes pass or fail. If all video data blocks in the video to be uploaded pass the third review, the video to be uploaded will be identified as the video to be reviewed.
3. The method according to claim 1, characterized in that, The process of obtaining the review data for the video to be reviewed includes: Based on text information extraction technology, the first text data of the video to be reviewed is determined; Based on speech recognition technology, the second text data of the video to be reviewed is determined; Based on the first text data and the second text data, determine the text data of the video to be reviewed; Based on image recognition algorithms, the video data of the video to be reviewed is determined.
4. The method according to claim 3, characterized in that, The process of determining the video data of the video to be reviewed based on the image recognition algorithm includes: Each video frame in the video to be reviewed is divided into N×N sub-regions; Determine the LBP value for each pixel in each sub-region; Based on the LBP value of each pixel in each sub-region, determine the statistical histogram of each sub-region; By connecting the histograms of each sub-region, the feature vector of each video frame in the video to be reviewed is obtained. Based on the timestamp of the feature vector of each video frame in the video to be reviewed, all video frames are concatenated sequentially to obtain the video data of the video to be reviewed.
5. The method according to claim 1, characterized in that, The algorithms of the first, second, and third AI reviewers are deep convolutional neural networks, among which, The third AI reviewer has more convolutional neural network layers than the first and second AI reviewers. The number of neural network layers in the first, second, and third AI reviewers is adjusted based on a self-organizing map algorithm.
6. A content review device, characterized in that, The content moderation device is used to implement the content moderation method according to any one of claims 1 to 5.
7. A content review device, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the content review method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, It stores executable instructions for causing a processor to execute, thereby implementing the content review method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Content management system
CN101558591A
Method for performing legal clearance review of digital content
CN113424204A