AI Recognition and Detection Method and System for Unapproved Videos in TV Drama Submissions
By maintaining consistent feature points across frames and applying adaptive image enhancement, the method addresses exposure-induced inaccuracies in video frame analysis, enhancing key frames and improving detection accuracy for rule-violating content.
Patent Information
- Application Number
- CN202411389328.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-09-30
AI Technical Summary
Existing video frame analysis methods for identifying rule-violating content in video sequences suffer from inaccuracies due to exposure issues causing irrelevant frames to be misidentified as key frames, leading to noise in feature extraction and reduced detection accuracy.
The method identifies and enhances key frames by maintaining consistent feature points across adjacent frames, filtering out irrelevant frames, and applying adaptive image enhancement techniques to behavior regions based on feature point density and contrast, using a modified double-edge filter.
This approach improves the accuracy of rule-violating video detection by ensuring key frames are accurately identified and enhanced, reducing noise and enhancing relevant behavior information, thereby improving the overall detection precision.
Smart Images

Figure CN119296001B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video recognition, and particularly relates to a method and system for AI recognition and detection of illegal videos for drama submission for review. Background Art
[0002] With the continuous increase in drama production and distribution, viewers can easily obtain rich film and television content through online platforms, meeting diverse cultural needs. However, the illegal content in some dramas may pose a threat to the growth of teenagers. Therefore, identifying and handling these illegal dramas has become a key task in content supervision. Considering that the drama library is usually relatively large, and the traditional manual review method has low supervision efficiency and may have missed reviews. Therefore, AI is usually used to perform automated illegal video recognition and detection using artificial intelligence technology to improve the review efficiency and accuracy.
[0003] The prior art usually extracts key frames in the video for drama submission for review based on the inter-frame difference method to perform illegal recognition and detection on the video. However, when the exposure is too high or too low, it will cause loss of image details or color distortion, resulting in video frames that are not key frames but are affected by the exposure being identified as key frames, that is, invalid key frames. These invalid key frames cannot provide effective context analysis, and the existence of invalid key frames will introduce noise in the process of key frame behavior analysis and feature extraction, resulting in the information of valid key frames being masked or misunderstood, thus affecting the accuracy of behavior analysis and reducing the accuracy of AI recognition and detection of illegal videos for drama submission for review. Summary of the Invention
[0004] The present application provides a method and system for AI recognition and detection of illegal videos for drama submission for review. By the characteristics that the features of valid key frames are consistent between adjacent frames and the feature points of valid key frames are relatively densely distributed, the validity degree of each video key frame is determined, and thus valid key frames are screened out according to the validity degree of key frames; and further, in order to enhance the behavior regions with higher information content in the valid key frames, the behavior feature points divided according to the characteristics that the contrast of the feature points corresponding to the behavior actions is relatively high and the feature points corresponding to the behavior actions are relatively concentrated are further used to determine each behavior feature region; and adaptive filtering is performed based on the overall behavior features and feature point distribution in the behavior feature region, so as to perform adaptive image enhancement on the important behavior information in each valid key frame, solving the problem that the information of valid key frames is masked or misunderstood, thus affecting the accuracy of behavior analysis, and making the accuracy of AI recognition and detection of illegal videos for drama submission for review higher.
[0005] The first aspect of the present application provides a method for AI recognition and detection of illegal videos for drama submission for review, including:
[0006] Extract all video key frames in the video for drama submission for review;
[0007] In the video for drama submission review, initial feature points are obtained in each video frame based on a feature extraction method; according to the similarity of the distribution of initial feature points between each video key frame and its adjacent video frames, the inter-frame stability of each video key frame is determined; according to the inter-frame stability and the density of initial feature points in each video key frame, the effectiveness of each video key frame is determined; effective key frames are selected according to the effectiveness of the key frames;
[0008] In each effective key frame, according to the concentrated distribution and contrast of feature points in the local area where each initial feature point is located, the behavior evaluation index of each initial feature point in each effective key frame is determined; all behavior feature points in each effective key frame are selected according to the behavior evaluation index; according to the concentrated distribution of behavior feature points in each effective key frame, the behavior feature regions in each effective key frame are divided;
[0009] According to the overall magnitude of the behavior evaluation indices of each behavior feature point and the density of behavior feature points in the behavior feature region, the modified bilateral filtering spatial weight of each behavior feature region is determined; after filtering the behavior feature region according to the modified bilateral filtering spatial weight, the enhanced key frame corresponding to each effective key frame is obtained; AI recognition detection of illegal videos for drama submission review is performed according to the enhanced key frame.
[0010] Furthermore, the process of obtaining the inter-frame stability includes:
[0011] In the video for drama submission review, the mean value between the number of initial feature points of each video key frame and the number of initial feature points of the corresponding previous video frame is used as the first reference quantity value; the number of first matching feature points between each video key frame and the corresponding previous video frame is obtained through the FLANN feature matching algorithm; the negative correlation mapping value of the difference between the number of first matching feature points and the first reference quantity value is used as the first stability degree of each video key frame;
[0012] After replacing the previous video frame in the process of obtaining the first stability degree with the subsequent video frame, the second stability degree between each video key frame and the subsequent video frame is determined; the minimum value between the first stability degree and the second stability degree is used as the feature point stability degree of each video key frame;
[0013] The difference between the gray value of each pixel point in each video key frame and the gray value of the pixel point at the same pixel position in the corresponding previous video frame is used as the first local gray difference of each pixel point in each video key frame; the negative correlation mapping of the mean value of the first local gray differences of all pixel points in each video frame is performed to determine the first overall gray stability between each video key frame and the previous video frame;
[0014] After replacing the previous video frame with the subsequent video frame in the process of obtaining the first overall gray stability, determine the second overall gray stability between each video key frame and the subsequent video frame; take the minimum value between the first overall gray stability and the second overall gray stability as the gray feature stability degree of each video key frame.
[0015] Determine the inter-frame stability of each video key frame according to the product between the feature point stability degree and the gray feature stability degree.
[0016] Further, the process of obtaining the key frame effectiveness degree includes:
[0017] Normalize the product between the number of initial feature points in each video key frame and the inter-frame stability to determine the key frame effectiveness degree of each video key frame.
[0018] Further, the process of obtaining the effective key frames includes:
[0019] Take the video key frames with the key frame effectiveness degree greater than a preset effective threshold as the effective key frames.
[0020] Further, the process of obtaining the behavior evaluation index includes:
[0021] Perform canny edge detection on each effective key frame, and obtain all edge connected regions corresponding to each effective key frame according to the image edges in the obtained edge image.
[0022] Take the mean value of the number of initial feature points of all edge connected regions in each effective key frame as the reference number mean value; take the ratio between the number of initial feature points in the edge connected region where each initial feature point is located in each effective key frame and the reference number mean value as the local feature point density of each initial feature point in each effective key frame.
[0023] In each effective key frame, normalize the product between the contrast of each initial feature point and the local feature point density to determine the behavior evaluation index of each initial feature point in each effective key frame.
[0024] Further, the process of obtaining the behavior feature points includes:
[0025] In each effective key frame, take the initial feature points with the behavior evaluation index greater than a preset behavior threshold as the behavior feature points.
[0026] Further, the process of obtaining the behavior feature region includes:
[0027] In each valid key frame, an edge connected region with the number of behavior feature points greater than a preset number threshold is used as the behavior feature region of each valid key frame.
[0028] Further, the process of obtaining the modified bilateral filtering spatial weight includes:
[0029] In each valid key frame, the mean value of the behavior evaluation indexes of all behavior feature points in each behavior feature region is used as the regional behavior feature value of each behavior feature region;
[0030] In each valid key frame, the mean value of the initial number of feature points of all behavior feature regions is used as the corresponding reference feature point density; the ratio between the initial number of feature points of each behavior feature region and the reference feature point density is used as the feature point density weight of each behavior feature region in each valid key frame;
[0031] The product of the regional behavior feature value, the feature point density weight and the prior bilateral filtering spatial weight is used as the modified bilateral filtering spatial weight of each behavior feature region in each valid key frame.
[0032] Further, the process of obtaining the enhanced key frame includes:
[0033] In each valid key frame, bilateral filtering is performed on the corresponding behavior feature region through the modified bilateral filtering spatial weight to obtain the filtered and enhanced region of each behavior feature region;
[0034] After each behavior feature region in each valid key frame is replaced with the corresponding filtered and enhanced region, the enhanced key frame corresponding to each valid key frame is obtained.
[0035] In a second aspect, the present application provides an AI recognition and detection system for drama submission violation videos, and the system includes:
[0036] A data acquisition module, configured to extract all video key frames in the drama submission video;
[0037] A valid key frame acquisition module, configured to, in the drama submission video, obtain initial feature points in each video frame based on a feature extraction method; determine the inter-frame stability of each video key frame according to the similarity of the initial feature point distribution between each video key frame and its adjacent video frames; determine the key frame effectiveness of each video key frame according to the inter-frame stability and the initial feature point density in each video key frame; and filter out valid key frames according to the key frame effectiveness;
[0038] The behavior feature region acquisition module is used to determine the behavior evaluation index of each initial feature point in each valid key frame according to the concentrated distribution and contrast of the feature points in the local region where each initial feature point is located; screen out all behavior feature points in each valid key frame according to the behavior evaluation index; and divide the behavior feature region in each valid key frame according to the concentrated distribution of the behavior feature points in each valid key frame.
[0039] The AI recognition and detection module is used to determine the modified bilateral filtering spatial weight of each behavior feature region according to the overall size of the behavior evaluation index and the behavior feature point density of each behavior feature point in the behavior feature region; after filtering the behavior feature region according to the modified bilateral filtering spatial weight, obtain the enhanced key frame corresponding to each valid key frame; and perform AI recognition and detection on the video for episode submission violation according to the enhanced key frame.
[0040] In a third aspect, the present application provides a computer device, including a memory and a processor. The memory is used to store computer program code, and the processor is used to call and run the computer program code from the memory to execute the method as described in the first aspect or any embodiment of the first aspect of the present application.
[0041] In a fourth aspect, the present application provides a computer program product, which includes computer program code. When the computer program code is executed, it is used to execute the method as described in the first aspect or any embodiment of the first aspect of the present application.
[0042] In a fifth aspect, the present application provides a computer-readable storage medium, which stores computer program code. When the computer program code is executed, it is used to execute the method as described in the first aspect or any embodiment of the first aspect of the present application.
[0043] The present application has the following beneficial effects:
[0044] First, since the feature points are consistent between adjacent frames through valid key frames, and the invalid key frames generated by exposure have lower feature point stability due to the static or less obvious change characteristics of the scene. Therefore, first, according to the similarity of the initial feature point distribution between each video key frame and the adjacent video frame, the inter-frame stability of each video key frame is determined.
[0045] Valid key frames usually represent important scene changes and details in the video, so they usually contain more feature points; while the scene changes corresponding to the invalid key frames generated by exposure are not obvious, and the corresponding feature point distribution is relatively sparse. Therefore, the present application measures the effectiveness of key frames by combining the initial feature point density in the video key frames on the basis of the inter-frame stability, and then screens out the valid key frames according to the effectiveness of the key frames.
[0046] In a valid key frame, there are usually a behavior area representing behavior information and a background area representing background information. The behavior area usually contains behavior information that is more important for violation detection, while the background information is less important for violation detection. Therefore, in order to improve the accuracy of subsequent violation detection, it is necessary to perform image enhancement on the behavior area.
[0047] In a valid key frame, the area corresponding to the action behavior often contains more motion details, so it usually corresponds to a higher feature point density. And dynamic actions will highlight edges and details, so the feature points in the action behavior area usually have a higher contrast. Therefore, first, according to the concentrated distribution of feature points and the contrast in the local area where each initial feature point is located, determine the behavior evaluation index of each initial feature point in each valid key frame. Then, based on the behavior evaluation index, filter out more accurate behavior feature points representing behavior information. And based on the behavior feature points, filter out the behavior feature area corresponding to the action behavior area.
[0048] Furthermore, it is necessary to perform local image enhancement on each behavior feature area to determine a more accurate enhanced key frame. Considering that bilateral filtering can effectively remove noise and can retain edge information to a certain extent, bilateral filtering is used to perform filtered image enhancement on each behavior feature area. For each behavior feature area, the density of feature points reflects the complexity and richness of details of the corresponding behavior feature area. High-density areas usually contain more information. Therefore, a higher weight is required during the filtering process. So, further, according to the overall size of the behavior evaluation index of each behavior feature point in the behavior feature area and the density of behavior feature points, determine the modified bilateral filtering spatial weight of each behavior feature area, and perform filtered enhancement on each behavior feature area in the valid key frame according to the obtained modified bilateral filtering spatial weight, so as to obtain an enhanced key frame with better filtering effect and clearer image information representation, making the accuracy of the AI recognition and detection of episode submission violation videos based on the enhanced key frame higher. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 Flowchart of a method for AI recognition and detection of episode submission violation videos provided by an embodiment of the present invention;
[0051] Figure 2 Structural diagram of an AI recognition and detection system for illegal videos in drama submission for review provided by an embodiment of the present invention;
[0052] Figure 3 Schematic diagram of the structure of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0053] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manners, structures, features and effects of an AI recognition and detection method and system for illegal videos in drama submission for review proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment, and specific features, structures or characteristics in one or more embodiments can be combined in any suitable form. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0055] The following specifically describes the specific solutions of an AI recognition and detection method and system for illegal videos in drama submission for review provided by the present invention with reference to the accompanying drawings.
[0056] An embodiment of the present application provides an AI recognition and detection method for illegal videos in drama submission for review. Please refer to Figure 1 , which shows a flowchart of an AI recognition and detection method for illegal videos in drama submission for review provided by an embodiment of the present invention. The method includes:
[0057] Step S101: Extract all video key frames in the drama submission video.
[0058] First, obtain the drama submission video, and extract the video key frames in the drama submission video through the inter-frame difference method; it should be noted that the inter-frame difference method is a well-known technical means to those skilled in the art, and in other specific implementation manners of the embodiments of the present invention, video key frames can also be extracted through other key frame extraction methods, such as using a trained convolutional neural network for video key frame extraction, which will not be further limited and elaborated herein.
[0059] Step S102: In the video submitted for review of the episode, obtain the initial feature points in each video frame based on the feature extraction method; determine the inter-frame stability of each video key frame according to the similarity of the distribution of the initial feature points between each video key frame and its adjacent video frames; determine the effectiveness of each video key frame according to the inter-frame stability and the density of the initial feature points in each video key frame; and filter out the effective key frames according to the effectiveness of the key frames.
[0060] Since the feature points are consistent between adjacent frames through the effective key frames, while the ineffective key frames generated by exposure have relatively low stability of feature points due to the characteristics of static or insignificant scene changes. Therefore, first, determine the inter-frame stability of each video key frame according to the similarity of the distribution of the initial feature points between each video key frame and its adjacent video frames; the greater the corresponding inter-frame stability, the more likely it is that the video key frame belongs to the effective key frames.
[0061] Effective key frames usually represent important scene changes and details in the video, so they usually contain more feature points; while the ineffective key frames generated by exposure have insignificant scene changes, and the corresponding feature point distribution is relatively sparse. Therefore, in this application, the effectiveness of the key frames is measured by combining the density of the initial feature points in the video key frames on the basis of the inter-frame stability, so as to filter out the effective key frames according to the effectiveness of the key frames.
[0062] In a specific implementation manner of the embodiment of the present invention, the initial feature points in each video frame are extracted by the SIFT feature extraction algorithm, and the SIFT feature extraction algorithm is a well-known technical means for those skilled in the art, and will not be further limited and described herein.
[0063] Preferably, the process of obtaining the inter-frame stability includes:
[0064] In the video submitted for review of the TV series, the mean value between the number of initial feature points of each video key frame and the number of initial feature points of the corresponding previous video frame is used as the first reference quantity value; the number of first matching feature points between each video key frame and the corresponding previous video frame is obtained through the FLANN feature matching algorithm; the negative correlation mapping value of the difference between the number of first matching feature points and the first reference quantity value is used as the first stability degree of each video key frame; after replacing the previous video frame in the process of obtaining the first stability degree with the subsequent video frame, the second stability degree between each video key frame and the subsequent video frame is determined; the minimum value between the first stability degree and the second stability degree is used as the feature point stability degree of each video key frame. The specific process of obtaining the second stability degree includes: in the video submitted for review of the TV series, the mean value between the number of initial feature points of each video key frame and the number of initial feature points of the corresponding subsequent video frame is used as the second reference quantity value; the number of second matching feature points between each video key frame and the corresponding subsequent video frame is obtained through the FLANN feature matching algorithm; the negative correlation mapping value of the difference between the number of second matching feature points and the second reference quantity value is used as the second stability degree of each video key frame.
[0065] First, considering that the feature points of valid key frames remain consistent among multiple adjacent frames, on the basis of each video key frame, the degree of feature point matching is calculated for the previous video frame and the subsequent video frame respectively, so as to more accurately reflect the stability degree of the feature points of the video key frame in the front and back adjacent directions and weaken the influence of accidental factors; and the mean value of the number of feature points of two adjacent video frames is introduced in the calculation process, so that the calculated stability degree of the feature points has better robustness; when the obtained stability degree of the feature points is larger, the corresponding video key frame more conforms to the characteristics of the valid key frame. It should be noted that the FLANN feature matching algorithm is a well-known technical means for those skilled in the art and will not be further defined and described here.
[0066] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the feature point stability degree is expressed by the formula: T k = Min((exp(-|n k,1 -n′ k |), exp(-|n k,2 -n″ k |)) where, T k is the feature point stability degree of the kth video key frame, n k,1 is the first reference quantity value of the kth video key frame; n′ k is the number of first matching feature points corresponding to the kth video key frame; n k,2 is the second reference quantity value of the kth video key frame; n″k is the number of second matching feature points corresponding to the k-th video key frame; exp(-|n k,1 - n′ k |) is the first stability degree corresponding to the k-th video key frame; exp(-|n k,2 - n″ k |) is the second stability degree corresponding to the k-th video key frame; Min() is the minimum value selection function, Min(exp(-|n k,1 - n′ k |), exp(-|n k,2 - n″ k |)) means to select a minimum value between exp(-|n k,1 - n′ k |) and exp(-|n k,2 - n″ k |); exp() is the exponential function with the natural constant as the base. Implementers can adopt other negatively correlated mapping methods according to the specific implementation environment, such as 1-Norm(), 1-tanh(), and reciprocal. Among them, tanh() is the hyperbolic tangent function, which will not be elaborated further here. It should be noted that for the video key frame that belongs to the first video frame or the last video frame of the drama submission video, since there is no previous video frame or next video frame, the embodiments of the present invention directly use the corresponding video key frame as an effective key frame for analysis, that is, no calculation of the inter-frame stability and the effectiveness of the key frame is performed, making the embodiments of the present invention more complete.
[0067] The difference between the grayscale value of each pixel in each video key frame and the grayscale value of the pixel at the same pixel position in the corresponding previous video frame is used as the first local grayscale difference of each pixel in each video key frame; the mean value of the first local grayscale differences of all pixels in each video frame is negatively correlated and mapped to determine the first overall grayscale stability between each video key frame and the previous video frame. For invalid key frames caused by exposure effects, there is usually a large difference in the overall grayscale between them and adjacent images. Therefore, the greater the change in the grayscale value at each pixel position, the more likely the corresponding video key frame belongs to an invalid key frame; conversely, the smaller the change in the grayscale value at each pixel position, the more likely the corresponding video key frame belongs to a valid key frame. Similarly, in order to reduce the influence of accidental factors, further, after replacing the previous video frame in the process of obtaining the first overall grayscale stability with the next video frame, the second overall grayscale stability between each video key frame and the next video frame is determined. Specifically: the difference between the grayscale value of each pixel in each video key frame and the grayscale value of the pixel at the same pixel position in the corresponding next video frame is used as the second local grayscale difference of each pixel in each video key frame; the mean value of the second local grayscale differences of all pixels in each video frame is negatively correlated and mapped to determine the second overall grayscale stability of each video key frame. Thus, further, the minimum value between the first overall grayscale stability and the second overall grayscale stability is used as the grayscale feature stability degree of each video key frame. Based on each video key frame, the overall grayscale stability is calculated for the previous video frame and the next video frame respectively, so as to more accurately reflect the grayscale feature stability degree of the video key frame in the two adjacent directions of the front and back, reduce the influence of accidental factors, and make it more likely that the corresponding video key frame belongs to a valid key frame when the grayscale feature stability degree is greater.
[0068] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the grayscale feature stability degree is expressed by the formula: where k H k,i is the grayscale feature stability degree of the kth video key frame; N is the number of pixels of the video key frame; g k,i is the grayscale value of the ith pixel in the kth video key frame; g′ k,i is the grayscale value of the ith pixel in the kth video key frame at the same pixel position in the previous video frame; g″ k,i is the grayscale value of the ith pixel in the kth video key frame at the same pixel position in the next video frame; |g k,i - g′ k,i | is the first local grayscale difference of the ith pixel in the kth video key frame; |gk,i -g″ k,i is the second local gray difference of the i-th pixel point in the k-th video key frame; is the first overall gray stability between the k-th video key frame and the previous video frame; is the second overall gray stability between the k-th video key frame and the next video frame; Min() is the minimum selection function, that is, and select a minimum value; exp() is the exponential function with the natural constant as the base. Implementers can adopt other negatively correlated mapping methods according to the specific implementation environment, such as 1-Norm(), 1-tanh(), and reciprocal. Among them, tanh() is the hyperbolic tangent function, which will not be elaborated further here.
[0069] Furthermore, in combination with the stability degree of feature points and the stability degree of gray features, the possibility of effective key frames is characterized in the dimension of stability. According to the product between the stability degree of feature points and the stability degree of gray features, the inter-frame stability of each video key frame is determined, so that the greater the inter-frame stability, the greater the possibility that the corresponding video key frame belongs to an effective key frame.
[0070] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the inter-frame stability is expressed by the formula: W k = T k × H k ; where, W k is the inter-frame stability of the k-th video key frame; H k is the stability degree of the gray feature of the k-th video key frame; T k is the stability degree of the feature points of the k-th video key frame. It should be noted that in addition to the product, implementers can also calculate the inter-frame stability by other methods according to the relevant relationship, such as the normalized values of the mean and sum, etc., which will not be elaborated further here.
[0071] Effective key frames usually represent important scene changes and details in the video, so that the included feature points are usually more; while the scene changes corresponding to the invalid key frames generated by exposure are not obvious, and the corresponding feature point distribution is relatively sparse; therefore, this application measures the effectiveness of key frames by combining the initial feature point density in the video key frame on the basis of the inter-frame stability, and thus filters out effective key frames according to the effectiveness of key frames.
[0072] Preferably, in some possible implementation manners of the embodiment of the present invention, the process of obtaining the effectiveness of key frames includes:
[0073] Since the sizes of video key frames are the same, the more the corresponding initial feature points, the greater the feature point density, and the greater the possibility of the corresponding effective key frame. Further, considering that when the inter-frame stability is greater, the possibility that the corresponding video key frame belongs to an effective key frame is greater, the product of the number of initial feature points and the inter-frame stability in each video key frame is normalized to determine the effectiveness degree of each video key frame. The greater the effectiveness degree of the key frame, the more likely the corresponding video key frame is an effective video frame. In a specific implementation manner of the embodiment of the present invention, the process of obtaining the effectiveness degree of the key frame is expressed by the formula: E k = Norm(W k ×M k ); where E k is the effectiveness degree of the k-th video key frame; W k is the inter-frame stability of the k-th video key frame; M k is the number of initial feature points of the k-th video key frame; Norm() is a linear normalization function, and the implementer can adjust the normalization method according to the specific implementation environment. It should be noted that in addition to the product, the implementer can also calculate the effectiveness degree of the key frame by other methods according to the relevant relationship, such as the normalized values of the mean and sum, etc., which will not be elaborated further here.
[0074] Preferably, in some possible implementation manners of the embodiment of the present invention, the process of obtaining the effective key frame includes:
[0075] Since the greater the effectiveness degree of the key frame, the more likely the corresponding video key frame is an effective video frame, further, the video key frames with the effectiveness degree of the key frame greater than the preset effective threshold are used as effective key frames. In a specific implementation manner of the embodiment of the present invention, the preset effective threshold is set to 0.5, which can be adjusted according to the specific implementation environment.
[0076] Step S103: In each effective key frame, determine the behavior evaluation index of each initial feature point according to the concentrated distribution of feature points and the contrast in the local area where each initial feature point is located; screen out all behavior feature points in each effective key frame according to the behavior evaluation index; and divide the behavior feature regions in each effective key frame according to the concentrated distribution of behavior feature points in each effective key frame.
[0077] In an effective key frame, there are usually a behavior region representing behavior information and a background region representing background information. The behavior region usually contains behavior information that is more important for violation detection, while the background information is less important for violation detection. Therefore, in order to improve the accuracy of subsequent violation detection, it is necessary to perform image enhancement on the behavior region.
[0078] In the valid key frames, the region corresponding to the action behavior often contains more motion details, so it usually corresponds to a higher feature point density. Moreover, dynamic actions highlight edges and details, so the feature points in the action behavior region usually have a higher contrast. Therefore, first, according to the concentrated distribution and contrast of the feature points in the local region where each initial feature point is located, the behavior evaluation index of each initial feature point in each valid key frame is determined. Then, more accurate behavior feature points representing the behavior information are selected according to the behavior evaluation index. And based on the behavior feature points, the behavior feature region corresponding to the action behavior region is selected.
[0079] Preferably, in some possible implementation manners of the embodiments of the present invention, the process of obtaining the behavior evaluation index includes:
[0080] Perform canny edge detection on each valid key frame, and obtain all edge connected regions corresponding to each valid key frame according to the image edges in the obtained edge image. In the valid key frame, the region corresponding to the action behavior feature usually has a large texture difference from the background region. Therefore, after dividing the valid key frame by the edges obtained by edge detection, a more accurate region corresponding to the action behavior feature can be selected based on the obtained edge connected regions.
[0081] Take the mean value of the number of initial feature points in all edge connected regions in each valid key frame as the reference number mean value; take the ratio between the number of initial feature points in the edge connected region where each initial feature point is located in each valid key frame and the reference number mean value as the local feature point density of each initial feature point in each valid key frame. For the edge connected region where each initial feature point is located, the denser the corresponding initial feature point distribution, the more it conforms to the characteristics of the action behavior region, that is, the more likely the corresponding edge connected region belongs to the behavior feature region, and the more likely the initial feature points in it belong to the behavior feature points. And calculating the mean value of the number of initial feature points in all edge connected regions in each valid key frame, that is, the reference number mean value, can use the reference number mean value as a reference to make the measurement of the feature point density of each edge connected region more accurate. Therefore, when the local feature point density obtained according to the ratio between the number of initial feature points in the edge connected region and the reference number mean value is larger, the corresponding initial feature point is more likely to belong to the behavior feature point. It should be noted that since the feature point density has been considered in the process of screening the valid key frames, the number of feature points in the valid key frame is not 0, so the corresponding reference number mean value cannot be 0, and there is no situation where the denominator is 0.
[0082] In each valid key frame, the product of the contrast of each initial feature point and the local feature point density is normalized to determine the behavior evaluation index of each initial feature point in each valid key frame. In addition, since the contrast of the initial feature points in the action behavior area is relatively higher, the contrast of the initial feature points is combined on the basis of the local feature point density for the co-representation of the behavior evaluation index, so that when the behavior evaluation index is larger, the corresponding initial feature point is more likely to represent the action behavior feature, that is, it is more likely to be a behavior feature point.
[0083] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the behavior evaluation index is expressed by the formula: Z r,c is the behavior evaluation index of the c-th initial feature point in the r-th valid key frame; U r,c is the number of initial feature points in the edge-connected domain where the c-th initial feature point is located in the r-th valid key frame; is the mean value of the number of initial feature points in all edge-connected domains in the r-th valid key frame, that is, the reference number mean value; is the local feature point density of the c-th initial feature point in the r-th valid key frame; D r,c is the contrast of the c-th initial feature point in the r-th valid key frame; Norm() is a linear normalization function, and the implementer can adjust the normalization method according to the specific implementation environment, and the normalization can limit the value range of the behavior evaluation index to be within 0 to 1, which is convenient for subsequent screening. It should be noted that the process of obtaining the contrast of the pixel points in the image or video frame is a well-known technical means for those skilled in the art and will not be further elaborated here.
[0084] Preferably, the process of obtaining the behavior feature points includes:
[0085] Since when the behavior evaluation index is larger, the corresponding initial feature point is more likely to represent the action behavior feature, that is, it is more likely to be a behavior feature point; therefore, in each valid key frame, the initial feature points with a behavior evaluation index greater than the preset behavior threshold are used as behavior feature points. In a specific implementation manner of the embodiment of the present invention, the preset behavior threshold is set to 0.5 and can be adjusted according to the specific implementation environment.
[0086] Preferably, the process of obtaining the behavior feature area includes:
[0087] The more the number of behavior feature points in the edge-connected domain, the more likely the corresponding area belongs to the area corresponding to the action behavior feature. Therefore, in each valid key frame, the edge-connected domain with the number of behavior feature points greater than the preset number threshold is used as the behavior feature area of each valid key frame. In a specific implementation manner of the embodiment of the present invention, the preset number threshold is set to 5 and can be adjusted according to the specific implementation environment.
[0088] Step S104: Determine the modified bilateral filtering spatial weight of each behavior feature region according to the overall size of the behavior evaluation indexes of each behavior feature point in the behavior feature region and the behavior feature point density; after filtering the behavior feature region according to the modified bilateral filtering spatial weight, obtain the enhanced key frame corresponding to each valid key frame; perform AI recognition detection on the drama submission violation videos according to the enhanced key frame.
[0089] Furthermore, local image enhancement needs to be performed on each behavior feature region to determine a more accurate enhanced key frame. Considering that bilateral filtering can effectively remove noise and can retain edge information to a certain extent, bilateral filtering is used to perform filtered image enhancement on each behavior feature region. For each behavior feature region, the density of feature points reflects the complexity and detail richness of the corresponding behavior feature region. High-density regions usually contain more information. Therefore, a higher weight is required during filtering; so further, according to the overall size of the behavior evaluation indexes of each behavior feature point in the behavior feature region and the behavior feature point density, determine the modified bilateral filtering spatial weight of each behavior feature region, and perform filtering enhancement on each behavior feature region in the valid key frame according to the obtained modified bilateral filtering spatial weight, so as to obtain an enhanced key frame with better filtering effect and clearer image information representation, making the accuracy of the AI recognition detection of the drama submission violation videos according to the enhanced key frame higher.
[0090] Preferably, the process of obtaining the modified bilateral filtering spatial weight includes:
[0091] In each valid key frame, take the mean value of the behavior evaluation indexes of all behavior feature points in each behavior feature region as the regional behavior feature value of each behavior feature region. For each behavior feature region, the larger its corresponding regional behavior feature value, the higher the behavior evaluation index of the behavior feature points in this behavior feature region, the higher the complexity and detail richness of the corresponding local region, and the more action behavior information it contains. Therefore, a larger filtering weight is required for enhancement to highlight its action behavior information.
[0092] In each valid key frame, the mean of the initial number of feature points in all behavior feature regions is used as the corresponding reference feature point density; the ratio between the initial number of feature points in each behavior feature region and the reference feature point density is used as the feature point density weight of each behavior feature region in each valid key frame. First, the more the initial number of feature points in the behavior feature region, the denser the feature distribution in this region, and thus a higher spatial weight is required for image enhancement; and taking the mean of the initial number of feature points in all behavior feature regions as a limit makes the robustness of the feature distribution density characterized by the obtained feature point density weight better; and the greater the feature point density weight, the greater the spatial weight requirement for the corresponding bilateral filtering.
[0093] Therefore, further, the product of the regional behavior feature value, the feature point density weight, and the prior bilateral filtering spatial weight is used as the modified bilateral filtering spatial weight of each behavior feature region in each valid key frame; when the regional behavior feature value and the feature point density weight are greater, the corresponding modified bilateral filtering spatial weight is greater, and the more image details are retained after enhancement based on this spatial weight, that is, the action behavior information is clearer; the prior bilateral filtering spatial weight serves as the weighted base, and in a specific implementation manner of the embodiment of the present invention, its value is 3 and can be adjusted according to the specific implementation environment.
[0094] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the modified bilateral filtering spatial weight is expressed by the formula: where, S′ r,x is the modified bilateral filtering spatial weight corresponding to the xth behavior feature region in the rth valid key frame; S is the prior bilateral filtering spatial weight; V r,x is the number of behavior feature points corresponding to the xth behavior feature region in the rth valid key frame; Z r,x,v is the behavior evaluation index of the uth behavior feature point in the xth behavior feature region in the rth valid key frame; P r,x is the initial number of feature points in the xth behavior feature region in the rth valid key frame; is the reference feature point density of the rth valid key frame, that is, the mean of the initial number of feature points in all behavior feature regions in the rth valid key frame, and its value is not 0; is the regional behavior feature value of the xth behavior feature region in the rth valid key frame; is the feature point density weight of the xth behavior feature region in the rth valid key frame.
[0095] Preferably, in some possible implementation manners of the embodiment of the present invention, the process of obtaining the enhanced key frame includes:
[0096] In each valid key frame, bilateral filtering is performed on the corresponding behavior feature region by correcting the bilateral filtering spatial weight to obtain the filtered enhanced region of each behavior feature region; after replacing each behavior feature region in each valid key frame with the corresponding filtered enhanced region, the enhanced key frame corresponding to each valid key frame is obtained. That is, in each valid key frame, different degrees of bilateral filtering enhancement are performed with the corrected bilateral filtering spatial weights corresponding to each behavior feature region, so as to obtain the enhanced key frame after enhancing the action behavior information, making the accuracy of subsequent AI recognition and detection of episode submission violation videos higher.
[0097] Finally, AI recognition and detection of episode submission violation videos is performed based on the enhanced key frames; in a specific implementation manner of the embodiment of the present invention, each enhanced key frame is input into a trained convolutional neural network to output the enhanced key frames with violations, and the loss function is selected as the cross-entropy loss function, which will not be further limited and elaborated here.
[0098] In summary, an AI recognition and detection method for episode submission violation videos determines the effectiveness of each video key frame based on the characteristics that the features are consistent between adjacent frames in the valid key frames and the feature points in the valid key frames are densely distributed, and thus filters out the valid key frames according to the effectiveness of the key frames; and further, in order to enhance the behavior regions with higher information content in the valid key frames, the behavior feature regions are determined based on the behavior feature points divided according to the characteristics that the contrast of the feature points corresponding to the behavior actions is high and the feature points corresponding to the behavior actions are relatively concentrated; and adaptive filtering is performed based on the overall behavior features and the feature point distribution in the behavior feature regions, so as to perform adaptive image enhancement on the important behavior information in each valid key frame, solving the problem that the information in the valid key frames is masked or misunderstood, thus affecting the accuracy of behavior analysis, and making the accuracy of AI recognition and detection of episode submission violation videos higher.
[0099] This application also provides an AI recognition and detection system for episode submission violation videos. Please refer to Figure 2 FIG. 11, which shows the structure diagram of an AI recognition and detection system for episode submission violation videos provided by an embodiment of the present invention. The system includes: a data acquisition module 201, a valid key frame acquisition module 202, a behavior feature region acquisition module 203, and an AI recognition and detection module 204.
[0100] The data acquisition module 201 is used to extract all video key frames in the episode submission video;
[0101] The effective key frame acquisition module 202 is configured to obtain initial feature points in each video frame in the video for drama review based on a feature extraction method; determine the inter-frame stability of each video key frame according to the similarity of the distribution of initial feature points between each video key frame and its adjacent video frames; determine the effectiveness of each video key frame according to the inter-frame stability and the density of initial feature points in each video key frame; and filter out effective key frames according to the effectiveness of the key frames.
[0102] The behavior feature region acquisition module 203 is configured to, in each effective key frame, determine the behavior evaluation index of each initial feature point in each effective key frame according to the concentrated distribution of feature points and the contrast in the local region where each initial feature point is located; filter out all behavior feature points in each effective key frame according to the behavior evaluation index; and divide the behavior feature region in each effective key frame according to the concentrated distribution of behavior feature points in each effective key frame.
[0103] The AI recognition and detection module 204 is configured to determine the modified bilateral filtering spatial weight of each behavior feature region according to the overall size of the behavior evaluation index and the density of behavior feature points in each behavior feature region; obtain the enhanced key frame corresponding to each effective key frame after filtering the behavior feature region according to the modified bilateral filtering spatial weight; and perform AI recognition and detection on the video for drama review with violations according to the enhanced key frame.
[0104] It should be noted that the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, an AI recognition and detection system for videos with violations in drama review and an embodiment of an AI recognition and detection method for videos with violations in drama review provided in the above embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0105] The embodiment of the present application further provides a computer device. Please refer to Figure 3 , which shows a schematic structural diagram of a computer device provided in an embodiment of the present invention. The computer device includes a memory 301, a processor 302, and a computer program 303 stored in the memory 301 and running on the processor 302. When the processor 302 executes the computer program 303, the computer device can execute any one of the above-described AI recognition and detection methods for videos with violations in drama review.
[0106] The embodiments of the present application also provide a computer program product. When the computer program product runs on a computer device, the computer device can execute any one of the above-described episode submission violation video AI recognition and detection methods.
[0107] The embodiments of the present application also provide a computer-readable storage medium. Computer program code is stored in the computer-readable storage medium. When the computer program code runs on a computer device, the computer device can execute any one of the above-described episode submission violation video AI recognition and detection methods.
[0108] In the embodiments provided by the present application, it should be understood that the provided computer device, computer program product, and computer-readable storage medium are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the methods provided above, and will not be elaborated here.
[0109] It should be noted that the above order of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0110] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.
Claims
1. A method for AI recognition and detection of illegal videos submitted for review, characterized in that: The method comprises: Extract all the video key frames from the video submitted for review; In the video submitted for review, the initial feature points in each video frame are obtained based on the feature extraction method; the inter-frame stability of each video key frame is determined based on the similarity of the initial feature point distribution between each video key frame and the adjacent video frame; the key frame validity of each video key frame is determined based on the inter-frame stability and the density of the initial feature points in each video key frame; and the valid key frames are screened out based on the key frame validity; In each valid key frame, according to the concentrated distribution of feature points and contrast in the local area where each initial feature point is located, a behavior evaluation index of each initial feature point in each valid key frame is determined; according to the behavior evaluation index, all behavior feature points in each valid key frame are screened out; according to the concentrated distribution of behavior feature points in each valid key frame, a behavior feature area in each valid key frame is divided; According to the overall size of the behavior evaluation index of each behavior feature point in the behavior feature area and the density of the behavior feature points, the modified bilateral filtering spatial weight of each behavior feature area is determined; after filtering the behavior feature area according to the modified bilateral filtering spatial weight, the enhanced key frame corresponding to each valid key frame is obtained; according to the enhanced key frame, AI recognition and detection of illegal videos submitted for review are performed.
2. According to claim 1, the method for detecting illegal videos submitted for review by AI is characterized in that: The process of obtaining the inter-frame stability includes: In the video submitted for review, the average value between the initial number of feature points of each video key frame and the initial number of feature points of the corresponding previous video frame is used as the first reference number value; the first matching number of feature points between each video key frame and the corresponding previous video frame is obtained by the FLANN feature matching algorithm; the negative correlation mapping value of the difference between the first matching number of feature points and the first reference number value is used as the first stability level of each video key frame; After replacing the previous video frame in the process of obtaining the first stability with the next video frame, determining the second stability between each video key frame and the next video frame; taking the minimum value between the first stability and the second stability as the stability of the feature point of each video key frame; The difference between the grayscale value of each pixel in each video key frame and the grayscale value of the pixel at the same pixel position in the corresponding previous video frame is used as the first local grayscale difference of each pixel in each video key frame; the mean of the first local grayscale differences of all pixels in each video frame is negatively correlated and mapped to determine the first overall grayscale stability between each video key frame and the previous video frame; After replacing the previous video frame in the process of obtaining the first overall grayscale stability with the next video frame, determining the second overall grayscale stability between each video key frame and the next video frame; taking the minimum value between the first overall grayscale stability and the second overall grayscale stability as the grayscale feature stability degree of each video key frame; The inter-frame stability of each video key frame is determined according to the product of the stability of the feature point and the stability of the grayscale feature.
3. According to claim 1, the AI recognition and detection method for illegal videos submitted for review is characterized in that: The process of obtaining the key frame validity degree includes: The product of the number of initial feature points in each video key frame and the inter-frame stability is normalized to determine the key frame effectiveness of each video key frame.
4. According to claim 1, the method for detecting illegal videos submitted for review by AI is characterized in that: The process of obtaining the valid key frame includes: The video key frames whose key frame validity degree is greater than a preset validity threshold are regarded as valid key frames.
5. According to claim 1, the method for detecting illegal videos submitted for review by AI is characterized in that: The process of obtaining the behavior evaluation index includes: Perform canny edge detection on each valid key frame, and obtain all edge connected domains corresponding to each valid key frame based on the image edges in the obtained edge image; The average number of initial feature points of all edge connected domains in each valid key frame is used as the reference number average; the ratio between the number of initial feature points in the edge connected domain where each initial feature point in each valid key frame is located and the reference number average is used as the local feature point density of each initial feature point in each valid key frame; In each valid key frame, the product of the contrast of each initial feature point and the density of the local feature points is normalized to determine the behavior evaluation index of each initial feature point in each valid key frame.
6. According to claim 1, the method for detecting illegal videos submitted for review by AI is characterized in that: The process of acquiring the behavior feature points includes: In each valid key frame, the initial feature point whose behavior evaluation index is greater than the preset behavior threshold is taken as the behavior feature point.
7. According to claim 5, the method for detecting illegal videos submitted for review by AI is characterized in that: The process of obtaining the behavior feature area includes: In each valid key frame, an edge connected domain whose number of behavior feature points is greater than a preset number threshold is used as the behavior feature region of each valid key frame.
8. According to claim 1, the method for detecting illegal videos submitted for review by AI is characterized in that: The process of obtaining the modified bilateral filtering spatial weight includes: In each valid key frame, the average of the behavior evaluation indexes of all behavior feature points in each behavior feature area is used as the regional behavior feature value of each behavior feature area; In each valid key frame, the average of the number of initial feature points in all behavior feature areas is used as the corresponding reference feature point density; the ratio between the number of initial feature points in each behavior feature area and the reference feature point density is used as the feature point density weight of each behavior feature area in each valid key frame; The product of the regional behavior feature value, the feature point density weight and the prior bilateral filtering spatial weight is used as the modified bilateral filtering spatial weight of each behavior feature region in each valid key frame.
9. According to claim 1, the method for detecting illegal videos submitted for review by AI is characterized in that: The process of acquiring the enhanced key frame includes: In each valid key frame, bilateral filtering is performed on the corresponding behavior feature region by using the modified bilateral filtering spatial weight to obtain a filter enhancement region of each behavior feature region; After each behavior feature region in each valid key frame is replaced with the corresponding filter enhancement region, an enhanced key frame corresponding to each valid key frame is obtained.
10. An AI recognition and detection system for illegal videos submitted for review, characterized in that: The system comprises: The data acquisition module is used to extract all the video key frames in the video submitted for review; The effective key frame acquisition module is used to obtain the initial feature points in each video frame in the video submitted for review based on the feature extraction method; determine the inter-frame stability of each video key frame according to the similarity of the initial feature point distribution between each video key frame and the adjacent video frame; determine the key frame effectiveness of each video key frame according to the inter-frame stability and the initial feature point density in each video key frame; and screen out the effective key frames according to the key frame effectiveness; The behavior feature area acquisition module is used to determine the behavior evaluation index of each initial feature point in each valid key frame according to the concentrated distribution of feature points in the local area where each initial feature point is located and the contrast; filter out all behavior feature points in each valid key frame according to the behavior evaluation index; and divide the behavior feature area in each valid key frame according to the concentrated distribution of behavior feature points in each valid key frame; The AI recognition and detection module is used to determine the modified bilateral filtering spatial weight of each behavior feature area according to the overall size of the behavior evaluation index of each behavior feature point in the behavior feature area and the density of the behavior feature points; after filtering the behavior feature area according to the modified bilateral filtering spatial weight, an enhanced key frame corresponding to each valid key frame is obtained; and AI recognition and detection of illegal videos submitted for review are performed according to the enhanced key frame.
Citation Information
Patent Citations
Synchronous key frame extraction video splicing method based on binocular camera
CN110120012A
Intelligent detection method for aerial video image of unmanned aerial vehicle
CN115439424A