Video auditing method and device, electronic equipment, storage medium and program product
By employing a layered and progressive video review mechanism and a feature library for borderline content, the problem of accurately identifying gray-area information in video review has been solved, improving review efficiency and accuracy, reducing the risk of misjudgment and omission, and providing a reliable basis for compliance penalties.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-04-07
AI Technical Summary
Existing video review systems struggle to accurately identify gray-area content, leading creators to deliberately adjust or hide inappropriate content, increasing the difficulty of review and impacting the efficiency of normal users.
A layered and progressive video review mechanism is adopted, including first-level review, multi-round progressive second-level review and third-level review. Combined with a pre-built feature library of borderline content, the mechanism gradually identifies borderline content by marking image frames according to the playback sequence and performing texture feature extraction and similarity matching.
It has achieved a dual improvement in the efficiency and accuracy of video review, reduced the risk of misjudgment and omission, provided a reliable basis for compliance penalties, and saved review costs.
Smart Images

Figure CN121814985A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of Internet technology, and in particular to a video review method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] Currently, cloud storage UGC content review is generally conducted using intelligent systems. Compared to the previous reliance on manual review, this has greatly saved costs and improved efficiency, while also contributing to social development. Existing review methods generally use prohibited images and sensitive features to train review models, and this approach is already widely used.
[0003] However, it is difficult to accurately judge some data that is in a gray area, especially since many creators and users will deliberately adjust or hide illegal content in the uploaded data in order to avoid cloud storage review. This is especially true for video review, which not only greatly increases the difficulty of content review, but also the lack of an accurate review system will affect the efficient use of normal users due to the slow review process. Summary of the Invention
[0004] In view of the above-mentioned technical problems, this disclosure provides a video review method, apparatus, electronic device, storage medium and program product.
[0005] Firstly, this disclosure provides a video review method, including: Obtain the set of image frames from the video containing borderline content, and sequentially label each image frame in the set with a frame number according to the playback sequence of the video containing borderline content; Based on the initial review intensity coefficient, the first image frame with the first target quantity is selected from the image frame set, and the first image frame is reviewed in a first-level manner in combination with the pre-built edge-trimming feature library; If the first-level review determines that there is no violation, the number of second-level reviews is determined based on the initial review intensity coefficient. Based on the number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames of the target review frame number corresponding to the increased intensity coefficient. Combined with the edge-checking feature library, the key image frames of the target review frame number selected each time are reviewed in the second level. If the second-level review detects the first non-compliant image frame, the first non-compliant image frame is spread, and the edge-trimming feature library is called to conduct a third-level review on the non-compliant associated image frame set formed after the spread. Based on the results of the three-level review, a decision will be made as to whether to penalize videos containing borderline content.
[0006] Furthermore, according to the method of the first aspect of this disclosure, the initial review intensity coefficient is determined based on the basic weight, the reporting rate, and the account's historical violation rate.
[0007] Furthermore, according to the method of the first aspect of this disclosure, based on an initial review intensity coefficient, a first number of first image frames of a first target quantity are selected from the image frame set, and the first image frames are subjected to a first-level review in conjunction with a pre-built edge-trimming feature library, including: The number of first targets is determined based on the initial review intensity coefficient and the number of image frames in the image frame set. Image frames with an odd number of frames and whose number is equal to the number of first targets are selected from the image frame set and used as the first image frames. Extract texture features from the first image frames of the first target quantity to obtain the texture feature set corresponding to each first image frame; Based on the texture feature set corresponding to each first image frame and the feature data in the pre-built edge-scratching feature library, similarity matching is performed on each first image frame to obtain the similarity matching result corresponding to each first image frame; Based on the similarity matching results and preset thresholds corresponding to each first image frame, it is determined whether each first image frame violates the rules.
[0008] Furthermore, according to the method of the first aspect of this disclosure, the method also includes: If the first-level review result indicates a violation, the video containing borderline content will be blocked, and the video will be penalized according to the preset initial penalty rules.
[0009] Furthermore, according to the method of the first aspect of this disclosure, the initial penalty rules include: If the initial violation precision is greater than 0.2, light restrictions will be imposed on videos containing borderline content. If the initial violation precision is greater than 0.4, moderate restrictions will be imposed on videos containing borderline content. If the initial violation precision is greater than 0.6, videos containing borderline content will be severely punished; The initial violation precision is determined based on the number of violation frames and the number of first targets discovered during the first-level review process.
[0010] Furthermore, according to the method of the first aspect of this disclosure, if the first-level review determines no violations, the number of reviews for the second-level review is determined based on the initial review intensity coefficient. Based on the number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames corresponding to the target review frame number after the increase in intensity coefficient. Combining the edge-checking feature library, the key image frames of the target review frame number selected each time are sequentially reviewed for the second-level review, including: Step 1: If the first-level review result is no violation, determine the number of second-level reviews based on 10 times the initial review intensity coefficient; Step 2: Based on the preset incremental rules and the initial review intensity coefficient, generate the incremental review intensity coefficient for this round; Step 3: Based on the incremental review intensity coefficient of this round, the number of image frames in the image frame set, and the odd-even frame alternation rule, extract the key image frames corresponding to the target review frame number from the image frame set; Step 4: Call the edge-scratching feature library to perform secondary review on key image frames of the target review frame number; Step 5: Repeat steps 2 to 4 above until the number of audits determined in step 1 is completed.
[0011] Furthermore, according to the method of the first aspect of this disclosure, if the second-level review detects a first non-compliant image frame, the first non-compliant image frame is diffused, and the edge-scratching feature library is invoked to perform a third-level review on the non-compliant associated image frame set formed after diffusion, including: If the second-level review detects the first non-compliant image frame, based on the review results of the second-level review, extract all first non-compliant image frames corresponding to the second-level review. Based on the preset number of diffusions and the preset diffusion range, the diffusion extends forward and backward with all the first violation image frames as the center, and determines the set of violation-related image frames generated by each diffusion. The edge-check feature library is invoked to perform a three-level review on each set of related image frames that violate the rules.
[0012] Furthermore, based on the method of the first aspect of this disclosure, and based on the judgment results of the three-level review, it is determined whether to penalize videos containing borderline content, including: Based on the results of the three-level review, the final violation precision is determined. The final violation precision is determined based on the number of violation frames found during the second-level review, the number of violation frames found during the third-level review, and the target review frame corresponding to the first violation image frame found during the second-level review. Based on the final precision of the violation, it will be determined whether to penalize the videos containing borderline content.
[0013] Furthermore, based on the method of the first aspect of this disclosure, and based on the final violation precision, it is determined whether to penalize videos containing borderline content, including: If the final level of violation exceeds the first penalty threshold, the first-level penalty will be applied. If the final level of violation exceeds the second penalty threshold, the second-level penalty will be imposed on the video containing borderline content. If the final level of violation exceeds the third penalty threshold, the video containing borderline content will be subject to a third-level penalty.
[0014] Furthermore, according to the method of the first aspect of this disclosure, the edge-trimming feature library is constructed based on the following: Based on historical review data, obtain samples of borderline content, illegal video frames, and illegal images; Multi-dimensional visual feature extraction is performed on content border-skimming samples, illegal video frames, and illegal images to obtain the first border-skimming feature set corresponding to the content border-skimming samples, the second border-skimming feature set corresponding to the illegal video frames, and the third border-skimming feature set corresponding to the illegal images. An edge-scratching feature library is constructed based on the first edge-scratching feature set, the second edge-scratching feature set, and the third edge-scratching feature set.
[0015] Furthermore, according to the method of the first aspect of this disclosure, the method also includes: If the third-level review determines that there are no violations, the videos containing borderline content will undergo manual review.
[0016] Secondly, this disclosure provides a video review device, including: The acquisition module is used to acquire a set of image frames from the video containing borderline content, and to sequentially label each image frame in the set with a frame number according to the playback sequence of the video containing borderline content. The first-level review module is used to select the first target number of first image frames from the image frame set based on the initial review intensity coefficient, and to conduct first-level review of the first image frames in combination with the pre-built edge-rubbing feature library; The secondary review module is used to determine the number of reviews for the secondary review based on the initial review intensity coefficient if the primary review determines that there is no violation. Based on the number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames of the target review frame number corresponding to the increased intensity coefficient. Combined with the edge-checking feature library, the key image frames of the target review frame number selected each time are reviewed in the secondary review. The Level 3 review module is used to spread the first illegal image frame if the Level 2 review detects the first illegal image frame, and call the edge-trimming feature library to perform Level 3 review on the illegal associated image frame set formed after the spread; The penalty module is used to determine whether to penalize videos containing borderline content based on the results of the three-level review.
[0017] Furthermore, according to the apparatus of the second aspect of this disclosure, the initial review intensity coefficient is determined based on the basic weight, the reporting rate, and the account's historical violation rate.
[0018] Furthermore, according to the apparatus of the second aspect of this disclosure, the primary review module is also used for: The number of first targets is determined based on the initial review intensity coefficient and the number of image frames in the image frame set. Image frames with an odd number of frames and whose number is equal to the number of first targets are selected from the image frame set and used as the first image frames. Extract texture features from the first image frames of the first target quantity to obtain the texture feature set corresponding to each first image frame; Based on the texture feature set corresponding to each first image frame and the feature data in the pre-built edge-scratching feature library, similarity matching is performed on each first image frame to obtain the similarity matching result corresponding to each first image frame; Based on the similarity matching results and preset thresholds corresponding to each first image frame, it is determined whether each first image frame violates the rules.
[0019] Furthermore, according to the apparatus of the second aspect of this disclosure, the primary review module is also used for: If the first-level review result indicates a violation, the video containing borderline content will be blocked, and the video will be penalized according to the preset initial penalty rules.
[0020] Furthermore, according to the apparatus of the second aspect of this disclosure, the initial penalty rules include: If the initial violation precision is greater than 0.2, light restrictions will be imposed on videos containing borderline content. If the initial violation precision is greater than 0.4, moderate restrictions will be imposed on videos containing borderline content. If the initial violation precision is greater than 0.6, videos containing borderline content will be severely punished; The initial violation precision is determined based on the number of violation frames and the number of first targets discovered during the first-level review process.
[0021] Furthermore, according to the apparatus of the second aspect of this disclosure, the secondary review module is also used for: Step 1: If the first-level review result is no violation, determine the number of second-level reviews based on 10 times the initial review intensity coefficient; Step 2: Based on the preset incremental rules and the initial review intensity coefficient, generate the incremental review intensity coefficient for this round; Step 3: Based on the incremental review intensity coefficient of this round, the number of image frames in the image frame set, and the odd-even frame alternation rule, extract the key image frames corresponding to the target review frame number from the image frame set; Step 4: Call the edge-scratching feature library to perform secondary review on key image frames of the target review frame number; Step 5: Repeat steps 2 to 4 above until the number of audits determined in step 1 is completed.
[0022] Furthermore, according to the apparatus of the second aspect of this disclosure, the three-level audit module is also used for: If the second-level review detects the first non-compliant image frame, based on the review results of the second-level review, extract all first non-compliant image frames corresponding to the second-level review. Based on the preset number of diffusions and the preset diffusion range, the diffusion extends forward and backward with all the first violation image frames as the center, and determines the set of violation-related image frames generated by each diffusion. The edge-check feature library is invoked to perform a three-level review on each set of related image frames that violate the rules.
[0023] Thirdly, this disclosure provides an electronic device, including: a memory for storing computer-readable instructions; and a processor for executing the computer-readable instructions, causing the electronic device to perform the method as described in any embodiment of the first aspect.
[0024] Fourthly, this disclosure provides a non-transitory computer-readable storage medium for storing computer-readable instructions that, when executed by a processor, cause the processor to perform the method as described in any embodiment of the first aspect.
[0025] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any embodiment of the first aspect.
[0026] This disclosure provides a video review method, apparatus, electronic device, storage medium, and program product. By marking image frames according to playback sequence, this disclosure constructs a hierarchical, progressive mechanism: "Level 1 initial screening – Level 2 progressively stronger multi-round review – Level 3 focused review of violation frames." Combined with a pre-built feature library, it achieves accurate feature matching. The Level 1 review, with its low intensity and few frames, ensures basic review efficiency and avoids impacting normal user experience. The Level 2 review, with its progressively stronger multi-round review, expands the video frame coverage, effectively addressing the review challenges caused by the "hiding / adjustment" of violation content. The Level 3 review, with its targeted review of the first violation image frame, enhances the accuracy of identifying borderline content. Ultimately, this achieves a dual improvement in efficiency and accuracy in video review scenarios, significantly reducing the risk of misjudgment and omission of borderline videos. Simultaneously, differentiated review resource allocation saves costs and provides a reliable basis for compliance penalties.
[0027] It should be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further illustration of the claimed technology. Attached Figure Description
[0028] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0029] Figure 1 A flowchart illustrating a video review method provided in this embodiment of the disclosure; Figure 2 A flowchart illustrating yet another video review method provided in this disclosure embodiment; Figure 3 A flowchart illustrating yet another video review method provided in this disclosure embodiment; Figure 4 A flowchart illustrating yet another video review method provided in this disclosure embodiment; Figure 5 A flowchart illustrating yet another video review method provided in this disclosure embodiment; Figure 6 A structural block diagram of a video review device provided in this embodiment of the present disclosure; Figure 7 A hardware block diagram of an electronic device provided in an embodiment of this disclosure; Figure 8 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of this disclosure. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0031] Mobile cloud storage is an application within a cloud storage system, which itself is an application developed from cloud computing technology. The core of cloud storage is data storage and management. It is configured with massive storage space on the basis of cloud computing systems. With the support of technologies such as cluster systems, grid technology, and distributed file systems, cloud storage systems can enable large-scale storage devices across regions to work together and provide services to the outside world. The existence of various application programming interfaces (APIs) in cloud storage systems allows developers to continuously expand the types of services that cloud storage systems can provide by developing different applications. Currently, the main services that cloud storage systems can provide include three categories: cloud storage, space rental services, and remote backup and disaster recovery. Among these, cloud storage applications are the most closely related to ordinary internet users.
[0032] Currently, the main task of content security operation for cloud storage is still handled by enterprise security departments. However, with the rapid development of cloud storage services, the traditional manual method of monitoring and reviewing large amounts of rich media content uploaded by users is clearly no longer efficient enough to meet the security needs of enterprises. Therefore, there is an urgent need for a dynamic frame extraction method in video content review to reduce the amount of manual work while ensuring the secure and efficient operation of cloud storage.
[0033] Currently, cloud storage UGC content review is widely conducted using intelligent systems. Compared to the previous reliance on manual review, this has significantly reduced costs and improved efficiency, while also contributing to social development. Existing review methods generally use prohibited images and sensitive features to train review models, which are widely used. However, for some data that falls into gray areas, it is difficult to judge accurately. In particular, many creators and users deliberately adjust or hide prohibited content in their uploaded data to avoid cloud storage review, especially for video review. This not only greatly increases the difficulty of content review, but the lack of an accurate review system also slows down the review process, affecting the efficient use of the platform by normal users.
[0034] Based on the aforementioned technical issues, this disclosure provides a video review method.
[0035] Figure 1 This is a flowchart illustrating a video review method provided in an embodiment of the present disclosure.
[0036] like Figure 1 As shown, the video review method specifically includes the following steps: Step 101: Obtain the set of image frames from the video containing borderline content, and sequentially label each image frame in the set with a frame number according to the playback sequence of the video containing borderline content.
[0037] In one embodiment of this disclosure, "borderline content videos" refer to short videos and other image content that skirt the line between law and platform rules, attracting attention through vulgarity, sexual innuendo, etc., but not reaching the explicit level of illegality such as obscenity or pornography. There is no unified legal definition standard for this type of content; it is judged more based on the context of dissemination and public order and good morals. The videos to be reviewed undergo routine review, including: user upload, content deduplication, feature extraction, feature fusion, sample library import feature matching, system review, and review result steps. The review results include content violation, approval, and borderline content. Videos judged as borderline content videos are then defined as such. A set of single-frame images containing complete visual information extracted from the borderline content videos is called an image frame set. Each image frame (i.e., a single-frame image) in the image frame set is sequentially labeled with frame number 1, 2, 3...n, where n is the total number of frames.
[0038] Step 102: Based on the initial review intensity coefficient, select the first image frame with the first target quantity from the image frame set, and conduct a first-level review on the first image frame in combination with the pre-built edge-trimming feature library.
[0039] In one embodiment of this disclosure, the initial review intensity coefficient is determined based on a base weight, a report rate, and the account's historical violation rate. The base weight, adjustable by the administrator, ranges from 0 to 1. The report rate is the ratio of reports to views for borderline content videos. The account's historical violation rate refers to the historical violation rate of the account to which the borderline content video belongs, determined by the ratio of the number of historical violations to the total number of historical uploads. For example, if the initial review intensity coefficient is set to m, then the initial review intensity coefficient m = base weight + report rate + (number of historical violations / total number of historical uploads). The first target quantity is a pre-set quantitative parameter used to define the "total number of key image frames to be screened in the first-level review stage," which can be determined in conjunction with the initial review intensity coefficient; the specific method is not limited.
[0040] In one embodiment of this disclosure, the edge-trimming feature library is constructed based on the following: Based on historical review data, obtain samples of borderline content, illegal video frames, and illegal images; Multi-dimensional visual feature extraction is performed on content border-skimming samples, illegal video frames, and illegal images to obtain the first border-skimming feature set corresponding to the content border-skimming samples, the second border-skimming feature set corresponding to the illegal video frames, and the third border-skimming feature set corresponding to the illegal images. An edge-scratching feature library is constructed based on the first edge-scratching feature set, the second edge-scratching feature set, and the third edge-scratching feature set.
[0041] Specifically, the "borderline" feature library is the core database of this publication. It differs from traditional feature libraries that store explicitly prohibited content (such as gory, violent, or pornographic images). Its design goal is to identify ambiguous, covert, and difficult-to-determine "gray" or "borderline" content. It is a structured collection of high-dimensional feature vectors. Simply put, it doesn't store raw prohibited images or videos, but rather mathematical features extracted from a large amount of known "borderline" content that represent its visual essence. These features typically include texture, color distribution, shape, and key points. Historical review data refers to past review records of videos and images. "Borderline" content samples refer to content that doesn't reach a clear level of violation but falls into a gray area; prohibited video frames and images refer to video frames and images confirmed to be prohibited through manual review. Multi-dimensional visual features include, but are not limited to, texture, color distribution, shape, and key points. The first set of borderline features refers to the set of high-dimensional feature vectors extracted from "content borderline samples" that represent "gray area borderline attributes"; the second set of borderline features refers to the set of high-dimensional feature vectors extracted from "violation video frames" that represent "confirmed violation borderline attributes"; and the third set of borderline features refers to the set of high-dimensional feature vectors extracted from "violation images" that represent "confirmed violation borderline attributes". Two types of seed samples are obtained: "content borderline samples that do not reach a clear level of violation but belong to the gray area" and "video frames and images confirmed to be in violation through manual review". Multi-dimensional visual features such as texture, color distribution, shape, and key points are extracted from these samples and transformed into high-dimensional feature vectors, forming three corresponding sets of borderline features. These sets are then stored in a structured manner, that is, stored in a data structure that facilitates rapid searching and comparison. This disclosure explicitly uses a KD-tree to organize these features, ultimately constructing a dedicated feature library for identifying ambiguous and concealed "borderline" content.
[0042] In one embodiment of this disclosure, a specified number (i.e., a first target number) of keyframes are selected from the image frame set as first image frames based on the aforementioned initial review strength coefficient. The first image frames are then compared with a border-skimming feature library to determine the border-skimming risk of the first image frames, thus completing the first-level review.
[0043] Step 103: If the first-level review determines that there is no violation, determine the number of reviews for the second-level review based on the initial review intensity coefficient. Based on the number of reviews, increase the initial review intensity coefficient round by round to determine the key image frames of the target review frame number corresponding to the increased intensity coefficient. Combined with the edge-checking feature library, conduct second-level review on the key image frames of the target review frame number selected each time.
[0044] In one embodiment of this disclosure, a key image frame refers to an image frame selected for review in each round of secondary review. The target number of review frames refers to the number of key image frames selected for review in each round of secondary review, determined based on the increasing review intensity coefficient for each round. During the primary review, secondary review is initiated only when all first image frames are deemed "no violation" after primary review. This is for a "secondary fallback check" (to prevent primary review from overlooking highly concealed borderline content due to insufficient frame selection or low intensity). Secondary review is not a single review but a "multi-round, progressively stronger" review. The number of reviews is determined by the initial review intensity coefficient, avoiding indiscriminate multiple rounds of review that waste computational resources. Specifically, the intensity of each round of secondary review is higher than the previous round, and this increase in intensity is directly reflected in an "increase in the number of review frames," ensuring that the progressive rounds are meaningful. The number of reviews for the second-level review is determined by the initial review intensity coefficient. Based on the number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames of the target review frame number corresponding to the increased intensity coefficient. Using the edge-checking feature library, the key image frames of the target review frame number selected each time are compared for features. The second-level review is completed according to the number of reviews.
[0045] Step 104: If the second-level review detects the first illegal image frame, the first illegal image frame is diffused, and the edge-trimming feature library is called to conduct a third-level review on the illegal associated image frame set formed after diffusion.
[0046] In one embodiment of this disclosure, the first violating image frame is an image frame detected and judged to be violating during the second-level review. Diffusion refers to using the first violating image frame as the core to mine other image frames that are related to it (i.e., a set of violating-related image frames). During the second-level review process, if a violating image frame is detected, diffusion is performed on all the violating image frames detected in the second-level review to obtain a set of violating-related image frames. A borderline feature library is then called to perform feature comparison on the set of violating-related image frames, completing the third-level review.
[0047] Step 105: Based on the results of the three-level review, determine whether to penalize the videos containing borderline content.
[0048] In one embodiment of this disclosure, a determination is made, based on the results of the three-level review, whether to impose penalties on videos containing borderline content.
[0049] In summary, based on the technical solution provided in the embodiments of this disclosure, this disclosure constructs a hierarchical and progressive mechanism by marking image frames according to playback sequence: "first-level initial screening - second-level progressive multi-round review - third-level focused review of violation frames." Combined with a pre-built feature library for borderline content, it achieves accurate feature matching. The first-level review, with its low intensity and few frames, ensures basic review efficiency and avoids affecting normal user experience. The second-level review, with its progressively stronger multi-round review, expands the coverage of video frames, effectively solving the review difficulties caused by the "hiding / adjustment" of violation content. The third-level review, with its targeted review of the first violation image frame, enhances the accuracy of identifying borderline content. Ultimately, this achieves a dual improvement in efficiency and accuracy in video review scenarios, significantly reducing the risk of misjudgment and omission of borderline videos. At the same time, it saves costs through differentiated review resource allocation and provides a reliable basis for compliance penalties.
[0050] Figure 2 This is a flowchart illustrating another video review method provided in an embodiment of the present disclosure.
[0051] like Figure 2 As shown, based on the initial review intensity coefficient, the first image frame with the first target quantity is selected from the image frame set, and the first image frame is subjected to first-level review in combination with the pre-built edge-trimming feature library. The specific steps include the following: Step 201: Determine the first target quantity based on the initial review strength coefficient and the number of image frames in the image frame set. Select image frames from the image frame set whose quantity is equal to the first target quantity and whose frame number is odd, and use them as the first image frames.
[0052] In one embodiment of this disclosure, the first target quantity is determined by the product of the initial review strength coefficient and the number of image frames in the image frame set, i.e., the first target quantity = m×n, and the m×n number of image frames at position 2n-1 (i.e., the first image frames) are selected from the image frame set.
[0053] Step 202: Extract texture features from the first image frames of the first target number to obtain the texture feature set corresponding to each first image frame.
[0054] In one embodiment of this disclosure, the texture features include: roughness (reflecting the graininess and fineness of an image. The texture of certain special clothing or censored areas may have a specific roughness), contrast (reflecting the difference in brightness and darkness of an image. Certain illegal content that deliberately adjusts brightness and contrast to hide details will exhibit abnormal contrast features), and orientation (reflecting the dominant direction of lines and edges in an image. For example, certain patterns or arrangements will exhibit strong orientation).
[0055] Specifically, the steps for texture feature extraction and roughness calculation are as follows: a. Calculate the image size as The formula for the average grayscale value of pixels in the active window of pixels is:
[0056] Where k = 0, 1, 2, ... 5, It is located in The grayscale value of the pixel.
[0057] b. For each pixel, calculate the average grayscale difference between windows that do not intersect in the vertical and horizontal directions, using the following formula:
[0058] For each pixel, the k value that maximizes the E value is used to set the optimal size. .
[0059] c. Roughness is obtained by calculating the average value of an image of size m×n (where m is the number of rows and n is the number of columns). The calculation formula is:
[0060] The formula for calculating contrast during texture feature extraction is as follows:
[0061] in, It is obtained through statistics on pixel intensity distribution. Defined by, among which It is the fourth moment of the pixel values at each point in the image. It is variance.
[0062] The steps for calculating the directionality in texture feature extraction are as follows: a. Gradient vector for each pixel:
[0063] The calculation of △H and △V involves convolving the image with the following two 3x3 formulas, where △H is the change in the gradient vector in the horizontal direction and △V is the change in the gradient vector in the vertical direction. Used to calculate ΔH:
[0064] Used to calculate ΔV:
[0065] Once the pixel gradient vector is calculated, the θ value can be expressed using a histogram HD. b. Discretize the range of θ values using a histogram; c. Count the number of pixels in each handle whose corresponding |△G| and |△G| are greater than a given threshold. Images with obvious directionality will have peaks, while images without obvious directionality will be flat. d. The overall directionality of an image can be obtained by calculating the sharpness of the peaks in its histogram, using the following formula:
[0066] In the above formula This represents the overall orientation of the image, where p represents the peak value in the histogram. For all the peaks in the histogram, for a given peak p, This represents all the discrete regions encompassed by the peak, while It is the center of the wave crest. It is the specific value of a discrete region in the histogram.
[0067] The aforementioned texture feature extraction method is also applicable to edge-scratching feature libraries.
[0068] Step 203: Based on the texture feature set corresponding to each first image frame and the feature data in the pre-built edge-trimming feature library, perform similarity matching on each first image frame to obtain the similarity matching result corresponding to each first image frame.
[0069] In one embodiment of this disclosure, the texture feature set corresponding to each first image frame is used as a query point. This query point is input into a pre-constructed edge-scratching feature library for feature comparison. Based on the Best-Bin-First (BBF) algorithm, the KD-tree of the edge-scratching feature library is traversed to find the one or more feature vectors closest to the query point. Here, "distance" refers to the Euclidean distance or other metric between two points (i.e., two feature vectors) in the feature space, representing the visual similarity of their corresponding images. This distance is used as the similarity matching result. The closer the distance, the more similar the two images are in terms of texture features.
[0070] Specifically, a KD-tree is a data structure that organizes points in a k-dimensional Euclidean space. A KD-tree partitions data in a k-dimensional feature space, with each node being a k-dimensional binary tree corresponding to a spatial range. When building a KD-tree, the dimension with the largest variance in the dataset is first determined, and the median of the dataset in that dimension is found. This median is used as the root node of the KD-tree, dividing the space into two parts: those smaller than the root node in that dimension are called the left child nodes, and those larger are called the right child nodes. The same method is then used to continue building binary trees for the left and right child nodes until the remaining dataset is empty. Traditional KD-tree queries are based on tree traversal, as follows: starting from the root node, search along the binary tree until a leaf node is reached. At this point, the leaf node is not necessarily the nearest point, but it is definitely a point near a leaf node, so backtracking is required until the root node is reached again.
[0071] The algorithm using KD trees is relatively simple. It adds nodes that might be traversed during backtracking to a queue, sorting them according to the distance from the search point to the hyperplane determined by that node. Each time, it traverses the node with the highest priority (i.e., the shortest distance), continuing until the queue is empty. The algorithm also has a time limit; if the algorithm's runtime exceeds this limit, it stops running regardless of whether the queue is empty. After the program finishes running, the latest nearest neighbor is returned as feedback. The logical flow of this disclosure using the BBF algorithm is as follows: The texture feature set corresponding to each first image frame is used as a query point. This "query point" is input into the pre-constructed edge-trimming feature library. When the process determines that the KD value is empty, the distance between the target and the KD is maximized, and the specific value takes values in an infinite range, and the result is returned. When the KD tree is not empty, the root node value of the KD tree is automatically obtained and added to the queue, with a priority level set. The process involves searching for the maximum value in the priority queue, reading the node corresponding to that maximum value, and determining whether the distance between that node and the target node is the minimum (compare each node sequentially, discarding the smaller one; after one loop, the node with the minimum distance is obtained). While calculating the minimum distance, it simultaneously checks if the search point's coordinates in the split dimension are minimum. If they are, the node's right subtree is added to the queue, and the left subtree is calculated using the same method. If the right subtree cannot be added to the queue, the left subtree is added. This process is repeated until all leaf nodes of the KD-tree have been calculated and checked, and finally, all... For nodes already in the priority queue, read the maximum value, which depends on the priority. Following the above method, proceed to the next round of loop judgment and calculation. The goal is to extract all nodes in the existing priority queue according to their priority. Once the loop ends, return the nearest neighbor node and its distance. For system efficiency, a maximum time limit is typically set for this loop. If the time limit expires before the loop ends, the loop still terminates, and the nearest neighbor node and its distance are returned. Finally, output the KD-tree node closest to the query target and its corresponding nearest distance value (D). This nearest distance value is used as the similarity matching result.
[0072] Step 204: Based on the similarity matching results and preset thresholds corresponding to each first image frame, determine whether each first image frame violates the rules.
[0073] In one embodiment of this disclosure, the similarity matching result is the "nearest distance value" in step S203 above. The preset threshold is a quantitative critical value used to determine whether "two objects are sufficiently similar," and it can be determined through extensive experimentation. Determining whether each first image frame violates a rule can be done in the following way: Set the preset threshold as T. If D < T, it indicates that the picture to be reviewed (i.e., the first image frame) is very similar to a certain known "borderline" content in the borderline video feature library, and it is marked as a violation. If D ≥ T, it means that the picture to be reviewed is not similar enough to all known "borderline" contents, and it is temporarily marked as "passed" or "no abnormality".
[0074] In actual applications, according to the review intensity coefficient m, when D < T, the picture to be reviewed can be marked as "suspected violation" or directly determined as "violation".
[0075] In summary, according to the technical solution provided by the embodiments of the present disclosure, the present disclosure screens the first image frames of odd frames based on the initial review intensity coefficient, extracts their texture features (roughness, contrast, directionality), then traverses the K-D tree of the borderline feature library through the best-first algorithm (BBF) to perform similarity matching on the first image frames, and finally determines whether there is a violation based on the comparison between the matching result and the preset threshold. The above method realizes accurate and efficient video review of borderline content through targeted image frame screening, multi-dimensional texture feature extraction, and efficient similarity retrieval algorithm. While ensuring the accuracy of the review, it improves the review efficiency, can effectively identify illegal borderline content, and reduce the situation of missed and misjudged reviews.
[0076] Further, the above video review method further includes: If the first-level review result is judged as a violation, intercept the borderline content video and impose a penalty on the borderline content video according to the preset initial penalty rule.
[0077] Specifically, the first-level review result being judged as a violation means that in the first-level review process of the above video review, after a series of operations such as image frame screening, texture feature extraction, and similarity matching with the borderline feature library, the judgment conclusion is obtained. Specifically, when extracting the texture features such as roughness, contrast, and directionality of the selected first image frame and using these features as query points to perform similarity matching with the pre-constructed borderline feature library through the BBF algorithm to traverse the K-D tree, if the obtained nearest distance value D is less than or equal to the preset threshold T, it indicates that the image frame of the video is highly similar to the known "borderline" content in the borderline feature library, and thus it is determined that the borderline content video is an illegal content in the first-level review. Furthermore, intercept the borderline content video and impose a penalty on the borderline content video according to the preset initial penalty rule.
[0078] Further, the initial penalty rule includes: If the initial violation precision is greater than 0.2, impose a mild restriction on the borderline content video; If the initial violation precision is greater than 0.4, impose a medium restriction on the borderline content video; If the initial violation precision is greater than 0.6, videos containing borderline content will be severely punished; The initial violation precision is determined based on the number of violation frames and the number of first targets discovered during the first-level review process.
[0079] Specifically, the initial violation precision is denoted by S, and is calculated as follows:
[0080] in, p This represents the number of violation frames discovered during the first-level review process, where m×n is the number of the first target.
[0081] When S>0.2: Implement mild restrictions, such as blurring out the video.
[0082] When S>0.4: Implement moderate restrictions, such as prohibiting forwarding / sharing.
[0083] When S > 0.6: Severe penalties will be imposed, such as deleting videos and potentially imposing more stringent functional restrictions on the account.
[0084] In summary, based on the technical solutions provided in the embodiments of this disclosure, this disclosure, through a refined violation judgment and graded penalty mechanism, not only effectively intercepts and regulates borderline illegal videos, but also implements differentiated policies according to the degree of violation. While ensuring content compliance, it also takes into account the rationality and flexibility of penalties, improves the accuracy of video review and governance efficiency, and effectively purifies the content ecosystem.
[0085] Figure 3 This is a flowchart illustrating another video review method provided in an embodiment of the present disclosure.
[0086] like Figure 3 As shown, if the first-level review determines no violations, the number of reviews for the second-level review is determined based on the initial review intensity coefficient. Based on the number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames corresponding to the target number of review frames after the increase in intensity coefficient. Combining the edge-checking feature library, the key image frames of the target number of review frames selected each time are sequentially reviewed for the second-level review. The specific steps include the following: Step 301: If the result of the first-level review is no violation, determine the number of reviews for the second-level review by multiplying the initial review intensity coefficient by 10.
[0087] In one embodiment of this disclosure, a first-level review result of "no violation" means that the first image frames of the first target number have no violations. In this case, the number of reviews for the second-level review is calculated based on the initial review intensity coefficient of 10, i.e., the number of reviews = 10m. Here, 10m is a dynamic safety threshold. For high-risk content (large m), more rounds (large 10m) of in-depth investigation are allowed; for low-risk content (small m), the investigation ends quickly to save resources.
[0088] Step 302: Based on the preset incremental rules and the initial review intensity coefficient, generate the incremental review intensity coefficient for this round.
[0089] In one embodiment of this disclosure, the preset incremental rule can be understood as a fixed incremental rule, that is, based on the previous review intensity coefficient, a fixed increment is added to obtain the review intensity coefficient after the current round of increment. The characteristic of this rule is that the increment value is fixed each time, the calculation is simple and intuitive, and it can clearly improve the intensity of the secondary review, conducting a second review of videos with stricter standards and reducing the risk of missed reviews. The determination of the increment is usually based on business needs, risk assessment, and experimental verification. Taking an initial review intensity coefficient of 0.1 increasing to 0.15 as an example, the logic for determining the increment of 0.05 is: the platform, through extensive historical data testing and risk assessment, found that increasing the initial intensity coefficient by 0.05 effectively improves the strictness of the secondary review (reducing the risk of missed reviews) while achieving a balance between review efficiency and resource consumption. This increment setting is an empirical preset rule, a fixed incremental value determined by the platform based on its own business scenarios (such as common violation characteristics of borderline content, user report data, etc.) through repeated experiments and effect verification, thereby achieving the goal of "precisely improving the intensity of the secondary review when there are no violations at the primary level." The incremental review intensity coefficient refers to the dynamic intensity parameter generated in each round of secondary review, based on the initial review intensity coefficient and according to the preset incremental rules, to increase the strictness of the review in that round. The intensity coefficient of each round of review will continue to increase based on the preset incremental rules.
[0090] Step 303: Based on the incremental review intensity coefficient of this round, the number of image frames in the image frame set, and the odd-even frame alternation rule, extract the key image frames corresponding to the target review frame number from the image frame set.
[0091] In one embodiment of this disclosure, the odd-even frame alternation rule refers to the order constraint of frame extraction (to avoid repetition of image information caused by consecutive frame extraction). Specifically, during extraction, frames are selected according to an alternating logic of "odd frame → even frame → odd frame → ..." or "even frame → odd frame → even frame → ..." to ensure that the extracted keyframes cover different time periods of the video, improving the comprehensiveness of the review. Since odd frames are extracted during the first-level review process, the second-level review follows the alternating logic of "even frame → odd frame → even frame → ...". The target review frame count refers to the actual number of images to be sampled in this round, and the target review frame count is based on the increasing review intensity coefficient of this round. The number of image frames (n) in the image frame set is determined (e.g., the target number of review frames = Key image frames refer to the core images selected in this round (i.e., the specific frames corresponding to the "target review frame number"). The initial review sampling scope for the second-level review shifts to the images at position 2n (i.e., even-numbered frames), and then the images that have already been sampled are filtered out.
[0092] Step 304: Call the edge-scratching feature library to perform secondary review on the key image frames of the target review frame number.
[0093] In one embodiment of this disclosure, texture features are extracted from key image frames to obtain texture feature sets for each key image frame. These texture feature sets are then input into an edge-rubbing feature library for similarity matching. Based on the similarity matching results and a preset threshold, it is determined whether each key image frame violates the rules. The specific process described above is the same as steps 202-204, and will not be repeated here.
[0094] Step 305: Repeat steps 302 to 304 above until the number of audits determined in step 301 is completed.
[0095] In one embodiment of this disclosure, steps 302 to 304 are executed repeatedly until the number of audits determined in step 301 is completed, thus completing the secondary audit.
[0096] For example, the initial sampling scope of the second-level review shifts to the images at position 2n, filtering out the already sampled images, and repeating this cycle 10m times until the second-level review is completed. This process is primarily based on the first-level review. Specifically, if the first-level review finds no problems, the system becomes "more concerned" and initiates a cycle, namely the second-level review. During this second-level review, the initial review intensity coefficient m is increased (e.g., if the initial review intensity coefficient m is 0.1, it increases in increments of 0.05).
[0097] In summary, according to the technical solution provided in the embodiments of this disclosure, after the first-level review determines that there is no violation, the second-level review is determined by 10 times the initial intensity coefficient, the review intensity coefficient is increased round by round by a preset incremental rule, key image frames are extracted by combining the number of image frames and the odd-even frame alternation rule, and the edge-checking feature library is called for cyclic review, thereby achieving precise progression of the second-level review intensity. This reduces the risk of missed review while balancing review efficiency and resource consumption, and performs secondary compliance verification of video content in a more rigorous and layered manner.
[0098] Figure 4 This is a flowchart illustrating another video review method provided in an embodiment of the present disclosure.
[0099] like Figure 4 As shown, if the second-level review detects the first non-compliant image frame, the first non-compliant image frame is diffused. The edge-scratching feature library is then used to perform a third-level review on the resulting set of non-compliant associated image frames. The specific steps include the following: Step 401: If the second-level review detects the first non-compliant image frame, extract all first non-compliant image frames corresponding to the second-level review based on the review results of the second-level review.
[0100] In one embodiment of this disclosure, the first non-compliant image frame refers to an image frame that is determined to be non-compliant during the secondary review process. After the secondary review is completed, all image frames determined to be non-compliant (i.e., the first non-compliant image frames) are extracted based on the review results.
[0101] Step 402: Based on the preset number of diffusions and the preset diffusion range, extend the diffusion forward and backward with all the first violation image frames as the center, and determine the set of violation-related image frames generated by each diffusion.
[0102] In one embodiment of this disclosure, the preset diffusion count refers to the number of times the operation of "extending outward from the first violation frame" is performed. For example, three diffusions mean extending forward and backward three times from the first violation image frame. The preset diffusion range refers to the number of image frames extended before and after the first violation image frame in each diffusion. For example, a range of 2 means extending forward 2 frames and backward 2 frames each time. The violation-related image frame set refers to the multiple sets of image frames that require further review formed by including all frames related to the first violation image frame in each diffusion operation (each diffusion generates a new set). Centered on all the first violation image frames that have been determined to be violations, the diffusion is extended to the image frames before and after these frames according to the preset diffusion count and diffusion range, thereby determining the set of image frames associated with the first violation frame generated after each diffusion.
[0103] For example, suppose the first violating image frame is the 10th frame, the preset number of diffusions is 2, and the preset diffusion range is 1 frame.
[0104] First diffusion: Centered on frame 10, extend forward by 1 frame (frame 9) and backward by 1 frame (frame 11). At this point, the set of associated frames is {9, 11}. Second diffusion: The set of associated frames is {8, 12}.
[0105] Through this diffusion, all possible related frames that may exist before and after the first illegal image frame can be found, further ensuring that no illegal content is missed.
[0106] The preset number of diffusion times can be determined based on the amount of violation (p) and the review intensity coefficient. The number of image frames n in the image frame set and the manual weighting coefficient q are determined. The violation quantity p is determined during a random check in a certain round of secondary review. Following the images, the number of non-compliant frames discovered in this round and the total number of frames sampled in this round. The ratio, where p is the value calculated for the first time. Specifically, the preset diffusion number = .
[0107] In practical applications, the determination of the preset diffusion times and manual weights can be based on the actual needs of the scenario, and no specific restrictions are imposed here.
[0108] Step 403: Call the edge-skimming feature library and perform a three-level review on each set of image frames associated with violations.
[0109] In one embodiment of this disclosure, texture features are extracted from image frames in each set of illegally associated image frames to obtain a texture feature set for each set. This texture feature set is then input into a border-skimming feature library for similarity matching. Based on the similarity matching result and a preset threshold, it is determined whether an image frame in each set of illegally associated image frames violates the rules. The specific process described above is the same as steps 202-204, and will not be repeated here.
[0110] In summary, based on the technical solutions provided in the embodiments of this disclosure, this disclosure, through the progressive logic of "violation frame location - related frame diffusion - edge-spot feature library matching," can not only accurately identify single-frame violations, but also systematically cover implicit violation scenarios related to time, significantly reducing the missed review rate. At the same time, through algorithmic feature comparison, it can achieve standardized and efficient identification of edge-spot content while reducing manual costs, ultimately achieving multi-dimensional optimization of review accuracy, efficiency, and cost, effectively solving pain points such as the difficulty in identifying edge-spot content in complex scenarios and the lack of uniformity in manual review standards.
[0111] Figure 5 This is a flowchart illustrating another video review method provided in an embodiment of the present disclosure.
[0112] like Figure 5 As shown, based on the results of the three-level review, determining whether to penalize videos containing borderline content involves the following steps: Step 501: Based on the audit results of the three-level audit, determine the final violation precision. The final violation precision is determined based on the number of violation frames found during the two-level audit, the number of violation frames found during the three-level audit, and the target audit frame corresponding to the first violation image frame found during the two-level audit.
[0113] In one embodiment of this disclosure, the number of violation frames identified during the Level 3 review is determined based on the review results. The final violation precision is calculated by combining the number of violation frames found during the Level 2 review, the number of violation frames found during the Level 3 review, and the target review frame number corresponding to the first violation image frame initially discovered during the Level 2 review. The final violation precision R is calculated using the following formula:
[0114] Where p is the total number of non-compliant frames, which is the sum of the number of non-compliant frames identified during the Level 3 review process and the number of non-compliant frames discovered during the Level 2 review process. It is the target review frame number corresponding to the first non-compliant image frame discovered during the second-level review process.
[0115] Step 502: Based on the final violation precision, determine whether to penalize the video containing borderline content.
[0116] In one embodiment of this disclosure, determining whether to penalize videos containing borderline content based on the final violation precision includes: If the final level of violation exceeds the first penalty threshold, the first-level penalty will be applied. If the final level of violation exceeds the second penalty threshold, the second-level penalty will be imposed on the video containing borderline content. If the final level of violation exceeds the third penalty threshold, the video containing borderline content will be subject to a third-level penalty.
[0117] Specifically, the first penalty threshold is s, the second penalty threshold is 2s, and the third penalty threshold is 3s, where s is the control coefficient, and preferably s=0.2.
[0118] When R>s: The first level of punishment is applied to videos containing borderline content, such as the video being censored.
[0119] When S>2s: Apply a second-level penalty to videos containing borderline content, such as prohibiting forwarding / sharing.
[0120] When S>3s: Third-level penalties will be imposed on videos containing borderline content, such as video deletion and potentially more severe functional restrictions on the account.
[0121] In summary, based on the technical solution provided in the embodiments of this disclosure, this disclosure quantifies the final precision of the violation through three-level review results, integrates dimensions such as the scale of the violation, the time of occurrence, and human weighting, and then implements a tiered penalty system of "masking - prohibiting forwarding - deletion + account restriction" based on the tiered penalty threshold. This achieves a precise match between the degree of violation and the severity of the penalty, which not only improves the objectivity and interpretability of identifying borderline content, but also effectively balances content norms and user experience through tiered governance, solving the pain points of inconsistent penalty standards and insufficient targeting of violation governance in traditional review.
[0122] Furthermore, the aforementioned video review methods also include: If the third-level review determines that there are no violations, the videos containing borderline content will undergo manual review.
[0123] Specifically, if the three-level review still cannot determine whether the borderline content video violates the rules, it will be transferred to manual review. However, the prerequisite for manual review is that the content has passed the gray-scale review (i.e., the first, second and third level reviews).
[0124] Furthermore, the above video review method also includes: continuously monitoring borderline content during the review process, conducting in-depth reviews (i.e., level 2 and level 3 reviews) when data anomalies are detected, re-importing the borderline feature library for matching, obtaining the review results, and conducting manual reviews if no match is found; The aforementioned data anomalies include controllable and monitorable content such as the number of reports, the spread of malicious users, abnormal downloads, and page views. User reports are collected, and when reports are received at different levels (based on the percentage of views), they can be imported into the manual review process for quick review according to the report type and time period.
[0125] It should be noted that the preliminary review (level 1 review), in-depth review (i.e., level 2 and level 3 review) and manual review are adjusted by increasing or decreasing the review intensity coefficient m, which ranges from 0 to 1. The larger the value, the more stringent the review.
[0126] This disclosure also provides a video review device. Figure 6 A structural block diagram of a video review device provided in this disclosure embodiment is shown below. Figure 6 As shown, the video review device 600 includes: an acquisition module 601, a first-level review module 602, a second-level review module 603, a third-level review module 604, and a penalty module 605.
[0127] In one exemplary embodiment, the acquisition module 601 is used to acquire a set of image frames of the video with borderline content, and to sequentially mark each image frame in the set with a frame number according to the playback sequence of the video with borderline content. In one exemplary embodiment, the first-level review module 602 is used to select a first target number of first image frames from the image frame set based on an initial review strength coefficient, and to perform first-level review on the first image frames in conjunction with a pre-built edge-trimming feature library. In one exemplary embodiment, the secondary review module 603 is used to determine the number of reviews for the secondary review based on the initial review intensity coefficient if the primary review determines that there is no violation. Based on the number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames of the target review frame number corresponding to the increased intensity coefficient. Combined with the edge-checking feature library, the key image frames of the target review frame number selected each time are reviewed for the secondary review. In one exemplary embodiment, the three-level review module 604 is used to diffuse the first illegal image frame if the second-level review detects it, and call the edge-trimming feature library to perform a three-level review on the illegal associated image frame set formed after diffusion; In one exemplary embodiment, the penalty module 605 is used to determine whether to penalize the video containing borderline content based on the judgment results of the three-level review.
[0128] In one exemplary embodiment, the initial review intensity coefficient is determined based on the base weight, the reporting rate, and the account's historical violation rate.
[0129] In one exemplary embodiment, the first-level review module 602 is further configured to: The number of first targets is determined based on the initial review intensity coefficient and the number of image frames in the image frame set. Image frames with an odd number of frames and whose number is equal to the number of first targets are selected from the image frame set and used as the first image frames. Extract texture features from the first image frames of the first target quantity to obtain the texture feature set corresponding to each first image frame; Based on the texture feature set corresponding to each first image frame and the feature data in the pre-built edge-scratching feature library, similarity matching is performed on each first image frame to obtain the similarity matching result corresponding to each first image frame; Based on the similarity matching results and preset thresholds corresponding to each first image frame, it is determined whether each first image frame violates the rules.
[0130] In one exemplary embodiment, the first-level review module 602 is further configured to: If the first-level review result indicates a violation, the video containing borderline content will be blocked, and the video will be penalized according to the preset initial penalty rules.
[0131] In one exemplary embodiment, the initial penalty rule includes: If the initial violation precision is greater than 0.2, light restrictions will be imposed on videos containing borderline content. If the initial violation precision is greater than 0.4, moderate restrictions will be imposed on videos containing borderline content. If the initial violation precision is greater than 0.6, videos containing borderline content will be severely punished; The initial violation precision is determined based on the number of violation frames and the number of first targets discovered during the first-level review process.
[0132] In one exemplary embodiment, the secondary review module 603 is further configured to: Step 1: If the first-level review result is no violation, determine the number of second-level reviews based on 10 times the initial review intensity coefficient; Step 2: Based on the preset incremental rules and the initial review intensity coefficient, generate the incremental review intensity coefficient for this round; Step 3: Based on the incremental review intensity coefficient of this round, the number of image frames in the image frame set, and the odd-even frame alternation rule, extract the key image frames corresponding to the target review frame number from the image frame set; Step 4: Call the edge-scratching feature library to perform secondary review on key image frames of the target review frame number; Step 5: Repeat steps 2 to 4 above until the number of audits determined in step 1 is completed.
[0133] In one exemplary embodiment, the three-level review module 604 is further configured to: If the second-level review detects the first non-compliant image frame, based on the review results of the second-level review, extract all first non-compliant image frames corresponding to the second-level review. Based on the preset number of diffusions and the preset diffusion range, the diffusion extends forward and backward with all the first violation image frames as the center, and determines the set of violation-related image frames generated by each diffusion. The edge-check feature library is invoked to perform a three-level review on each set of related image frames that violate the rules.
[0134] The above is subject to the attached document. Figure 1 -Appendix Figure 6This disclosure describes in detail the video review method provided in the embodiments of this disclosure. The method involves converting videos (i.e., videos with borderline content) into labeled images, then uniformly sampling a certain number of images at specific locations, extracting texture features, importing them into a borderline feature library, and then, if no violations are found, increasing the review intensity coefficient m and shifting the sampling range to images at position 2n. The already sampled images are then filtered out. This process is repeated 10m times. If violations are detected within 10m times, the increasing sampling targets are concentrated on the violation point, and then the review is expanded to both sides, advancing a certain number of reviews. When the final pass rate (i.e., initial violation precision) is greater than 0.2, different levels of functional restrictions and measures are then implemented. This allows for more accurate review of borderline videos, ensuring cloud storage security.
[0135] Figure 7 This is a hardware block diagram of an electronic device provided according to an embodiment of the present disclosure. The electronic device 700 according to an embodiment of the present disclosure includes at least a processor and a memory for storing computer-readable instructions. When the computer-readable instructions are loaded and executed by the processor, the processor performs the video review method described in any of the preceding embodiments of the present disclosure.
[0136] Figure 7 The illustrated electronic device 700 specifically includes a central processing unit (CPU) 701, a graphics processing unit (GPU) 702, and a memory 703. These units are interconnected via a bus 704. The CPU 701 and / or GPU 702 can function as the aforementioned processors, and the memory 703 can function as the aforementioned memory storing computer-readable instructions. Furthermore, the electronic device 700 may also include a communication unit 705, a storage unit 706, an output unit 707, an input unit 708, and an external device 709, all of which are also connected to the bus 704.
[0137] Figure 8 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of this disclosure. (As shown...) Figure 8 As shown, a computer-readable storage medium 800 according to an embodiment of the present disclosure stores computer-readable instructions 801 thereon. When the computer-readable instructions 801 are executed by a processor, the video review method described above with reference to any embodiment of the present disclosure is performed. The computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.
[0138] This disclosure further provides a computer program product, including a computer program that, when executed by a processor, implements the video review method described in any of the preceding embodiments of this disclosure.
[0139] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0140] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0141] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0142] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.
[0143] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0144] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.
[0145] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0146] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A video review method, characterized in that, include: Obtain a set of image frames from the video containing borderline content, and sequentially label each image frame in the set with a frame number according to the playback sequence of the video containing borderline content. Based on the initial review intensity coefficient, the first image frame of the first target number is selected from the image frame set, and the first image frame is reviewed in a first-level manner in combination with the pre-built edge-trimming feature library; If the first-level review determines that there is no violation, the number of second-level reviews is determined based on the initial review intensity coefficient. Based on the number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames of the target review frame number corresponding to the increased intensity coefficient. Combined with the edge-checking feature library, the key image frames of the target review frame number selected each time are sequentially reviewed for the second-level review. If the second-level review detects the first illegal image frame, the first illegal image frame is spread, and the edge-trimming feature library is called to perform a third-level review on the set of illegal associated image frames formed after the spread. Based on the results of the three-level review, it is determined whether to penalize the video containing borderline content.
2. The video review method according to claim 1, characterized in that, The initial review intensity coefficient is determined based on the basic weight, the reporting rate, and the account's historical violation rate.
3. The video review method according to claim 1, characterized in that, The step of selecting a first target number of first image frames from the image frame set based on the initial review intensity coefficient, and performing a first-level review on the first image frames in conjunction with a pre-built edge-trimming feature library, includes: The first target quantity is determined based on the initial review intensity coefficient and the number of image frames in the image frame set. Image frames with an odd number of frames and a number equal to the first target quantity are selected from the image frame set and used as the first image frames. Extract texture features from the first image frames of the first target number to obtain a texture feature set corresponding to each first image frame; Based on the texture feature set corresponding to each first image frame and the feature data in the pre-constructed edge-rubbing feature library, similarity matching is performed on each first image frame to obtain the similarity matching result corresponding to each first image frame; Based on the similarity matching results and preset thresholds corresponding to each first image frame, it is determined whether each first image frame violates the rules.
4. The video review method according to claim 1, characterized in that, The method further includes: If the first-level review result indicates a violation, the video containing borderline content will be blocked, and the video will be penalized according to the preset initial penalty rules.
5. The video review method according to claim 4, characterized in that, The initial penalty rules include: If the initial violation precision is greater than 0.2, the video containing borderline content will be subject to slight restrictions. If the initial violation precision is greater than 0.4, the video containing borderline content will be subject to moderate restrictions. If the initial violation precision is greater than 0.6, the video containing borderline content will be severely punished; The initial violation precision is determined based on the number of violation frames discovered during the first-level review process and the first target quantity.
6. The video review method according to claim 1, characterized in that, If the first-level review determines there is no violation, the number of second-level reviews is determined based on the initial review intensity coefficient. Based on this number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames corresponding to the target review frame number after the increase in intensity coefficient. Combining this with the edge-checking feature library, the key image frames selected for each target review frame number are sequentially subjected to second-level review, including: Step 1: If the first-level review result is no violation, determine the number of reviews for the second-level review by multiplying the initial review intensity coefficient by 10. Step 2: Based on the preset incremental rules and the initial review intensity coefficient, generate the current round of incremental review intensity coefficient; Step 3: Based on the incremental review intensity coefficient of this round, the number of image frames in the image frame set, and the odd-even frame alternation rule, extract the key image frames corresponding to the target review frame number from the image frame set; Step 4: Call the edge-rubbing feature library to perform secondary review on the key image frames of the target review frame number; Step 5: Repeat steps 2 to 4 above until the number of audits determined in step 1 is completed.
7. The video review method according to claim 6, characterized in that, If the second-level review detects a first non-compliant image frame, the first non-compliant image frame is diffused, and the edge-trimming feature library is used to perform a third-level review on the diffused set of non-compliant associated image frames, including: If the second-level review detects a first non-compliant image frame, based on the review result of the second-level review, extract all first non-compliant image frames corresponding to the second-level review; Based on a preset number of diffusions and a preset diffusion range, the diffusion extends forward and backward with all the first illegal image frames as the center, and determines the set of illegal associated image frames generated by each diffusion. The edge-check feature library is invoked to perform a three-level review on each set of related image frames that violate the rules.
8. The video review method according to claim 6, characterized in that, The determination of whether to penalize the borderline content video based on the judgment results of the three-level review includes: Based on the review results of the three-level review, the final violation precision is determined. The final violation precision is determined based on the number of violation frames found during the second-level review, the number of violation frames found during the third-level review, and the target review frame number corresponding to the first violation image frame found during the second-level review. Based on the final violation precision, it is determined whether to penalize the video containing borderline content.
9. The video review method according to claim 8, characterized in that, The process of determining whether to penalize the borderline content video based on the final violation precision includes: If the final violation precision exceeds the first penalty threshold, the first-level penalty shall be applied. If the final violation precision is greater than the second penalty threshold, the second level of penalty will be applied to the video containing borderline content. If the final violation precision is greater than the third penalty threshold, the third-level penalty will be applied to the video containing borderline content.
10. The video review method according to claim 1, characterized in that, The edge-scratching feature library is constructed based on the following method: Based on historical review data, obtain samples of borderline content, illegal video frames, and illegal images; Multi-dimensional visual feature extraction is performed on the content edge-scratching sample, the illegal video frame, and the illegal image to obtain a first edge-scratching feature set corresponding to the content edge-scratching sample, a second edge-scratching feature set corresponding to the illegal video frame, and a third edge-scratching feature set corresponding to the illegal image; The edge-scratching feature library is constructed based on the first edge-scratching feature set, the second edge-scratching feature set, and the third edge-scratching feature set.
11. The video review method according to claim 1, characterized in that, The method further includes: If the three-level review determines that there is no violation, the video containing borderline content will undergo manual review.
12. A video review device, characterized in that, include: The acquisition module is used to acquire a set of image frames of the video with borderline content, and to sequentially mark each image frame in the set with a frame number according to the playback sequence of the video with borderline content. The first-level review module is used to select a first target number of first image frames from the image frame set based on the initial review intensity coefficient, and to conduct a first-level review of the first image frames in combination with the pre-built edge-trimming feature library. The secondary review module is used to determine the number of secondary reviews based on the initial review intensity coefficient if the primary review determines that there is no violation. Based on the number of reviews, the initial review intensity coefficient is increased round by round to determine the key image frames of the target review frame number corresponding to the increased intensity coefficient. Combined with the edge-checking feature library, the key image frames of the target review frame number selected each time are sequentially reviewed in the secondary review. The three-level review module is used to spread the first illegal image frame if the second-level review detects the first illegal image frame, and call the edge-trimming feature library to perform a three-level review on the illegal associated image frame set formed after the spread; The penalty module is used to determine whether to penalize the borderline content video based on the judgment results of the three-level review.
13. The video review device according to claim 12, characterized in that, The initial review intensity coefficient is determined based on the basic weight, the reporting rate, and the account's historical violation rate.
14. The video review device according to claim 12, characterized in that, The primary review module is also used for: The first target quantity is determined based on the initial review intensity coefficient and the number of image frames in the image frame set. Image frames with an odd number of frames and a number equal to the first target quantity are selected from the image frame set and used as the first image frames. Extract texture features from the first image frames of the first target number to obtain a texture feature set corresponding to each first image frame; Based on the texture feature set corresponding to each first image frame and the feature data in the pre-constructed edge-rubbing feature library, similarity matching is performed on each first image frame to obtain the similarity matching result corresponding to each first image frame; Based on the similarity matching results and preset thresholds corresponding to each first image frame, it is determined whether each first image frame violates the rules.
15. The video review device according to claim 12, characterized in that, The primary review module is also used for: If the first-level review result indicates a violation, the video containing borderline content will be blocked, and the video will be penalized according to the preset initial penalty rules.
16. The video review device according to claim 15, characterized in that, The initial penalty rules include: If the initial violation precision is greater than 0.2, the video containing borderline content will be subject to slight restrictions. If the initial violation precision is greater than 0.4, the video containing borderline content will be subject to moderate restrictions. If the initial violation precision is greater than 0.6, the video containing borderline content will be severely punished; The initial violation precision is determined based on the number of violation frames discovered during the first-level review process and the first target quantity.
17. The video review device according to claim 12, characterized in that, The secondary review module is also used for: Step 1: If the first-level review result is no violation, determine the number of reviews for the second-level review by multiplying the initial review intensity coefficient by 10. Step 2: Based on the preset incremental rules and the initial review intensity coefficient, generate the current round of incremental review intensity coefficient; Step 3: Based on the incremental review intensity coefficient of this round, the number of image frames in the image frame set, and the odd-even frame alternation rule, extract the key image frames corresponding to the target review frame number from the image frame set; Step 4: Call the edge-rubbing feature library to perform secondary review on the key image frames of the target review frame number; Step 5: Repeat steps 2 to 4 above until the number of audits determined in step 1 is completed.
18. The video review device according to claim 17, characterized in that, The three-level review module is also used for: If the second-level review detects a first non-compliant image frame, based on the review result of the second-level review, extract all first non-compliant image frames corresponding to the second-level review; Based on a preset number of diffusions and a preset diffusion range, the diffusion extends forward and backward with all the first illegal image frames as the center, and determines the set of illegal associated image frames generated by each diffusion. The edge-check feature library is invoked to perform a three-level review on each set of related image frames that violate the rules.
19. An electronic device, characterized in that, include: Memory, used to store computer-readable instructions; as well as A processor for executing the computer-readable instructions, causing the electronic device to perform the video review method as described in any one of claims 1 to 11.
20. A non-transitory computer-readable storage medium for storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by a processor, the processor performs the video review method as described in any one of claims 1 to 11.
21. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the video review method as described in any one of claims 1 to 11.