A video defect detection method, a detection model training method and a device
By segmenting the video frame into image blocks and using training models to detect defects, the problem of low local defect detection efficiency in the prior art is solved, high-precision defect recognition and positioning is achieved, and a variety of business needs are supported.
Patent Information
- Application Number
- CN202410742356.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-06-07
AI Technical Summary
The prior art is difficult to efficiently detect and locate local defects in videos, such as color levels, color blocks, and color spots, during the capture, compression, transmission and playback of video content, resulting in inefficient video quality evaluation.
By dividing each video frame of the video to be detected into multiple image blocks and using the trained detection model for object detection, the defect area and defect type in each image block are determined. The detection model is based on multiple image block training and can identify multiple types of defective regions in different image blocks.
It realizes fast and high-precision detection and positioning defect areas in the video, can accurately identify multiple defect types, improves the detection capabilities of local defects, and supports media delivery, ultra-high-definition video film selection, old film repair and other services.
Smart Images

Figure CN118485659B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and in particular, to a video defect detection method, a detection model training method, and an apparatus therefor. Background Art
[0002] As an important part of entertainment culture, the video quality of film and television works directly affects the viewing experience of audiences. However, in multiple links such as video content capture, compression, transmission, and playback, various defects are easily introduced. For example, local defects such as color levels, color blocks, and color spots.
[0003] In scenarios such as media delivery, ultra-high-definition video selection, and old film restoration, it mainly relies on manual viewing of the entire film to discover defects in the video. This method is inefficient, error-prone, and difficult to meet the requirements for high-quality and high-efficiency production. Summary of the Invention
[0004] In view of this, the present disclosure provides a video defect detection method, a detection model training method, an apparatus, an electronic device, and a storage medium.
[0005] According to one aspect of the present disclosure, there is provided a video defect detection method, the method comprising:
[0006] Dividing each video frame in the video to be detected into at least one image block;
[0007] For each of the at least one image block, performing object detection through a trained detection model to determine a defect region in each image block and a defect type corresponding to the defect region; wherein the detection model is trained based on a plurality of image blocks, and the defect types corresponding to the defect regions in each of the plurality of image blocks are not all the same;
[0008] Determining a defect region in each video frame and a defect type corresponding to the defect region according to the defect regions in the at least one image block and the defect types corresponding to the defect regions;
[0009] Determining a video segment with defects in the video to be detected, a defect region in the video segment, and a defect type corresponding to the defect region according to the defect regions in each video frame and the defect types corresponding to the defect regions.
[0010] In a possible implementation manner, the method further comprises:
[0011] Determining a defect score corresponding to each video frame according to the defect regions in each video frame and the defect types corresponding to the defect regions, the defect score being used to represent the degree of influence of the defect region on the visual effect;
[0012] Determine the defect score corresponding to the video to be detected according to the defect scores corresponding to each video frame.
[0013] In a possible implementation manner, for each of the at least one image block, performing object detection through a trained detection model to determine the defect region and the defect type corresponding to the defect region in each image block, including:
[0014] For any image block, generate a high-frequency sub-image and a low-frequency sub-image corresponding to the image block; wherein, the high-frequency sub-image is used to characterize the texture information of the image block, and the low-frequency sub-image is used to characterize the structural information of the image block;
[0015] After splicing and processing the high-frequency sub-image and the low-frequency sub-image corresponding to the image block, input them into the detection model to obtain the defect region and the defect type corresponding to the defect region in the image block.
[0016] In a possible implementation manner, the method further includes:
[0017] Generate a defect degree map for each video frame according to the defect regions in each of the at least one image block; the defect degree map is used to characterize the defect degrees corresponding to different image blocks;
[0018] Perform fusion processing on the defect degree map and the low-frequency sub-image corresponding to each image block to generate a mask for the defect region in each video frame; the mask is used to characterize the defect region.
[0019] In a possible implementation manner, the defect type includes at least one of color level, color block, and color spot.
[0020] According to another aspect of the present disclosure, a method for training a detection model is provided, and the method includes:
[0021] Construct a training sample set; wherein, the training sample set includes a plurality of image blocks and the label of each image block in the plurality of image blocks, and the label of each image block includes a label indicating whether there is a defect region in each image block, a label of the defect region in each image block, and a label of the defect type corresponding to the defect region; the defect types corresponding to the defect regions in the various image blocks in the plurality of image blocks are not completely the same;
[0022] Use the training sample set to train a preset model to obtain a trained detection model, and the detection model is used to detect the defect region and the defect type corresponding to the defect region in the image blocks obtained by segmenting the video.
[0023] In a possible implementation manner, the constructing the training sample set includes:
[0024] Obtain the original dataset; the original dataset includes defective videos and / or non-defective videos;
[0025] Screen out candidate videos that meet the preset diversity requirements from the original dataset;
[0026] Extract sample images from the candidate videos, and obtain the defective regions and the corresponding defect types in the sample images;
[0027] Segment the sample images into the multiple image patches;
[0028] Determine the label of each image patch in the multiple image patches according to the defective regions and the corresponding defect types in the sample images.
[0029] In a possible implementation, the extracting sample images from the candidate videos includes:
[0030] When the candidate video is a non-defective video, perform encoding processing on the candidate video so that there are defects in the encoded video;
[0031] Extract the sample images from the encoded video.
[0032] In a possible implementation, the training the preset model with the training sample set to obtain a trained detection model includes:
[0033] For any one of the multiple image patches, generate a high-frequency sub-image and a low-frequency sub-image corresponding to the image patch; wherein, the high-frequency sub-image is used to represent the texture information of the image patch, and the low-frequency sub-image is used to represent the structural information of the image patch;
[0034] Input the high-frequency sub-image and the low-frequency sub-image corresponding to the image patch into the preset model to obtain a defect prediction result of the image patch, and the defect prediction result includes the defective region and the corresponding defect type in the image patch;
[0035] Adjust the parameters in the preset model based on the defect prediction result of the image patch and the label of the image patch to obtain the detection model.
[0036] In a possible implementation, the defect type includes at least one of color level, color block, and color spot.
[0037] According to another aspect of the present disclosure, there is provided a video defect detection device, and the device includes:
[0038] A segmentation module, configured to segment each video frame in the video to be detected into at least one image patch;
[0039] An image block defect detection module, which is used to perform object detection on each of the at least one image block through a trained detection model to determine the defect region and the corresponding defect type in each image block; wherein, the detection model is trained based on multiple image blocks, and the defect types corresponding to the defect regions in each of the multiple image blocks are not completely the same;
[0040] A video frame defect detection module, which is used to determine the defect region and the corresponding defect type in each video frame according to the defect region and the corresponding defect type in the at least one image block;
[0041] A video defect detection module, which is used to determine the defective video segment in the video to be detected, the defect region in the video segment, and the corresponding defect type according to the defect region and the corresponding defect type in each video frame.
[0042] According to another aspect of the present disclosure, there is provided a training device for a detection model, the device includes:
[0043] A sample construction module, which is used to construct a training sample set; wherein, the training sample set includes multiple image blocks and the label of each image block in the multiple image blocks, and the label of each image block includes the label of whether there is a defect region in each image block, the label of the defect region in each image block, and the label of the defect type corresponding to the defect region; the defect types corresponding to the defect regions in each of the multiple image blocks are not completely the same;
[0044] A training module, which is used to train a preset model by using the training sample set to obtain a trained detection model, and the detection model is used to detect the defect region and the corresponding defect type in the image blocks obtained by segmenting the video.
[0045] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above method when executing the instructions stored in the memory.
[0046] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein, the computer program instructions implement the above method when executed by a processor.
[0047] According to another aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, when the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above method.
[0048] Through the present disclosure, each video frame in a video to be detected is segmented into at least one image block; for each of the at least one image block, object detection is performed through a trained detection model to determine a defect region in each image block and a defect type corresponding to the defect region; wherein, the detection model is trained based on a plurality of image blocks, and the defect types corresponding to the defect regions in each of the plurality of image blocks are not completely the same; according to the defect regions and the defect types corresponding to the defect regions in the at least one image block, the defect regions and the defect types corresponding to the defect regions in each video frame are determined; according to the defect regions and the defect types corresponding to the defect regions in each video frame, a video segment with defects in the video to be detected, the defect regions in the video segment, and the defect types corresponding to the defect regions are determined. In this way, through the trained detection model, the defect regions in the video to be detected can be quickly and highly accurately detected and located. Since the detection model is trained based on a plurality of image blocks with defect types corresponding to the defect regions in each image block not being completely the same, the defect types can be accurately identified, enabling fast and precise detection for various types of defects. At the same time, by using the image blocks obtained by segmenting the video frame to analyze and process the local part of the video frame, local defects such as color gradation, color patches, and color spots can be accurately detected, improving the detection ability for local defects and effectively supporting services such as media delivery, selection of ultra-high-definition videos, and restoration of old films. As an example, detailed defect region visualization information can also be generated. For example, a video segment with defects in the video to be detected and a mask of the defect region that is highly consistent with the subjective perception of the human eye are generated, further improving the efficiency and accuracy of subsequent defect correction.
[0049] Other features and aspects of the present disclosure will become clear from the following detailed description of the exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification and are used to explain the principles of the present disclosure.
[0051] Figure 1 FIG. shows a flowchart of a video defect detection method according to an embodiment of the present disclosure;
[0052] Figure 2 FIG. shows a flowchart of a video defect detection method according to an embodiment of the present disclosure;
[0053] Figure 3 FIG. shows a schematic flowchart of a video defect detection method according to an embodiment of the present disclosure;
[0054] Figure 4The flowchart of a method for training a detection model according to an embodiment of the present disclosure is shown;
[0055] Figure 5 The flowchart of a method for constructing a training sample set according to an embodiment of the present disclosure is shown;
[0056] Figure 6 The schematic flowchart of a method for video defect detection according to an embodiment of the present disclosure is shown;
[0057] Figure 7 The schematic structural diagram of a video defect detection device according to an embodiment of the present disclosure is shown;
[0058] Figure 8 The schematic structural diagram of a training device for a detection model according to an embodiment of the present disclosure is shown;
[0059] Figure 9 The block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. Detailed implementation manners
[0060] The various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specifically noted.
[0061] The reference to "an embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, the statements "exemplary", "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0062] In this application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can mean: including the case where A exists alone, where A and B exist simultaneously, and where B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or a similar expression refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0063] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0064] In scenarios such as media delivery, selection of ultra-high-definition videos, and restoration of old films, it relies on manual viewing of the entire film to detect defects in the video. This method is inefficient, error-prone, and difficult to meet the requirements for high-quality and high-efficiency production. Among them, the media delivery scenario means that the video platform party needs to conduct quality inspections on the video media provided by the production party to determine whether the video media meets the corresponding online standards to ensure the quality of the film source. In the ultra-high-definition video selection scenario, it is necessary to evaluate the videos uploaded by users and screen out high-quality videos to improve the viewing experience of users. In the old film restoration scenario, it is necessary to conduct quality inspections on the restored video media to determine whether the video media meets the restoration standards.
[0065] In the related art, video defect detection algorithms can be used to automatically detect defects in videos. However, traditional video defect detection algorithms can only detect a single type of defect, but videos in actual business scenarios often have multiple types of defects, among which some defects are often irregularly distributed in local areas of video frames, such as color scale, color blocks, color spots, etc. Among them, color scale (Banding): refers to the abrupt color bands that appear in smooth and gradual color transitions; due to many reasons such as poor lighting conditions during shooting, incorrect camera settings, or excessive compression during post-processing, color scale problems are usually manifested as "color bands" (visible jumps or stripes that interrupt the smooth transition of colors), causing the image to look low-quality or unnatural. Color block (Blocking): refers to the block structure in the image; caused by compression distortion, encoding errors or damaged data, it usually appears as unnatural large monochrome areas in the image. These monochrome areas lack details and form a sharp contrast with the surrounding image content, destroying visual continuity and realism. Quantization: Due to sensor noise, poor lighting conditions or errors in the image processing process, the image appears "blocky", jagged or unsharp, with randomly distributed color dots or spots. Therefore, traditional video defect detection algorithms cannot meet the needs of fast and accurate detection of various types of defects.
[0066] In order to solve the above technical problems, the present disclosure proposes a video defect detection method (see below for detailed description), which can quickly and accurately detect defective areas in the video, and can accurately identify the defect types, thereby realizing fast and accurate detection of various types of defects. As an example, local defects such as color levels, color blocks, and color spots can be accurately detected, which improves the detection capability of local defects and effectively supports media delivery, ultra-high-definition video selection, old film restoration and other businesses.
[0067] The video defect detection method provided by the present invention is described in detail below.
[0068] Figure 1 A flowchart of a video defect detection method according to an embodiment of the present disclosure is shown. By way of example, the method can be applied to various types of electronic devices such as computers, servers, mobile terminals, wearable devices, etc., without limitation. Figure 1 As shown, the method may include the following steps:
[0069] Step 101: divide each video frame in the video to be detected into at least one image block.
[0070] Exemplarily, the video to be detected may include one or more video frames; it can be understood that when the video to be detected includes one video frame, the video to be detected is a single image.
[0071] In this step, each video frame in the video to be detected is divided into at least one image block for local analysis of the video frame, so as to accurately detect local defects in the video frame. Exemplarily, each video frame in the video to be detected can be divided into a preset number of image blocks, and the value of the preset number can be set according to requirements. Among them, the sizes of different image blocks can be the same or different. For example, each video frame can be divided into a preset number of image blocks with a size of 256x256 (that is, the size of each image block is 256 pixels x 256 pixels).
[0072] Exemplarily, there can be overlap or no overlap between different image blocks. For example, a sliding window with a preset scale can be slid on the video frame at a fixed step size, so as to divide the video frame into one image block after another. If the step size is less than the size of the image block, there is overlap between adjacent image blocks divided out. If the step size is not less than the size of the image block, there is no overlap between adjacent image blocks divided out.
[0073] In a possible implementation, each video frame in the video to be detected can be preprocessed, and the preprocessed video frame is divided into at least one image block. Exemplarily, the spatial resolution and frame rate of each video frame in the video to be detected can be adjusted to a preset size. For example, the spatial resolution can be adjusted to 1920x1080 (that is, the size of each video frame is 1920 pixels x 1080 pixels); the frame rate is adjusted to 25fps; the pixel format of each video frame in the video to be detected can also be adjusted to a preset format.
[0074] Step 102: For each of the at least one image block, perform object detection through a trained detection model to determine the defect area and the corresponding defect type in each image block; wherein, the detection model is trained based on a plurality of image blocks, and the defect types corresponding to the defect areas in each of the plurality of image blocks are not completely the same.
[0075] Among them, the detection model is used to detect the defective area in the image patch and can automatically identify the corresponding defect type of the image patch. The specific training process of the detection model is described below. Exemplarily, the detection model can be any object detection model. For example, it can be the YOLOv6 object detection model. Among them, the backbone network of the YOLOv6 object detection model is the lightweight network EfficientRep, and the neck network is the Rep-PAN structure. Any image patch is input into the trained YOLOv6 object detection model. Among them, the feature maps output by the network blocks ERBlock3, ERBlock4, and ERBlock5 in the lightweight network EfficientRep are respectively input into the neck network RepBi-PAN structure. The RepBi-PAN structure extracts and combines the effective information for subsequent object localization and classification, so as to accurately detect the defective area and the corresponding defect type in the image patch.
[0076] It can be understood that for any image patch, there may or may not be a defective area in the image patch; if there is no defective area in the image patch, the defective area and the corresponding defect type determined by the detection model for object detection are empty.
[0077] Among them, the defective area represents the location and range of the defect in the image; for an image patch with a defect, the number of defective areas in each image patch can be one or more, that is, there may be one defective area or multiple defective areas in the same image patch. Exemplarily, if there are multiple defective areas in an image patch, through the trained detection model for object detection, the defective areas and the corresponding defect types in the image patch can be determined.
[0078] Exemplarily, the defect type includes at least one of color level, color block, and color spot.
[0079] Among them, when there is one defective area in the same image patch, there is one type of defect in the image patch, and its type is the defect type corresponding to the defective area; when there are multiple defective areas in the same image patch, the defect types corresponding to different defective areas can be the same or different. If the defect types corresponding to different defective areas are the same, it means that there is one type of defect in the image patch. If the defect types corresponding to different defective areas are not completely the same, it means that there are multiple types of defects in the image patch.
[0080] In a possible implementation, for any image block, a high-frequency sub-image (HFM, High Frequency Modulation) and a low-frequency sub-image (LFM, Low Frequency Modulation) corresponding to the image block are generated; wherein, the high-frequency sub-image is used to characterize the texture information of the image block, and the low-frequency sub-image is used to characterize the structural information of the image block; after splicing the high-frequency sub-image and the low-frequency sub-image corresponding to the image block, the spliced result is input into the detection model to obtain the defect region in the image block and the defect type corresponding to the defect region. For example, after splicing the high-frequency sub-image and the low-frequency sub-image corresponding to the image block, the obtained spliced image can be input into the YOLOv6 object detection model, and the YOLOv6 object detection model outputs the defect region in the image block and the defect type corresponding to the defect region.
[0081] Exemplarily, the Sobel operator can be used to extract high-frequency detail information from the image block to generate a high-frequency sub-image, which can also be called a block-level high-frequency map; specifically, the Sobel operator can be used to perform edge detection on the image block, and then perform high-pass filtering or morphological operations (such as dilation, erosion, etc.) on the image block after edge detection to extract the high-frequency information in the image block. Finally, the average gradient, variance, etc. of the high-frequency information in the image block can be calculated, and the calculation results can be used as pixel values to generate the high-frequency sub-image.
[0082] Exemplarily, a piecewise smoothing algorithm can be applied to process the image block to obtain a low-frequency sub-image representing the overall structure and the change of the smooth region, which can also be called a block-level low-frequency map; specifically, smoothing algorithms such as mean filtering, Gaussian filtering, and median filtering can be applied within each image block to smooth the image block by reducing the difference between pixel values, eliminating high-frequency noise and details within the image block, and retaining the low-frequency overall structural information, thereby generating the low-frequency sub-image.
[0083] Step 103: Determine the defect region in each video frame and the defect type corresponding to the defect region according to the defect region in the at least one image block and the defect type corresponding to the defect region.
[0084] Among them, the number of defective regions in each video frame can be one or more, that is, there may be one defective region in the same video frame, or there may be multiple defective regions. The defective types corresponding to different defective regions can be the same or different. When there is one defective region in the same video frame, there is one type of defect in this video frame, and its type is the defective type corresponding to this defective region; while when there are multiple defective regions in the same video frame, the defective types corresponding to different defective regions can be the same or different. If the defective types corresponding to different defective regions are the same, it indicates that there is one type of defect in this video frame. If the defective types corresponding to different defective regions are not completely the same, it indicates that there are multiple types of defects in this video frame.
[0085] In a possible implementation manner, for any video frame, the defective regions corresponding to the same defective type in each image block in this video frame can be fused, so as to determine the defective regions in this video frame and the defective types corresponding to the defective regions. As an example, for the defective regions corresponding to the same defective type in each image block, the defective regions of adjacent or overlapping image blocks in these image blocks can be merged. For example, the defective regions of adjacent or overlapping image blocks can be merged through a feature fusion and adaptive weighted average algorithm. The merged defective region is a defective region in this video frame, and the defective type corresponding to the defective region before merging is the defective type corresponding to the merged defective region in this video frame; for example, the defective regions with a color level defect in all image blocks in this video frame can be fused, and the fused defective region is a defective region in this video frame, and the defective type corresponding to this defective region is the color level; in this way, by traversing each defective type and performing a fusion process on the defective regions corresponding to each defective type in each image block in this video frame, accurate defective regions in the video frame and the defective types corresponding to each defective region can be obtained; exemplarily, non-defective regions in the video frame can also be obtained.
[0086] In a possible implementation manner, the method further includes: generating a defect degree map for each video frame according to the defective regions in each image block of the at least one image block; the defect degree map is used to characterize the defect degree corresponding to different image blocks; performing a fusion process on the defect degree map and the low-frequency sub-map corresponding to each image block to generate a mask of the defective regions in each video frame; the mask is used to characterize the defective regions.
[0087] Exemplarily, the number of defect degree maps for each video frame can be multiple, where each defect degree map corresponds to one defective type; the size of each defect degree map is the same as the size of the video frame.
[0088] As an example, when the detection model performs object detection on any image patch, it can detect each defective region in the image patch and output the probability values corresponding to each defective region for each preset defect type (such as color level, color block, color spot). For any defective region, the preset defect type corresponding to the maximum probability value is the defect type corresponding to the defective region. Exemplarily, the defect degree map is a grayscale value. For any preset defect type, the probability value that the defect type corresponding to the defective region in the image patch with defects in the video frame is this preset defect type can be used as the grayscale value of each pixel point in the image patch, that is, the grayscale values of each pixel point in the same image patch are the same, and the grayscale value of each pixel point in the image patch without defects is 0, so that the defect degree map of the video frame can be generated. The magnitude of the grayscale value in the defect degree map represents the degree of the existence of this preset defect in each image patch of the video frame. For example, an image patch contains defective region A with a corresponding defect type of X and a probability of a%, and defective region B with a corresponding defect type of Y and a probability of b%. Then, for the defect degree map corresponding to the defect type X, the pixel value in this image patch is set to a%, and for the defect degree map corresponding to the defect type Y, the pixel value in this image patch is set to b%.
[0089] In this way, by traversing each preset defect type, the defect degree maps corresponding to each preset defect type for each video frame can be generated. Furthermore, for any preset defect type, the defect degree map corresponding to this preset defect type in the video frame and the low-frequency sub-map corresponding to each image patch in the video frame can be fused to generate a mask for the defective region corresponding to this preset defect type in the video frame. Exemplarily, during the fusion process, the product of the pixel value of any pixel point in the low-frequency sub-map corresponding to the image patch and the grayscale value of this pixel point in the defect degree map can be calculated as the new pixel value of this pixel point in the mask. Since the image patches obtained by segmenting the video frame can represent the local information in the video frame, by fusing the defect degree map and the low-frequency sub-map corresponding to each image patch, a mask for the defective region that is highly consistent with the subjective perception of the human eye is generated, that is, the defective region in the video frame represented by the mask is consistent with the defective region determined by the human eye through subjective viewing. In this way, through the visual information of the high-precision mask of the defective region, it is convenient for the user to view the defective regions in the video frame, better guiding subsequent work such as image defect repair, and greatly improving the efficiency and accuracy of subsequent correction of defects in the video frame.
[0090] Step 104: Determine the video segments with defects, the defective regions in the video segments, and the defect types corresponding to the defective regions in the video to be detected according to the defective regions in each video frame and the defect types corresponding to the defective regions.
[0091] Exemplarily, the number of video segments can be one or more, that is, there may be multiple defective video segments in the video to be detected. For example, consecutive video frames detected with the same type of defect can be regarded as a defective video segment, and the number of video frames included in the consecutive video frames can be set according to requirements, which is not limited herein. Among them, the defective area corresponding to each video segment can be one or more, and the defect types corresponding to different defective areas can be the same or different, that is, one type of defect or multiple types of defects existing in each video segment can be detected. For example, if multiple consecutive video frames all have a color level defect, then these multiple video frames are regarded as a video segment with a color level defect. For another example, if multiple consecutive video frames all have a color level defect and a color block defect, then these multiple video frames are regarded as a video segment with a color level defect and a color block defect. For yet another example, if multiple consecutive video frames all have a color level defect and a color block defect, and some of the video frames also have a color spot defect, then these multiple video frames are regarded as a video segment with a color level defect and a color block defect.
[0092] As an example, when the video to be detected includes one video frame, the detection and recognition of one or more types of local defects existing in the video frame can be achieved. As another example, when the video to be detected includes multiple video frames, the defective video segments existing in the video to be detected can be determined by combining the defective areas and defect types in these multiple video frames, and the defective areas existing in the video segment and the defect types corresponding to the defective areas can be determined. In this way, by determining the visual information of the defective video segments existing in the video to be detected, it is convenient for the user to view the defective video segments in the video to be detected, greatly improving the efficiency and accuracy of subsequent correction of the defects in the video to be detected.
[0093] In the embodiments of the present disclosure, each video frame in the video to be detected is segmented into at least one image block; for each of the at least one image block, object detection is performed through a trained detection model to determine the defective area and the corresponding defect type in each image block; wherein, the detection model is trained based on a plurality of image blocks, and the defect types corresponding to the defective areas in each of the plurality of image blocks are not completely the same; according to the defective areas and the corresponding defect types in the at least one image block, the defective areas and the corresponding defect types in each video frame are determined; according to the defective areas and the corresponding defect types in each video frame, the video segments with defects in the video to be detected, the defective areas in the video segments and the corresponding defect types are determined. In this way, through the trained detection model, the defective areas in the video to be detected can be quickly and accurately detected and located. Since the detection model is trained based on a plurality of image blocks with defect types corresponding to the defective areas in each image block not being completely the same, the defect types can be accurately identified, realizing fast and accurate detection for various types of defects. At the same time, by using the image blocks obtained by segmenting the video frames to analyze and process the local parts of the video frames, local defects such as color gradation, color patches, and color spots can be accurately detected, improving the detection ability of local defects and effectively supporting services such as media delivery, selection of ultra-high-definition videos, and restoration of old videos. As an example, detailed visual information of the defective areas can also be generated. For example, the video segments with defects in the video to be detected and the masks of the defective areas that are highly consistent with the subjective perception of the human eye are further improved the efficiency and accuracy of subsequent defect correction.
[0094] Further, in order to accurately analyze and evaluate the impact of the defects in the video to be detected on the video quality, the severity of the defects in the video to be detected can be accurately evaluated on the basis of detecting the defective areas and the corresponding defect types in the video to be detected.
[0095] Figure 2 FIG. shows a flowchart of a video defect detection method according to an embodiment of the present disclosure. Exemplarily, the method can be applied to various types of electronic devices, such as Figure 2 As shown, the method may include the following steps:
[0096] Step 201, segment each video frame in the video to be detected into at least one image block;
[0097] Step 202, for each of the at least one image block, perform object detection through a trained detection model to determine the defective area and the corresponding defect type in each image block; wherein, the detection model is trained based on a plurality of image blocks, and the defect types corresponding to the defective areas in each of the plurality of image blocks are not completely the same;
[0098] Step 203: Determine the defect regions and the corresponding defect types in each video frame according to the defect regions and the corresponding defect types in the at least one image block.
[0099] Step 204: Determine the video segments with defects, the defect regions in the video segments, and the corresponding defect types in the video to be detected according to the defect regions and the corresponding defect types in each video frame.
[0100] Among them, the above steps 201-204 are the same as steps 101-104 in the above Figure 1 and will not be elaborated here.
[0101] Step 205: Determine the defect score corresponding to each video frame according to the defect regions and the corresponding defect types in each video frame, where the defect score is used to represent the degree of influence of the defect region on the visual effect.
[0102] Exemplarily, for any video frame, the defect score corresponding to the video frame can be determined by analyzing the distribution of the defect regions in the video frame and the number of the corresponding defect types; among them, the distribution can include the size and the position. For example, the defect score can be positively correlated with the size of the defect region in the video frame, that is, the larger the defect region in the video frame, the higher the corresponding defect score, and vice versa, the lower the corresponding defect score; for another example, the defect score can be positively correlated with the position of the defect region in the video frame relative to the center of the video frame, that is, the closer the defect region in the video frame is to the center of the video frame, the higher the corresponding defect score, and vice versa, the lower the corresponding defect score. For another example, the defect score can be positively correlated with the number of the defect types corresponding to the defect regions in the video frame, that is, the more the defect types corresponding to all the defect regions in the video frame, the more the defect types existing in the video frame, and the higher the corresponding defect score, and vice versa, the lower the corresponding defect score. In this way, by determining the defect score corresponding to the video frame, the severity of the defects in a single-frame image is quantitatively evaluated. The specific calculation method of the defect score can be selected according to needs and is not limited here.
[0103] Exemplarily, the severity level of the defects in the video frame can be determined according to the relative size of the defect score and a preset threshold, where the number and the specific values of the preset thresholds can be set according to requirements; as an example, three preset thresholds can be set, and then the video frames with detected defects are divided into four grades: not obvious, slight, severe, and extremely severe based on the relative size of the defect score and these three preset thresholds.
[0104] Step 206: Determine the defect score corresponding to the video to be detected according to the defect score corresponding to each video frame.
[0105] Exemplarily, a weight assignment mechanism can be adopted to determine the weight corresponding to each video frame in the video to be detected, and the defect scores corresponding to each video frame in the video to be detected are weighted and summed to generate the defect score of the entire video to be detected.
[0106] As an example, various factors such as the defect types corresponding to the defect regions in each video frame, the visual impact of the defects in each video frame, and the occurrence frequency of each defect type in the video to be detected can be considered to integrate the defect detection results of each video frame in the video to be detected. Among them, the defect detection result of each video frame includes the defect region in each video frame and the defect type corresponding to the defect region, so as to calculate the defect score corresponding to the video to be detected by combining the detection results of each video frame in the video dimension; Exemplarily, the weights of each video frame can be determined based on the preset weights of different defect types and the preset weights of the occurrence frequencies of different defect types in the video to be detected. For example, for a video frame, if the preset weight of the defect type corresponding to the defect region in this video frame is higher and the frequency of the defect region of this defect type appearing in all video frames of the video to be detected is higher, then the weight of this video frame is higher; on the contrary, the weight is lower. The specific calculation method of the weight can be selected according to needs and is not limited here. The visual impact of the defects in each video frame is represented by the defect score corresponding to the video to be detected, so as to perform weighted summation on the defect scores corresponding to each video frame based on the weights of each video frame to obtain the defect score of the video to be detected. In this way, the defect detection results such as the defect region in each video frame of the video to be detected and the defect type corresponding to the defect region are integrated, and the defect score corresponding to the video to be detected is calculated through a scientific and reasonable weight assignment mechanism, and the severity of the defects in the whole video is quantitatively evaluated.
[0107] Exemplarily, a continuous plurality of video frames detected with defects can be used as a defective video segment, and the defect score of this video segment is determined based on the defect scores of these multiple video frames. If there is a defective video segment in the video to be detected, the defect score of this video segment is the defect score corresponding to this video to be detected; if there are multiple defective video segments in the video to be detected, the defect scores of these multiple video segments are weighted and summed as the defect score corresponding to this video to be detected.
[0108] Exemplarily, the severity level of the defects in the video to be detected can be determined according to the relative magnitude between the defect score and a preset threshold, where the number and specific values of the preset thresholds can be set according to requirements; As an example, three preset thresholds can be set, and then the severity of the defects in the video to be detected is divided into four levels: not obvious, minor, severe, and extremely severe based on the relative magnitude between the defect score of the video to be detected and these three preset thresholds, which can provide a strong quantitative reference for subsequent optimization of video quality and improvement of user viewing experience.
[0109] In an embodiment of the present disclosure, according to the defect regions in each video frame and the corresponding defect types of the defect regions, the defect scores corresponding to each video frame are determined, and according to the defect scores corresponding to each video frame, the defect score corresponding to the video to be detected is determined. In this way, the severity of defects in a single-frame image is quantitatively evaluated through the defect scores corresponding to the video frames; and the severity of defects in the entire video is quantitatively evaluated through the defect score corresponding to the video to be detected; thus comprehensively covering the quantitative evaluation of the severity of defects in both single frames of images and the entire video, providing effective reference information for manual judgment. For example, it can provide highly accurate data support for subsequent defect repair work, ensuring the pertinence and effectiveness of repair measures.
[0110] For example, Figure 3 shows a schematic flowchart of a video defect detection method according to an embodiment of the present disclosure, as Figure 3As shown, it may include the following stages. First, the preprocessing stage. In this stage, for each video frame in the video frame to be detected, the video frame is segmented into multiple image blocks. Then, for each image block, a high-frequency subgraph (HFM) and a low-frequency subgraph (LFM) corresponding to the image block are generated, and the high-frequency subgraph and the low-frequency subgraph are concatenated to obtain a concatenated graph. Second, the object detection stage. In this stage, the trained YOLOv6 object detection model is used. For each image block, the concatenated graph corresponding to the image block is input into the lightweight network EfficientRep in the YOLOv6 object detection model. The feature maps output by the network block ERBlock3, the network block ERBlock4, and the network block ERBlock5 in the lightweight network EfficientRep are respectively input into the neck network RepBi-PAN structure in the YOLOv6 object detection model. The RepBi-PAN structure extracts and combines effective information for subsequent object localization and classification, so as to accurately detect the defect area and the corresponding defect type in the image block. The YOLOv6 object detection model is used to process the concatenated graphs of all image blocks in the video frame, so as to detect the defect area and the corresponding defect type in all image blocks. Third, the post-processing stage. For each video frame, the post-processing stage obtains the defect area and the corresponding defect type (i.e., the predicted label) in all image blocks segmented from the video frame. Then, through the fusion post-processing module, the defect areas and the corresponding defect types in all image blocks are fused to generate a defect degree map, and the defect degree map and the low-frequency subgraphs corresponding to all image blocks are fused to generate a mask of the defect area in the video to be detected. At the same time, according to the defect area and the corresponding defect type in the video frame, the defect score corresponding to the video frame (i.e., 0.75 shown in the figure) can be determined. In this way, using the fast and high-precision video defect detection model YOLOv6, the defect area and the corresponding defect type in the video frame are accurately identified, realizing the high-precision detection of common local defects. At the same time, a mask of the defect area that is highly consistent with the subjective perception of the human eye is provided, and the severity of the defect is scored.
[0111] An exemplary description of the training process of the above detection model is given below.
[0112] Figure 4 The flowchart showing a training method of a detection model according to an embodiment of the present disclosure. Exemplarily, the method can be applied to various types of electronic devices. As Figure 4 shown, the method may include the following steps:
[0113] Step 401, construct a training sample set; wherein, the training sample set includes a plurality of image patches and the label of each image patch in the plurality of image patches, and the label of each image patch includes the label of whether there is a defective area in each image patch, the label of the defective area in each image patch, and the label of the defective type corresponding to the defective area; the defective types corresponding to the defective areas in the plurality of image patches are not all the same.
[0114] Exemplarily, for an image patch, if there is no defective area in the image patch, the label of the defective area and the label of the defective type corresponding to the defective area in the image patch are empty; if there is one defective area in the image patch, the label of the defective type corresponding to the defective area in the image patch is the defective type of the defective area, and if there are multiple defective areas in the image patch, the label of the defective type corresponding to the defective area in the image patch is the defective types of the multiple defective areas.
[0115] In a possible implementation, the defective type includes at least one of: color level, color block, color spot. Exemplarily, the defective type may include color level, color block, color spot, and the defective type corresponding to at least one image patch in the plurality of image patches is color level, the defective type corresponding to at least one image patch is color block, and the defective type corresponding to at least one image patch is color spot; in this way, there are various types of local defects in the constructed training sample set, so that the detection model trained by the training sample set has the ability to accurately identify various types of local defects.
[0116] Exemplarily, the training sample set can be constructed based on an open-source data set, or can be constructed based on a data set of a video platform.
[0117] Figure 5 Show a flowchart of constructing a training sample set according to an embodiment of the present disclosure, as Figure 5 shown, it may include the following steps:
[0118] Step 40101, obtain an original data set; the original data set includes videos with defects and / or videos without defects.
[0119] For example, video segments with various spatial resolutions (such as 4096x2160, 3840x2160, 1920x1080) and frame rates (such as 60fps, 50fps, 30fps, 25fps) can be collected from a video platform as the original data set.
[0120] In a possible implementation, the spatial resolution, frame rate, pixel format, etc. of each video in the original dataset can be unified. For example, the spatial resolution of the video can be adjusted to 1920x1080, and the frame rate of the video can be converted to 25fps. Then, some non-compliant videos can be cropped instead of unevenly scaled down, and the spatial resolution of the cropped video can be adjusted to the same resolution as above. For example, the spatial resolution of the cropped video can be downsampled to 1920x1080 to ensure the homogeneity and effectiveness of subsequent processing.
[0121] Step 40102: Screen out candidate videos in the original dataset that meet the preset diversity requirements.
[0122] Exemplarily, to ensure rich video content and data balance, videos that are too dark or bright, overly blurred, or colorful can be removed from the original dataset. Then, candidate videos that meet the preset diversity requirements can be screened out from the remaining videos.
[0123] Exemplarily, the diversity requirements can be that the characteristics such as contrast, brightness, clarity, and colorfulness of different candidate videos screened out satisfy a specific distribution, where the specific distribution can be set according to requirements.
[0124] In a possible implementation, contrast, brightness, clarity, and colorfulness can be selected as indicators to describe video diversity. The statistics and analysis of these characteristics are performed on the videos in the original dataset or the remaining videos mentioned above, so as to screen out candidate videos whose characteristics satisfy a specific distribution, which can ensure that the selected candidate videos have a wide coverage range in these visual characteristics, thereby improving the robustness of subsequent evaluation of the severity of defects in images and videos.
[0125] As an example, the candidate videos screened out can include videos with defects and videos without defects.
[0126] Step 40103: Extract sample images from the candidate videos, and obtain the defective regions in the sample images and the corresponding defect types of the defective regions.
[0127] In a possible implementation, extracting a sample image from the candidate video includes: when the candidate video is a video without defects, performing encoding processing on the candidate video so that defects exist in the encoded video; and extracting the sample image from the encoded video. In this way, for a candidate video without local defects, local defects are introduced into the candidate video by performing encoding processing on the candidate video. Exemplarily, for a candidate video without defects, mainstream video codecs such as H.264 / AVC, H.265 / HEVC, VP9, and bit-depth quantization can be introduced for encoding processing. When these encoders compress the video, defects or other types of visual distortions will occur under certain specific conditions (such as high-speed motion, low bit rate). For example, for the "color level" defect, it can be synthesized by smoothing the gradient. For example, a Gaussian filter can be used on the candidate video to smooth the random noise to generate rich gradient diversity, and these gradients can be quantified by a preset transformation factor, thereby introducing the "color level" defect into the original image. For the "color block" defect, it can be created by performing transform-domain quantization and deblocking filtering on the candidate video by the encoder. For the "color spot" defect, it can be created by performing a transformation on the candidate video by a preset transformation factor.
[0128] It can be understood that when performing encoding processing on a candidate video without defects, any one type of defect can be introduced into the candidate video, or any multiple types of defects can be introduced simultaneously. The same type of defect can also be at different positions in the same video, and this is not limited.
[0129] Exemplarily, in order to simulate the video defects existing in the real viewing scenario, accurate defect regions and defect types are marked for the candidate video with defects; then, video frames are extracted from the candidate video with defects as sample images, and at the same time, the defect regions in the marked sample images and the defect types corresponding to the defect regions are obtained.
[0130] Among them, the number of sample images can be set according to requirements. Exemplarily, sample images can be extracted from the candidate video with defects; among them, the candidate video with defects includes: the original candidate video with defects and the candidate video with local defects introduced after the above encoding processing; in this way, the extracted sample images all have defect regions. As an example, the spatial resolution of the sample image is 1920x1080, and the frame rate is 25fps.
[0131] Step 40104: Segment the sample image into the multiple image blocks.
[0132] In this step, each sample image is segmented into multiple image patches for local analysis of the sample image, and then local defects in the sample image are detected. Exemplarily, each sample image can be segmented into a preset number of image patches, where the value of the preset number can be set according to requirements. Among them, the sizes of different image patches can be the same or different. For example, each sample image can be divided into a preset number of image patches with a size of 256x256.
[0133] Exemplarily, for any sample image, there may or may not be an overlap between different segmented image patches. For example, a sliding window with a preset scale can be slid on the sample image at a fixed step size, so as to divide the sample image into one image patch after another. If the step size is less than the size of the image patch, there is an overlap between adjacent divided image patches. If the step size is not less than the size of the image patch, there is no overlap between adjacent divided image patches.
[0134] As an example, the sample image can be roughly segmented into a defective area and a non-defective area, and then a method of sliding window is used to extract a preset number of image patches from the roughly segmented sample image.
[0135] Step 40105: Determine the label of each image patch in the multiple image patches according to the defective area in the sample image and the corresponding defective type of the defective area.
[0136] In a possible implementation manner, the sample image can be first roughly segmented, and the segmentation result is labeled as a defective area and a non-defective area. Furthermore, each image patch segmented from the sample image is labeled with whether there is a defective area, the label of the defective area, and the label of the defective type corresponding to the defective area, so as to complete the image patch annotation. Exemplarily, if the overlapping part of an image patch and the defective area in the sample image exceeds 30% of the image patch, the image patch is marked as having a defective area. In this way, through automatic marking, compared with the method of manual labeling, the influence of subjective bias is avoided.
[0137] Step 402: Use the training sample set to train a preset model to obtain a trained detection model, where the detection model is used to detect the defective area and the corresponding defective type of the defective area in the image patches obtained by segmenting the video.
[0138] Among them, the training of the preset model can be carried out in an existing manner. In a possible implementation, the image patches in the training sample set can be sequentially input into the preset model. The preset model outputs the predicted defect regions and the corresponding defect types in the image patch. The loss function value is determined according to the predicted defect regions and the corresponding defect types in the image patch and the label of the image patch, and through backpropagation, the values of the parameters in the preset model are adjusted until the training termination condition is met, such as reaching the preset number of iterations or the preset model converges, etc., so as to obtain a trained detection model; through the high-precision detection model, various defects in the image patch can be accurately detected and located, and then various defects in the image and video can be determined. In particular, various local defects can be quickly and accurately detected. The trained detection model can provide comprehensive guarantee for the production quality of video content, and at the same time, the detailed defect classification is more in line with the subjective perception of defects by the human eye.
[0139] Exemplarily, the preset model can be any deep learning network with object detection function. For example, it can be the YOLOv6 object detection model. Among them, the backbone network of the YOLOv6 object detection model is the lightweight network EfficientRep, and the neck network is the Rep-PAN structure. The image patch is input into the trained YOLOv6 object detection model. Among them, the feature maps output by the network block ERBlock3, the network block ERBlock4, and the network block ERBlock5 in the lightweight network EfficientRep are respectively input into the neck network RepBi-PAN structure. The RepBi-PAN structure extracts and combines the effective information for subsequent object localization and classification, so as to accurately detect the defect regions and the corresponding defect types in the image patch.
[0140] In a possible implementation, the training of the preset model using the training sample set to obtain a trained detection model includes: for any one of the multiple image patches, generating the corresponding high-frequency sub-image and low-frequency sub-image of the image patch; wherein, the high-frequency sub-image is used to represent the texture information of the image patch, and the low-frequency sub-image is used to represent the structural information of the image patch; inputting the high-frequency sub-image and the low-frequency sub-image corresponding to the image patch into the preset model to obtain the defect prediction result of the image patch, and the defect prediction result includes the defect regions and the corresponding defect types in the image patch; based on the defect prediction result of the image patch and the label of the image patch, adjusting the parameters in the preset model to obtain the detection model. In this way, using the high-frequency sub-image and the low-frequency sub-image as the input of the preset model, representing the texture information and the structural information of the image patch respectively, aims to simulate the cognitive mechanism of the human brain, so as to more effectively identify the defect regions in the image patch.
[0141] As an example, taking the preset model as the YOLOv6 object detection model, for each image patch, the Sobel operator can be used to extract high-frequency detail information from the image patch to generate a high-frequency sub-image, and the piecewise smoothing algorithm can be used to process the image patch to obtain a low-frequency sub-image representing the overall structure and the change of the smooth region, so as to extract rich visual information from different frequency domains. Furthermore, after splicing the high-frequency sub-image and the low-frequency sub-image corresponding to the image patch, the obtained spliced image can be input into the YOLOv6 object detection model, and the YOLOv6 object detection model outputs the defect region and the corresponding defect type in the image patch. In this way, the defect in the image patch is detected by using the YOLOv6 object detection model and the high-low frequency image feature fusion technology, the detection accuracy is improved, and the high-precision defect region and the corresponding defect type in the image patch are effectively identified and located, especially when dealing with complex video content, the performance is particularly excellent.
[0142] In the embodiment of the present disclosure, a training sample set is constructed. Wherein, the training sample set includes a plurality of image patches and the label of each image patch in the plurality of image patches. The label of each image patch includes the label indicating whether there is a defect region in each image patch, the label of the defect region in each image patch, and the label of the corresponding defect type of the defect region. The corresponding defect types of the defect regions in each image patch in the plurality of image patches are not completely the same. The preset model is trained by using the training sample set to obtain a trained detection model, and the detection model is used to detect the defect region and the corresponding defect type of the defect region in the image patch obtained by segmenting the video. In this way, a training sample set with rich and balanced content is constructed. The training sample set contains image patches with various different types of defects, which can better reflect the problems that may occur after video coding and transmission in the real environment, and helps to improve the robustness and adaptability of the prediction model in complex scenarios. As an example, the training sample set can be established through operations such as raw data collection, video preprocessing, compression and quantization (bit-depth quantization), sample image segmentation, and image patch labeling, which solves the problem that the defect data set constructed by the existing method only contains a single defect type, so that a detection model driven by big data can be trained. The detection model trained by this training sample set can achieve fast and high-precision local defect detection, and can accurately identify different defect regions and the corresponding defect types of the defect regions.
[0143] For example, Figure 6 FIG. shows a schematic flowchart of a video defect detection method according to an embodiment of the present disclosure, as Figure 6As shown in the figure, it can be divided into a video data collection and preprocessing part, a defect dataset construction part, a high-precision defect detection model part, and a video defect scoring part. Among them: 1. Video data collection and preprocessing part: First, the original dataset can be collected from a video platform, and then the original video set is preprocessed. For example, through spatial resolution downsampling, videos with spatial resolutions of 4096x2160, 3840x2160, and 1920x1080 are uniformly adjusted to 1920x1080, and videos with frame rates of 60fps, 50fps, 30fps, and 25fps are unified to 25fps. 2. Defect dataset construction part: To establish a comprehensive defect database, first, the preprocessed videos are screened to select defective videos and non-defective videos, and the defective areas and defect types of the defective videos can be marked. Then, non-defective videos are compressed and quantified through H.264 / AVC, H.265 / HEVC, VP9, and bit-depth quantization to introduce local defects. Finally, the defective videos and the videos with local defects introduced are roughly segmented, the defective areas and the corresponding defect types of the roughly segmented defective areas are marked, and a sliding window is used to perform block segmentation on the roughly segmented videos, and the defective areas and the corresponding defect types of the segmented image blocks are marked. 3. High-precision defect detection model part: In the training stage, first, each segmented image block is preprocessed. For example, a high-frequency sub-image corresponding to the image block is generated through the Sobel operator, and a low-frequency sub-image corresponding to the image block is generated through a piecewise smoothing algorithm. After splicing the high-frequency sub-image and the low-frequency sub-image, it is input into the YOLOv6 object detection model to be trained (including the lightweight network EfficientRep and the neck network RepBi-PAN structure) for object detection, and the YOLOv6 object detection model is trained based on the prediction results and the annotation information of the image blocks. In the inference stage, first, each video frame in the video to be detected is segmented, and each segmented image block is preprocessed. Then, a high-frequency sub-image and a low-frequency sub-image corresponding to each image block are generated. After splicing the high-frequency sub-image and the low-frequency sub-image, they are input into the trained YOLOv6 object detection model for object detection, and the defective areas and the corresponding defect types in each image block are output. Finally, through a post-processing module, the defective areas and the corresponding defect types in each image block are summarized, and a defect score corresponding to each video frame is generated. 4. Video defect scoring part: In the inference stage, according to the defect scores corresponding to each video frame in the video to be detected, the defect score corresponding to the video to be detected is determined, and according to the relative size of the defect score corresponding to the video to be detected and a preset threshold, the defect severity of the video to be detected is divided into four levels: not obvious, slight, serious, and severe.
[0144] Based on the same inventive concept as the above method embodiments, embodiments of the present disclosure further provide a video defect detection device and a training device for a detection model, which can be used to execute the technical solutions described in the above method embodiments.
[0145] Figure 7 The structural schematic diagram of a video defect detection device according to an embodiment of the present disclosure is shown as Figure 7 shown. The device includes: a segmentation module 701, configured to segment each video frame in the video to be detected into at least one image block; an image block defect detection module 702, configured to perform object detection on each of the at least one image block through a trained detection model to determine the defect region and the defect type corresponding to the defect region in each image block; wherein the detection model is trained based on a plurality of image blocks, and the defect types corresponding to the defect regions in each of the plurality of image blocks are not completely the same; a video frame defect detection module 703, configured to determine the defect region and the defect type corresponding to the defect region in each video frame according to the defect regions and the defect types corresponding to the defect regions in the at least one image block; a video defect detection module 704, configured to determine the video segment with defects, the defect regions in the video segment, and the defect types corresponding to the defect regions in the video to be detected according to the defect regions and the defect types corresponding to the defect regions in each video frame.
[0146] In a possible implementation manner, the device further includes an evaluation module, configured to determine the defect score corresponding to each video frame according to the defect region and the defect type corresponding to the defect region in each video frame, where the defect score is used to represent the influence degree of the defect region on the visual effect; and determine the defect score corresponding to the video to be detected according to the defect scores corresponding to each video frame.
[0147] In a possible implementation manner, the image block defect detection module 702 is further configured to: generate a high-frequency sub-image and a low-frequency sub-image corresponding to any image block; wherein the high-frequency sub-image is used to characterize the texture information of the image block, and the low-frequency sub-image is used to characterize the structural information of the image block; after splicing the high-frequency sub-image and the low-frequency sub-image corresponding to the image block, input them into the detection model to obtain the defect region and the defect type corresponding to the defect region in the image block.
[0148] In a possible implementation, the video frame defect detection module 703 is further configured to generate a defect degree map for each video frame according to the defect regions in each of the at least one image block; the defect degree map is used to characterize the defect degrees corresponding to different image blocks; perform a fusion process on the defect degree map and the low-frequency sub-map corresponding to each image block to generate a mask for the defect region in each video frame; the mask is used to characterize the defect region.
[0149] In a possible implementation, the defect type includes at least one of color level, color block, and color spot.
[0150] Figure 8 The structural schematic diagram of a training device for a detection model according to an embodiment of the present disclosure is shown, as Figure 8 shown, the device includes: a sample construction module 801, configured to construct a training sample set; wherein, the training sample set includes a plurality of image blocks and labels for each of the plurality of image blocks, and the label for each image block includes a label indicating whether there is a defect region in each image block, a label for the defect region in each image block, and a label for the defect type corresponding to the defect region; the defect types corresponding to the defect regions in each of the plurality of image blocks are not completely the same; a training module 802, configured to train a preset model using the training sample set to obtain a trained detection model, and the detection model is used to detect the defect region and the defect type corresponding to the defect region in the image blocks obtained by segmenting the video.
[0151] In a possible implementation, the sample construction module 801 is further configured to: obtain an original data set; the original data set includes videos with defects and / or videos without defects; screen out candidate videos that meet a preset diversity requirement from the original data set; extract sample images from the candidate videos, and obtain the defect regions and the defect types corresponding to the defect regions in the sample images; segment the sample images into the plurality of image blocks; determine the label for each of the plurality of image blocks according to the defect regions and the defect types corresponding to the defect regions in the sample images.
[0152] In a possible implementation, the sample construction module 801 is further configured to: when the candidate video is a video without defects, perform encoding processing on the candidate video to make the encoded video have defects; extract the sample images from the encoded video.
[0153] In a possible implementation, the training module 802 is further configured to: for any one of the multiple image patches, generate a high-frequency sub-image and a low-frequency sub-image corresponding to the image patch; wherein, the high-frequency sub-image is used to represent the texture information of the image patch, and the low-frequency sub-image is used to represent the structural information of the image patch; input the high-frequency sub-image and the low-frequency sub-image corresponding to the image patch into the preset model to obtain a defect prediction result of the image patch, where the defect prediction result includes a defect area in the image patch and a defect type corresponding to the defect area; based on the defect prediction result of the image patch and the label of the image patch, adjust the parameters in the preset model to obtain the detection model.
[0154] In a possible implementation, the defect type includes at least one of: color level, color block, and color spot.
[0155] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0156] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0157] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above methods when executing the instructions stored in the memory.
[0158] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.
[0159] Figure 9 The block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 9 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above methods.
[0160] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0161] In an exemplary embodiment, there is also provided a non-transitory computer-readable storage medium, such as the memory 1932 including computer program instructions, and the computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0162] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0163] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punch card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0164] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0165] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0166] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0167] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions which implement various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0168] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0169] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0170] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A video defect detection method, characterized in that: The method comprises: Divide each video frame in the video to be detected into at least one image block; For each image block in the at least one image block, target detection is performed using a trained detection model to determine a defect area in each image block and a defect type corresponding to the defect area; wherein the detection model is trained based on multiple image blocks, and the defect types corresponding to the defect areas in each image block in the multiple image blocks are not completely the same; Determine the defective area and the defective type corresponding to the defective area in each video frame according to the defective area and the defective type corresponding to the defective area in the at least one image block; According to the defective area in each video frame and the defect type corresponding to the defective area, determine the video segment with defects in the video to be detected, the defective area in the video segment and the defect type corresponding to the defective area, The method further comprises: generating a defect level map of each video frame according to a defect area in each image block of the at least one image block, wherein the defect level map is used to characterize defect levels corresponding to different image blocks; The defect level map of each video frame and the low-frequency sub-map corresponding to each image block are fused to generate a mask of the defect area in each video frame, wherein the low-frequency sub-map is used to characterize the structural information of the image block, and the mask is used to characterize the defect area.
2. The method according to claim 1, characterized in that The method further comprises: Determine a defect score corresponding to each video frame according to the defect area in each video frame and the defect type corresponding to the defect area, wherein the defect score is used to indicate the degree of influence of the defect area on the visual effect; The defect score corresponding to the video to be detected is determined according to the defect score corresponding to each video frame.
3. The method according to claim 1, characterized in that The step of performing target detection on each image block of the at least one image block by using a trained detection model to determine a defect area in each image block and a defect type corresponding to the defect area includes: For any image block, a high-frequency sub-image and a low-frequency sub-image corresponding to the image block are generated; wherein the high-frequency sub-image is used to represent the texture information of the image block, and the low-frequency sub-image is used to represent the structural information of the image block; After the high-frequency sub-image and the low-frequency sub-image corresponding to the image block are spliced, they are input into the detection model to obtain the defect area in the image block and the defect type corresponding to the defect area.
4. The method according to claim 1, characterized in that: The defect type includes: at least one of: color scale, color block, and color spot.
5. A method for training a detection model, characterized in that: The method comprises: Constructing a training sample set; wherein the training sample set includes a plurality of image blocks and a label of each of the plurality of image blocks, wherein the label of each image block includes a label of whether there is a defective area in each image block, a label of the defective area in each image block, and a label of a defect type corresponding to the defective area; the defect types corresponding to the defective areas in each of the plurality of image blocks are not completely the same; The preset model is trained using the training sample set to obtain a trained detection model, and the detection model is applied to the method described in any one of claims 1-4 to detect defective areas and defect types corresponding to the defective areas in image blocks obtained by segmenting the video.
6. The method according to claim 5, characterized in that The step of constructing a training sample set includes: Acquire an original data set; the original data set includes videos with defects and / or videos without defects; Screening out candidate videos that meet preset diversity requirements from the original data set; Extracting a sample image from the candidate video, and obtaining a defect area in the sample image and a defect type corresponding to the defect area; dividing the sample image into the plurality of image blocks; A label of each image block in the plurality of image blocks is determined according to the defect area in the sample image and the defect type corresponding to the defect area.
7. The method according to claim 6, characterized in that The step of extracting a sample image from a candidate video includes: When the candidate video is a video without defects, encoding the candidate video so that defects exist in the encoded video; The sample image is extracted from the encoded video.
8. The method according to claim 5, characterized in that The method of training a preset model using the training sample set to obtain a trained detection model includes: For any image block among the multiple image blocks, generate a high-frequency sub-image and a low-frequency sub-image corresponding to the image block; wherein the high-frequency sub-image is used to represent texture information of the image block, and the low-frequency sub-image is used to represent structural information of the image block; Inputting the high-frequency sub-image and the low-frequency sub-image corresponding to the image block into the preset model to obtain a defect prediction result of the image block, wherein the defect prediction result includes a defect area in the image block and a defect type corresponding to the defect area; Based on the defect prediction result of the image block and the label of the image block, the parameters in the preset model are adjusted to obtain the detection model.
9. The method according to claim 5, characterized in that The defect type includes: at least one of: color scale, color block, and color spot.
10. A video defect detection device, characterized in that: The device comprises: A segmentation module, used for segmenting each video frame in the video to be detected into at least one image block; An image block defect detection module, configured to perform target detection on each image block of the at least one image block by using a trained detection model to determine a defect area in each image block and a defect type corresponding to the defect area; wherein the detection model is trained based on a plurality of image blocks, and the defect types corresponding to the defect areas in each of the plurality of image blocks are not completely the same; a video frame defect detection module, configured to determine the defect area and the defect type corresponding to the defect area in each video frame according to the defect area and the defect type corresponding to the defect area in the at least one image block; A video defect detection module, used to determine a video segment with defects in the video to be detected, a defective area in the video segment, and a defect type corresponding to the defective area according to the defective area in each video frame and the defect type corresponding to the defective area; The video frame defect detection module is further used to generate a defect degree map for each video frame according to the defect area in each image block in the at least one image block, wherein the defect degree map is used to characterize the defect degrees corresponding to different image blocks; and to fuse the defect degree map for each video frame and the low-frequency sub-image corresponding to each image block to generate a mask of the defect area in each video frame, wherein the low-frequency sub-image is used to characterize the structural information of the image block, and the mask is used to characterize the defect area.
11. A detection model training device, characterized in that: The device comprises: A sample construction module, used to construct a training sample set; wherein the training sample set includes a plurality of image blocks and a label of each of the plurality of image blocks, wherein the label of each image block includes a label of whether there is a defective area in each image block, a label of the defective area in each image block, and a label of a defect type corresponding to the defective area; the defect types corresponding to the defective areas in each of the plurality of image blocks are not completely the same; A training module is used to train a preset model using the training sample set to obtain a trained detection model. The detection model is applied to the method described in any one of claims 1 to 4 to detect defective areas and defect types corresponding to the defective areas in image blocks obtained by segmenting the video.
12. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method described in any one of claims 1 to 4 or the method described in any one of claims 5 to 9 when executing the instructions stored in the memory.
13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 4 or the method described in any one of claims 5 to 9 are implemented.
Citation Information
Patent Citations
Image defect detection method and related equipment
CN113781391A
Burr detection precision enhancing method and system for precision standard component
CN115564705A