Video Defect Detection Method and Device, Electronic Device, and Storage Medium

The shallow and deep feature maps of video frames are extracted through deep learning methods, combined with defect maps and attention mechanisms, and the problems of low efficiency and insufficient accuracy of video defect detection in the existing technology are solved, efficient and accurate video defect detection is achieved, and multiple business scenarios are supported.

CN118570698BActive Publication Date: 2025-07-08BEIJING YOUKU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410658844.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-24
Publication Date
2025-07-08
Estimated Expiration
2044-05-24

AI Technical Summary

Technical Problem

In the prior art, video defect detection relies on manual viewing, is inefficient and is susceptible to subjective factors, has insufficient detection accuracy and accuracy, and has a high error detection rate, making it difficult to meet the needs of large-scale high-resolution video defect detection.

Method used

The video defect detection method based on deep learning is adopted, and the shallow and deep feature maps of the video frame are extracted, combined with the defect map and attention mechanism, the defect areas in the video frame are accurately identified, and the accuracy of the detection results is improved by using feature fusion and feature map splicing.

Benefits of technology

It realizes efficient and accurate video defect detection, reduces the error detection rate, can identify defect areas at pixel level, supports video media delivery, ultra-high-definition video film selection and old film repair, and improves work efficiency and audience viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118570698B_ABST
    Figure CN118570698B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video defect detection method and apparatus, an electronic device, and a storage medium. The method includes: obtaining video frames of a video to be detected; extracting a shallow feature map and a deep feature map of the video frames; determining a defect map corresponding to the video frames according to the shallow feature map and the deep feature map, where the defect map represents a defect area where a defect exists in the predicted video frames; and determining a defect detection result corresponding to the video frames according to the shallow feature map, the deep feature map, and the defect map, where the defect detection result is used to indicate whether a defect exists in the video frames. According to the embodiments of the present disclosure, the efficiency and accuracy of video defect detection can be improved, and the false detection rate of video defect detection can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to a method and apparatus for video defect detection, an electronic device, and a storage medium. Background Art

[0002] In modern multimedia communication, video has become one of the main carriers of information dissemination. However, during processes such as video encoding, transcoding, transmission, and decoding, various factors such as device performance and network conditions can cause various defects in the video. For example, defects such as screen distortion, ghosting, and streaking, and for the case where old films are often stored on film, during long-term storage, due to physical and chemical changes and environmental factors, various damages such as dirt spots and scratches will inevitably occur, seriously affecting the artistic expressiveness and viewing experience of the film.

[0003] Video defect detection is to detect the defects existing in the video without the original video reference. These defects usually significantly reduce the quality of the video and seriously affect the user's viewing experience. Therefore, various video platforms need to conduct an overall evaluation of the video media provided by the film studio to determine the overall quality of the video, whether there are obvious defects, etc., and finally decide whether to accept the video media or produce ultra-high-definition videos, or repair the defective videos, such as repairing the video media of old films.

[0004] Currently, in the media production quality control system, the detection work (i.e., control work) for video defects mainly relies on manual viewing of a large number of video clips to make decisions, identify and mark various defects existing in the video frames. This process often takes a lot of time and has low efficiency. On the other hand, most of the evaluations given by manual assessment are for the overall viewing experience of the video, which is easily affected by subjective factors. Moreover, various defects are irregularly distributed in the picture and appear randomly in time. Some severely defective segments may be missed, resulting in low detection accuracy and high false detection rate. Summary of the Invention

[0005] In view of this, the present disclosure proposes a method and apparatus for video defect detection, an electronic device, and a storage medium, which can improve the efficiency and accuracy of video defect detection and reduce the false detection rate of video defect detection.

[0006] According to an aspect of the present disclosure, a method for video defect detection is provided, including: obtaining video frames of a video to be detected; extracting a shallow feature map and a deep feature map of the video frames; determining a defect map corresponding to the video frames according to the shallow feature map and the deep feature map, where the defect map represents a defect area where a defect in the predicted video frame is located; and determining a defect detection result corresponding to the video frames according to the shallow feature map, the deep feature map, and the defect map, where the defect detection result is used to indicate whether there is a defect in the video frames.

[0007] In a possible implementation, the extraction of the shallow feature map and the deep feature map of the video frame includes: dividing the video frame into a plurality of image blocks; inputting the plurality of image blocks into a converter module to obtain multi-scale feature maps extracted by N layers of converter units in the converter module, where the first-scale feature map extracted by the first layer of converter units is used as the shallow feature map, and N is a positive integer; performing a fusion process on the second-scale feature map to the N-scale feature map output by the second layer of converter units to the Nth layer of converter units in the converter module to obtain a first fusion feature map; and determining the deep feature map according to the first fusion feature map.

[0008] In a possible implementation, the determining of the deep feature map according to the first fusion feature map includes: performing a dimensionality reduction process on the first fusion feature map to obtain a second fusion feature map; respectively performing self-attention extraction processes on the second fusion feature map by using a plurality of window attention mechanisms to obtain a plurality of self-attention feature maps, where the window sizes used for extracting self-attention in different window attention mechanisms are different; performing a pooling process on the second fusion feature map to obtain a pooled feature map; performing channel concatenation on the second fusion feature map, the plurality of self-attention feature maps, and the pooled feature map to obtain a first concatenated feature map; and performing a dimensionality reduction process on the first concatenated feature map to obtain the deep feature map.

[0009] In a possible implementation, the determining of the defect map corresponding to the video frame according to the shallow feature map and the deep feature map includes: performing a dimensionality reduction process on the shallow feature map to obtain a first shallow feature map; performing an upsampling process on the deep feature map based on the resolution of the first shallow feature map to obtain a first deep feature map, where the resolution of the first deep feature map is the same as that of the first shallow feature map; performing channel concatenation on the first shallow feature map and the first deep feature map to obtain a second concatenated feature map; performing a dimensionality reduction process on the second concatenated feature map to obtain a defect map; and performing an upsampling process on the defect map based on the original resolution of the video frame to obtain a target defect map that matches the original resolution of the video frame.

[0010] In a possible implementation, determining the defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map, and the defect map includes: extracting features from the shallow feature map to obtain a second deep feature map, and performing dimensionality reduction processing on the second deep feature map to obtain a first intermediate feature map; based on the resolution of the shallow feature map, performing upsampling processing on the first intermediate feature map to obtain a second intermediate feature map, where the resolution of the second intermediate feature map is the same as that of the shallow feature map; respectively performing dimensionality reduction processing on the shallow feature map, the first deep feature map, and the defect map to obtain a second shallow feature map, a third deep feature map, and a first defect map; performing channel splicing on the second intermediate feature map, the second shallow feature map, the third deep feature map, and the first defect map to obtain a third spliced feature map; performing dimensionality reduction processing on the third spliced feature map to obtain a defect confidence level, where the defect detection result includes the defect confidence level, and the defect confidence level represents the probability that the video frame has a defect.

[0011] In a possible implementation, the method further includes: when the defect detection result indicates that the video frame has a defect, determining the defect severity of the video frame according to at least one of the position, size, and quantity of the defect area characterized by the target defect map.

[0012] In a possible implementation, the method further includes: determining a video segment with continuously existing defects in the video according to the defect detection results of multiple video frames in the video.

[0013] In a possible implementation, the method further includes: repairing the video frame with a defect according to the target defect map of the video frame with a defect to obtain a repaired video frame.

[0014] In a possible implementation, the method is implemented using a defect detection model, where the defect detection model includes a feature extractor, a defect locator, and a defect classifier; among them, the feature extractor is used to extract the shallow feature map and the deep feature map of the video frame, the defect locator is used to determine the defect map corresponding to the video frame according to the shallow feature map and the deep feature map, and the defect classifier is used to determine the defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map, and the defect map.

[0015] In a possible implementation, the defect detection model is trained using a sample data set; wherein, the sample data set is obtained through the following process: obtaining a plurality of original video samples; respectively performing random damage processing on the plurality of original video samples to obtain damaged video samples corresponding to the respective original video samples, wherein each damaged video sample includes at least one damaged video frame with a defect; according to the differences between the respective original video samples and the damaged video samples corresponding thereto, obtaining the damaged video frames in each damaged video sample and the mask graphs corresponding to the damaged video frames, the mask graphs being used to indicate the defect regions where the defects are located in the damaged video frames; wherein, the sample data set includes the plurality of damaged video samples corresponding to the plurality of original video samples, the mask graphs of the damaged video frames in each damaged video sample, and the class labels of the respective video frames in each damaged video sample, the class labels being used to indicate whether the video frame is a damaged video frame or an original video frame that has not been damaged.

[0016] In a possible implementation, the defects in the video frames include mosaic defects caused by partial data damage during video transmission; wherein, when the defect detection model is used to detect the mosaic defects, the step of respectively performing random damage processing on the plurality of original video samples to obtain damaged video samples corresponding to the respective original video samples includes: based on a preset damage model, performing data damage processing on the original video streams of the plurality of original video samples to obtain damaged video streams corresponding to the respective original video samples; using a video parser to parse the damaged video streams corresponding to the respective original video samples to obtain damaged video samples corresponding to the respective original video samples; wherein, the damage model is used to indicate the damage ratio, the damage position, and the damage length, the damage ratio representing the proportion of the damaged video frames in the original video sample, the damage position representing the position of the damaged video frame in the original video sample, and the damage length representing the length of the defect region in the damaged video frame.

[0017] In a possible implementation, the defects in the video frames include ghosting defects caused by adjacent frame overlap during video format conversion; wherein, when the defect detection model is used to detect the ghosting defects, the step of respectively performing random damage processing on the plurality of original video samples to obtain damaged video samples corresponding to the respective original video samples includes: for any one original video sample, estimating the motion regions of the respective video frames in the original video sample, the objects in the motion regions having a motion state, and determining at least one candidate ghosting frame according to the sizes of the motion regions of the respective video frames in the original video sample; overlapping each candidate ghosting frame in the original video sample with the video frames adjacent to each candidate ghosting frame to obtain a damaged video sample.

[0018] In a possible implementation, the defects in the video frame include the moiré defects caused by interpolating the odd and even lines of different frames of the original frame rate video into the same frame of the new frame rate video during the frame rate conversion of the interlaced video frame. The moiré defects have a horizontal comb shape. Among them, when the defect detection model is used to detect the moiré defects, the random corruption processing of the multiple original video samples respectively to obtain the damaged video samples corresponding to the respective original video samples includes: referring to the generation principle of the moiré defects, randomly introducing the moiré defects into at least one video frame of each original video sample to obtain the damaged video samples corresponding to the respective original video samples.

[0019] In a possible implementation, the defects in the video frame include physical defects caused by physical damage to the video film. The physical defects include linear traces caused by physical scratches in the video film, and / or irregular spots caused by stains or impurities in non-image structures on the chemical coating of the video film. Among them, when the defect detection model is used to detect the physical defects, the random corruption processing of the multiple original video samples respectively to obtain the damaged video samples corresponding to the respective original video samples includes: obtaining multiple film scan images and marking the region boundaries of the physical defects in each film scan image, where the film scan image is an image obtained by scanning the video film with the physical damage; dividing each film scan image into multiple pixel regions of a specified size, and respectively counting the number of defects of each physical defect that appears in each pixel region based on the region boundaries of the physical defects marked in each film scan image; fitting a gamma distribution model corresponding to each physical defect by performing maximum likelihood estimation on the number of defects of each physical defect that appears in the multiple pixel regions, where the gamma distribution model is used to characterize the quantity distribution characteristics of each physical defect; for any original video sample, sampling the target number of defects of each physical defect to be added to the video frames in the original video sample based on the gamma distribution model corresponding to each physical defect; based on the target number of defects of each physical defect, extracting position information matching the target number of defects from a preset spatial noise distribution for each physical defect, and extracting a rotation angle matching the target number of defects from a preset rotation angle range for each physical defect, where the position information is used to indicate the pixel position of the physical defect, and the rotation angle is used to indicate the direction of the physical defect; based on the position information and rotation angle matching the target number of defects of each physical defect, fusing the physical defects matching the target number of defects of each physical defect into at least one video frame of the original video sample to obtain the damaged video sample.

[0020] According to another aspect of the present disclosure, there is provided a video defect detection device, including: an acquisition module configured to acquire video frames of a video to be detected; a feature extraction module configured to extract a shallow feature map and a deep feature map of the video frames; a defect map determination module configured to determine a defect map corresponding to the video frames according to the shallow feature map and the deep feature map, where the defect map characterizes a defect area where a defect exists in the predicted video frames; and a result determination module configured to determine a defect detection result corresponding to the video frames according to the shallow feature map, the deep feature map, and the defect map, where the defect detection result is used to indicate whether there is a defect in the video frames.

[0021] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory configured to store instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0022] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor.

[0023] According to another aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0024] According to various aspects of the present disclosure, a pixel-level defect area can be predicted based on the shallow feature map and the deep feature map of video frames. This process is equivalent to a preliminary screening of video frames with defects. Then, by using the shallow feature map, the deep feature map, and the defect map, the defect detection result can be determined more accurately, that is, it can be more accurately determined whether there is a defect in the video frames. In particular, by introducing the defect map, it is equivalent to introducing an attention mechanism to pay more attention to the areas with defects, which can effectively improve the efficiency and accuracy of the defect detection result, thereby helping to reduce the false detection rate of defect detection.

[0025] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.

[0027] Figure 1 The flowchart showing a video defect detection method according to an embodiment of the present disclosure.

[0028] Figure 2 A schematic diagram showing the structure of a Transformer unit according to an embodiment of the present disclosure.

[0029] Figure 3 A schematic diagram showing the structure of a defect detection model according to an embodiment of the present disclosure.

[0030] Figure 4 A schematic diagram showing the structure of a defect detection model according to an embodiment of the present disclosure.

[0031] Figure 5 A schematic diagram showing the structure of a defect detection model according to an embodiment of the present disclosure.

[0032] Figure 6 A schematic diagram showing the process of detecting a screen freeze defect according to an embodiment of the present disclosure.

[0033] Figure 7 A schematic diagram showing the process of global defect detection according to an embodiment of the present disclosure.

[0034] Figure 8 A schematic diagram showing the process of physical defect detection according to an embodiment of the present disclosure.

[0035] Figure 9 A block diagram showing a video defect detection device according to an embodiment of the present disclosure.

[0036] Figure 10 A block diagram showing an electronic device 1900 according to an embodiment of the present disclosure. Detailed implementation manners

[0037] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0038] The special term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not necessarily have to be construed as superior to or better than other embodiments.

[0039] The term "and / or" herein merely describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0040] It should be understood that the terms "first", "second", "third", etc. in the claims, the specification and the drawings of the present disclosure are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the specification and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0041] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements and circuits well-known to those skilled in the art are not described in detail, so as to highlight the gist of the present disclosure.

[0042] To better understand the solutions of the embodiments of the present disclosure, the following first introduces the relevant terms and concepts that may be involved in the embodiments of the present disclosure.

[0043] (1) Media Production Quality Control (MQC): Refers to the process of verifying whether a media file meets the standards and specifications.

[0044] (2) H.264 Advanced Video Coding (AVC): Refers to a video bitstream compressed using the H.264 coding standard, which has the advantages of efficient transmission and storage. The H.264 encoder has a high compression ratio and video quality, and has become the mainstream choice in the field of video transmission and storage. Among them, frame data occupies a major proportion (usually more than 99.9%) in the H.264 bitstream and is most likely to be damaged.

[0045] (3) H.265 High Efficiency Video Coding (HEVC): Is a new video coding standard and the successor of H.264 / AVC. It has the advantage of a higher data compression ratio and can reduce the data traffic by about 50% while maintaining the same video quality. Due to its higher compression ratio of video data, any data damage encountered during the transmission or storage of an H.265-encoded video may cause more obvious screen distortion problems.

[0046] (4) MPEG-2 (Moving Picture Experts Group 2): It is an early digital video compression standard mainly used in television broadcasting and DVD video formats. Although its compression efficiency is not as good as H.264 and H.265, it is still widely used in some specific applications and early video materials. Compared with newer coding standards, when MPEG-2 encoded video is damaged during transmission, it will also cause visible distortion and defects in image quality.

[0047] (5)FFmpeg is an open source computer program that can be used to record, convert digital audio and video, and convert them into streams.

[0048] (6) Video Hit: During video transmission, due to software or hardware failures or network environment problems, video playback may experience video hits due to data packet loss or errors, manifested as abnormal color, blocky or striped distortion, etc., which seriously affects the audience's viewing experience.

[0049] (7) Video Global Artifacts: refers to defects whose formation principle is related to the overall overlap of the previous and next frames of the video, such as ghosting and drawing.

[0050] (8) Ghosting: Ghosting is caused by errors in shooting settings or format conversion after shooting. Due to the overlap of previous and next frames, the video will become jittery and blurry when viewed.

[0051] (9) Combing: When frame rate conversion is performed on interlaced video frames, combing defects are caused by interpolating odd and even lines of different frames of the original frame rate video into the same frame of the new frame rate video. Combing defects have horizontal comb teeth shape, and the audience will perceive obvious silk-like horizontal stripes.

[0052] (10) Scratches: Film damage may cause physical damage such as scratches, which usually appear as regular or irregular linear marks on the image. These physical defects will destroy the continuity and integrity of the original image and reduce the image quality.

[0053] (11) Dirt: Non-image structure stains or impurities that appear on the chemical coating of the film (usually the emulsion layer). After imaging, they appear as physical defects such as small black spots, white spots or irregular spots of other colors on the photo or movie screen.

[0054] (12) Perlin Noise: A natural noise generation algorithm widely used in computer graphics. This algorithm is mainly used to create smooth, continuous, and seemingly random textures or patterns to simulate many phenomena in nature, such as terrain undulations, cloud distributions, water flow fluctuations, etc.

[0055] (13) Transformer: The Transformer uses the self-attention mechanism to process sequence data. Without relying on sequential calculations, it can process each element of the input sequence in parallel, thus greatly improving the training efficiency.

[0056] (14) Window Attention (WinAttn): An attention mechanism that divides the input feature map into different windows so that the model can perform self-attention calculations on each window.

[0057] As described above, in the early quality control work of the media production quality control system, for video medium delivery and whether to produce ultra-high-definition videos, defects were mainly screened by manually watching the entire video and subjective judgment decisions were made. The traditional way of manually watching the entire video not only consumed a huge amount of time and the related processes were time-consuming and laborious, but also difficult to meet the needs of large-scale high-resolution video defect detection; and although some existing defect detection methods can identify and locate the defect positions to a certain extent, the accuracy is often not high, making it impossible to use the accurate position information of the defects in the defect video repair process, resulting in poor repair effects.

[0058] The video defect detection method of the embodiments of the present disclosure can provide the media production quality control system with the ability to detect video defects (such as screen freezing, ghosting, streaking, scratches, dirt spots, etc.), effectively support services such as video medium delivery, ultra-high-definition video selection, and old film repair, can accurately identify and evaluate various defects in video media and video streams, achieve pixel-level defect area detection, can provide defect areas at the pixel scale and corresponding accurate defect parameters (such as the position, size, quantity, etc. of the defect area), as well as the severity evaluation of defective video frames, and reduce the cost of manual review.

[0059] The video defect detection method of the embodiments of the present disclosure has efficient data processing capabilities and accurate image analysis capabilities, can perform pixel-level analysis on each frame, accurately identify the defect area, reduce the workload, and provide efficient and accurate video defect detection and evaluation for the media production quality control system. This not only significantly improves the performance and accuracy of defect detection, but also enables the rapid identification of defect quality problems in multiple business scenarios such as video medium delivery and ultra-high-definition video selection, effectively guaranteeing the high-quality viewing experience of the audience.

[0060] In practical applications, the video defect detection method of the embodiments of the present disclosure can be deployed on various terminal devices through software or hardware transformation. The terminal devices involved in this application may refer to devices with wireless connection functions. The wireless connection function means that it can be connected to other terminal devices through wireless connection methods such as Wi-Fi and Bluetooth. The terminal devices of this application can also have the function of wired connection for communication. The terminal devices of this application can be touch-screen, non-touch-screen, or without a screen. Touch-screen devices can be controlled by clicking, swiping, etc. on the display screen with fingers, styluses, etc. Non-touch-screen devices can be connected to input devices such as mice, keyboards, and touch panels to control the terminal devices. Devices without a screen can be, for example, Bluetooth speakers without a screen. For example, the terminal devices of the present disclosure can be computers, smart phones, netbooks, tablet computers, laptop computers, wearable electronic devices (such as smart bracelets, smart watches, etc.), TVs, virtual reality devices, etc.

[0061] The video defect detection method of the embodiments of the present disclosure can also be deployed on a server. The server can be located in the cloud or locally, and can be a physical device or a virtual device such as a virtual machine or a container, and has a wireless communication function. Among them, the wireless communication function can be set in the chip (system) or other components or assemblies of the server. It can refer to a device with a wireless connection function. The wireless connection function means that it can be connected to other servers or terminal devices through wireless connection methods such as Wi-Fi and Bluetooth. The servers of the present disclosure can also have the function of wired connection for communication. For example, the server of the present disclosure can be located in the cloud and communicate with terminal devices. The server receives the video to be detected sent by the terminal device, and uses the video defect detection method deployed on the server to perform defect detection on the video, obtains the defect detection results of the video frames in the video, and can return the defect detection results and the video frames with defects to the terminal device to display the defect detection results of the video frames through the terminal device.

[0062] The following is a detailed introduction to the video defect detection method of the embodiments of the present disclosure through Figures 1 to 8 the following.

[0063] Figure 1 FIG. shows a flowchart of a video defect detection method according to an embodiment of the present disclosure. This method can be executed by the above terminal device or server. As Figure 1 shown, the video defect detection method includes: step S11 to step S14.

[0064] In step S11, obtain video frames of the video to be detected.

[0065] Among them, the video to be detected can be, for example, a video captured by a video shooting device (such as a camera), or a video transmitted from local storage or other electronic devices, or a video obtained by digitizing the film of a movie. The embodiments of the present disclosure do not limit the acquisition method of the video. It should be understood that the embodiments of the present disclosure do not limit the content, format (such as mp4, avi, mkv, etc.), type (such as video stream or video frame sequence), duration, etc. of the video to be detected.

[0066] Among them, the video frames of the video can be each video frame in the video, that is, the defect detection can be performed on the video frames in the video frame by frame. Considering that in some scenarios, the video quality control objects are mostly medium and long videos, that is, the duration of the video to be detected may be long. For example, the number of video frames of a movie is about in the order of one hundred thousand. To achieve the balance between detection accuracy and detection efficiency, it is also possible not to perform defect detection on all video frames of the video, or to sample the video to extract some video frames for defect detection.

[0067] For example, after obtaining the video to be detected, the video can be analyzed to obtain information such as the frame rate, duration, resolution, bit rate, etc. of the video, and then based on this information and a preset sampling rule, an adaptive sampling interval is generated, and the video is frame-extracted according to the adaptive sampling interval. For example, the FFmpeg tool can be used to extract frames according to the sampling interval to obtain a set of video frames for subsequent defect detection. This can significantly improve the video defect detection performance while covering the entire video, greatly reduce the running cost of defect detection, and improve the detection efficiency.

[0068] In step S12, the shallow feature map and the deep feature map of the video frame are extracted.

[0069] Among them, the deep feature map represents the depth features of the video frame extracted in the deeper layer (such as the middle layer or the layer close to the output layer) of the feature extractor, and the shallow feature map represents the shallow features of the video frame extracted in the shallower layer (such as the first layer or the first few layers) of the feature extractor. In practical applications, known encoders in the art can be used to extract the shallow feature map and the deep feature map of the video frame. For example, the encoder of DeepLabV3+ (a semantic segmentation model of deep learning) can be used as the feature extractor, and the feature map output by the first layer of the feature extractor can be used as the shallow feature map (for example, a feature map with 1 / 4 resolution of the input video frame), and the feature map output by the last layer of the feature extractor can be used as the deep feature map.

[0070] Optionally, embodiments of the present disclosure further provide a feature extraction method implemented using a Transformer architecture, which can capture more fine-grained image context information and extract rich multi-scale features. Specifically, the above-mentioned extraction of the shallow feature map and the deep feature map of the video frame includes:

[0071] Dividing the video frame into multiple image patches;

[0072] Inputting the multiple image patches into a converter module to obtain a multi-scale feature map extracted by N converter units in the converter module. Among them, the first-scale feature map extracted by the first converter unit is used as the shallow feature map, and N is a positive integer;

[0073] Fusing the second-scale feature map to the N-scale feature map output by the second converter unit to the Nth converter unit in the converter module to obtain a first fused feature map;

[0074] Determining the deep feature map according to the first fused feature map.

[0075] Among them, for example, the OPM (Overlapped Patch Merging) mechanism can be used to process the input video frame, and the video frame is divided into multiple image patches as the input of the converter module. Among them, the multiple image patches divided from the video frame can be overlapped as the input of the converter module, which is equivalent to processing the video frame into serialized data, so that the Transformer can capture context information and extract more detailed multi-scale features.

[0076] In practical applications, each converter unit (i.e., Transformer Block) in the converter module can adopt an existing Transformer unit structure. For example, Figure 2 A shown Transformer unit structure (i.e., the structure of a Transformer Block) is adopted, which mainly includes an attention module (Attention Block), a feed-forward neural network (Feed Forward Network), and a normalization layer (LayerNorm). The embodiments of the present disclosure do not limit this. By using the converter module, it can be more sensitive to extracting the detailed features of the video frame, which is beneficial to extracting a feature map with more details. For example, multi-scale feature maps (i.e., multi-level features) with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original video frame resolution can be obtained through multiple converter units.

[0077] Among them, the higher the level of the converter unit in the converter module (the closer to the output), the smaller the scale (i.e., resolution) of the features extracted by the converter unit and the higher the channel dimension. Therefore, the first-scale feature map extracted by the first-layer converter unit can be used as the shallow feature map; then, multi-scale feature fusion processing is performed on the second-scale feature map to the N-scale feature map output by the second-layer converter unit to the Nth-layer converter unit in the converter module. For example, different-scale feature maps can be uniformly mapped to the size of 1 / 8 of the original image resolution through a linear layer (such as MLP), and then feature fusion is performed by combining channel splicing and dimensionality reduction mapping. In this way, in-depth fusion of context information between different-scale feature maps can be achieved, thereby obtaining a more refined and feature-rich first fusion feature map. That is, the first fusion feature map contains context information between different-scale feature maps, which is conducive to obtaining a deep feature map with rich depth features based on the first fusion feature map, and is conducive to improving the accuracy of the subsequent defect map and defect detection results.

[0078] Among them, known feature fusion methods in the art can be used to implement the fusion processing of the second-scale feature map to the N-scale feature map output by the second-layer converter unit to the Nth-layer converter unit in the converter module. For example, through dimensionality reduction processing and upsampling processing, the second-scale feature map to the N-scale feature map can be first reduced to a low dimension and unified to the same scale to obtain N - 1 intermediate feature maps with the same scale and low dimension, and then the N - 1 intermediate feature maps are subjected to channel splicing to obtain the first fusion feature map.

[0079] Among them, for example, by performing processing methods such as dimensionality reduction processing, pooling processing, and adding an attention mechanism on the first fusion feature map, a deep feature map can be obtained. In order to improve the richness of the deep feature map, the deep feature map can include global features and local features of various granularities. Thus, in a possible implementation manner, the above-mentioned determination of the deep feature map according to the first fusion feature map includes:

[0080] Perform dimensionality reduction processing on the first fusion feature map to obtain a second fusion feature map;

[0081] Use multiple window attention mechanisms to perform self-attention extraction processing on the second fusion feature map respectively to obtain multiple self-attention feature maps, where the window sizes used for extracting self-attention in different window attention mechanisms are different;

[0082] Perform pooling processing on the second fusion feature map to obtain a pooled feature map;

[0083] Perform channel splicing on the second fusion feature map, multiple self-attention feature maps, and the pooled feature map to obtain a first spliced feature map;

[0084] Perform dimensionality reduction on the first spliced feature map to obtain a deep feature map.

[0085] As described above, the first fused feature map can be obtained by splicing multiple feature maps. Therefore, dimensionality reduction can be first performed on the first fused feature map with a higher dimension. For example, the first fused feature map with 512 channels can be reduced to a second fused feature map with 128 channels. Here, dimensionality reduction can be understood as mapping features in a high-dimensional space to a low-dimensional space, which can reduce the computational complexity. At the same time, the second fused feature map still contains the context information between different-scale feature maps.

[0086] It should be understood that the embodiments of the present disclosure do not limit the specific implementation manner of dimensionality reduction. For example, a multilayer perceptron (MLP) can be used, or a convolutional network can also be used. The embodiments of the present disclosure do not limit this.

[0087] Among them, the window sizes of the window attention mechanisms are different, and the granularity of the local features extracted by them is different. Specifically, the smaller the window, the finer the granularity of the local features extracted. By using multiple window attention mechanisms with different window sizes, multiple self-attention feature maps can be extracted, and each self-attention mechanism can represent local features of various granularities. It should be understood that the embodiments of the present disclosure do not limit the number of window attention mechanisms and the window sizes adopted by each of them. For example, window sizes with ratios R = 2, 4, and 8 can be respectively adopted. A window with a ratio of R contains R feature patches. For example, a window with R = 2 contains 2 feature patches, and a window with R = 4 contains 4 feature patches.

[0088] Among them, known pooling methods in the art can be used, such as average pooling, max pooling, random pooling, etc., to perform pooling on the second fused feature map. The embodiments of the present disclosure do not limit this. Global features can be extracted while reducing the computational complexity through pooling.

[0089] It should be understood that by splicing the second fused feature map, multiple self-attention feature maps, and the pooled feature map in channels, and then performing dimensionality reduction on the first spliced feature map to obtain a deep feature map, the deep feature map can fuse the context information contained in the second fused feature map, the local features of various granularities contained in multiple self-attention feature maps, and the global features contained in the pooled feature map, improving the fineness and richness of the deep feature map, and at the same time facilitating the reduction of the computational complexity of the subsequent deep feature map.

[0090] It can be known that for a feature map of size PxP (i.e., the number of feature patches in the feature map is PxP), each feature patch needs to calculate the attention with all other feature patches. Therefore, the computational cost required to calculate the window self-attention is proportional to the size PxP of the feature map. For example, for window attention with a ratio R (R = 2, 4, 8), the feature map of PxP will be expanded to RPxRP to calculate the self-attention, and then the computational cost is increased by RxR times. Therefore, in order to reduce the computational cost, when calculating the self-attention using the window attention mechanism, the expanded RPxRP feature map can be pooled to the PxP feature map first, and then the window attention is calculated, which reduces the computational cost and enhances the global perception ability at the same time.

[0091] In step S13, according to the shallow feature map and the deep feature map, a defect map corresponding to the video frame is determined, and the defect map represents the defect area where the defect is located in the predicted video frame.

[0092] Optionally, a known decoder in the art can be used. For example, the decoder of DeepLabV3+ can be adopted to implement determining the defect map corresponding to the video frame according to the shallow feature map and the deep feature map. The embodiments of the present disclosure are not limited thereto.

[0093] Optionally, the shallow feature map and the deep feature map can also be fused, and then the defect map is determined based on the fused feature map. This is equivalent to supplementing the shallow feature information in the shallow feature map to the deep feature map, thereby enhancing the features of the deep feature map and making the fused feature map enhance the shallow feature information of the image, so as to facilitate more accurate pixel-level defect area prediction. Specifically, the above determining the defect map corresponding to the video frame according to the shallow feature map and the deep feature map may include:

[0094] Perform dimensionality reduction processing on the shallow feature map to obtain a first shallow feature map;

[0095] Based on the resolution of the first shallow feature map, perform upsampling processing on the deep feature map to obtain a first deep feature map, and the resolution of the first deep feature map is the same as that of the first shallow feature map;

[0096] Perform channel splicing on the first shallow feature map and the first deep feature map to obtain a second spliced feature map;

[0097] Perform dimensionality reduction processing on the second spliced feature map to obtain a defect map;

[0098] Based on the original resolution of the video frame, perform upsampling processing on the defect map to obtain a target defect map that matches the original resolution of the video frame. Among them, the target defect map can be used to implement subsequent tasks such as evaluating the defect verification degree and repairing the defect.

[0099] Among them, dimensionality reduction methods known in the art can be adopted. For example, an MLP can be used to perform dimensionality reduction processing on the shallow feature map to obtain a first shallow feature map. It should be understood that the first shallow feature map obtained after dimensionality reduction still contains the shallow feature information of the video frame. Then, through upsampling processing on the deep feature map (for example, it can be upsampled to 1 / 4 of the size of the input video frame), a first deep feature map with the same resolution as the first shallow feature map is obtained, which is convenient for subsequent channel splicing of the first shallow feature map and the first deep feature map with the same resolution to obtain a second spliced feature map. The second spliced feature map fuses the shallow information in the shallow feature map and the deep information in the deep feature map, realizing feature enhancement of the deep feature map.

[0100] Since the channel dimension of the second spliced feature map after splicing is relatively high, dimensionality reduction processing can be performed on the second spliced feature map. For example, it can be reduced to a single-channel defect map. At this time, the output defect map can represent the defect area where the predicted video frame defect is located. Since the deep feature map and the shallow feature map are extracted by splitting the video frame into multiple image blocks, the resolution of the defect map is usually smaller than the original resolution of the video frame. Therefore, upsampling processing can be performed on the defect map to obtain a target defect map with the same resolution as the original resolution of the video frame, which can make the target defect map represent a more accurate predicted pixel-level defect area. The above process can be briefly described as that the deep feature map undergoes upsampling processing, feature splicing with the shallow feature map after dimensionality reduction processing is performed, realizing feature enhancement of the deep feature map, and then on the basis of the second spliced feature map after shallow feature enhancement, a predicted defect map is obtained.

[0101] In step S14, according to the shallow feature map, the deep feature map, and the defect map, determine the defect detection result corresponding to the video frame, and the defect detection result is used to indicate whether the video frame has a defect.

[0102] Optionally, the shallow feature map, the deep feature map, and the defect map can be fused, and then based on the fused feature map, determine the defect detection result corresponding to the video frame. Among them, fusing the shallow feature map in the deep feature map can achieve feature enhancement. Since the defect map represents the defect area where the defect in the predicted video frame is located, introducing the defect map is equivalent to introducing an attention mechanism, which can make the fused feature map pay more attention to the defect area indicated by the defect map, thereby being beneficial to improving the accuracy of the defect detection result.

[0103] As described above, in the above step S13, the first deep feature map can be extracted to determine the defect map. Thus, in step S14, the first deep feature map and the defect map can be used to more accurately determine the defect detection result. Specifically, in a possible implementation manner, the above determining the defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map, and the defect map may include:

[0104] Feature extraction is performed on the shallow feature map to obtain a second deep feature map, and dimensionality reduction processing is performed on the second deep feature map to obtain a first intermediate feature map;

[0105] Based on the resolution of the shallow feature map, upsampling processing is performed on the first intermediate feature map to obtain a second intermediate feature map, and the resolution of the second intermediate feature map is the same as that of the shallow feature map;

[0106] Dimensionality reduction processing is respectively performed on the shallow feature map, the first deep feature map, and the defect map to obtain a second shallow feature map, a third deep feature map, and a first defect map;

[0107] Channel splicing is performed on the second intermediate feature map, the second shallow feature map, the third deep feature map, and the first defect map to obtain a third spliced feature map;

[0108] Dimensionality reduction processing is performed on the third spliced feature map to obtain a defect confidence level. The defect detection result includes the defect confidence level, and the defect confidence level represents the probability that a video frame has a defect.

[0109] Among them, for example, a multi-layer convolutional network or a multi-layer transformer unit can be used to perform feature extraction on the shallow feature map to obtain a second deep feature map. The second deep feature map may include deep feature information representing the category of the video frame, where the video frame category may indicate whether the video frame has a defect.

[0110] Among them, by performing dimensionality reduction processing on the second deep feature map and upsampling processing on the first intermediate feature map, it is convenient to perform channel splicing in the subsequent process. Performing dimensionality reduction processing on the shallow feature map, the first deep feature map, and the defect map can reduce the computational complexity; by performing channel splicing on the second intermediate feature map, the second shallow feature map, the third deep feature map, and the first defect map, the third spliced feature map obtained by channel splicing can fuse feature information from all aspects. Then, based on the third spliced feature map that fuses feature information from all aspects, a defect confidence level is generated, which can effectively improve the accuracy of the defect confidence level, and thus is beneficial to reducing the false detection rate of defect detection.

[0111] In practical applications, a confidence threshold (such as 95%) can be preset. When the detected defect confidence level is greater than the confidence threshold, it is considered that there is a defect in the video frame. Conversely, when the detected defect confidence level is less than or equal to the confidence threshold, it is considered that there is no defect in the video frame. It should be understood that the specific value of the confidence threshold in the embodiments of the present disclosure is not limited. By this means, video frames with low defect confidence levels can be filtered out, which is beneficial to reducing the false detection rate of defects.

[0112] After obtaining the defect detection results of video frames, the severity of the defects in the detected defective video frames can be evaluated based on the above-mentioned defect confidence levels and the defective areas characterized in the target defect map. In one possible implementation, the method may further include: when the defect detection results indicate that a video frame has defects, determining the severity of the defects in the video frame according to at least one of the position, size, and quantity of the defective areas characterized in the target defect map. Specifically, it is possible to first determine whether there are defects in the video frame based on the defect confidence levels, and then determine the severity of the defects in the video frame according to at least one of the position, size, and quantity of the defective areas characterized in the target defect map of the defective video frame.

[0113] Among them, the position of the defective area can represent the relative position of the defective area in the video frame, the size of the defective area can include area, length, width, etc., and the quantity of the defective areas represents the number of defective areas appearing in the video frame. It should be understood that those skilled in the art can implement the determination of the severity of the defects in the video frame according to the actual evaluation rules for the severity of the defects according to actual needs. For example, it can be set that the more centered the position of the defective area, the larger the area, and the more the quantity, the higher the severity of the defects is considered. The embodiments of the present disclosure are not limited thereto. Exemplarily, the severity of the defects can be quantitatively scored based on at least one of the position, size, and quantity of the defective areas characterized in the target defect map, and the severity of the defects in the video frame can be divided into the following levels: not obvious / slight, severe, and abominable, etc.

[0114] In practical applications, it is also possible to combine the defect detection results of multiple video frames in the video dimension and regard the time period during which defects are continuously detected as a defective video segment. Thus, in one possible implementation, the method may further include: determining the defective video segments continuously having defects in the video according to the defect detection results of multiple video frames in the video. Among them, determining the defective video segments continuously having defects in the video means determining the defective video segments composed of multiple consecutive video frames continuously having defects in the video.

[0115] In practical applications, the video frames can also be repaired based on the target defect map of the defective video frames. Thus, in one possible implementation, the method further includes: repairing the defective video frames according to the target defect map of the defective video frames to obtain the repaired video frames. In particular, repairing the defects in the video obtained by scanning the film of old films can improve the viewing experience of old film videos. It should be understood that those skilled in the art can use image repair techniques known in the art to repair the defective video frames, and the embodiments of the present disclosure do not limit this. In this way, not only can the severity of the defective areas in the video frames be accurately measured, but also the repair scheme can be effectively guided, which is beneficial to ensuring the quality of the normal areas while repairing the defective areas to the greatest extent.

[0116] In practical applications, the embodiments of the present disclosure can not only output the single-frame defect confidence, the target defect map, and the defect severity, but also output the representative defect frames that are obvious to the human eye (such as the video frames with the highest defect severity), the defective video segments with high defect severity (such as the defective video segments with the highest average defect severity), etc. by combining the defect detection results and defect severity of all video frames in the video. Moreover, based on the detected defect detection results, target defect maps, defect severity, and the above-mentioned representative defect frames and defective video segments and other detection information, a defect detection report can be generated. The report can include various detection information, video production suggestions, and possible video repair schemes, etc. It should be understood that those skilled in the art can design the information included in the defect detection report according to actual needs, and the embodiments of the present disclosure do not limit this. In this way, it provides high-value reference for video quality control, greatly improves work efficiency, and effectively supports various business scenarios such as video medium delivery services, ultra-high-definition video film selection services, and old film repair services.

[0117] According to the defect detection method of the embodiments of the present disclosure, the pixel-level defective areas can be predicted based on the shallow feature map and deep feature map of the video frames. This process is equivalent to a preliminary screening of the defective video frames. Then, by using the shallow feature map, deep feature map, and defect map, the defect detection results can be determined more accurately, that is, it can be more accurately determined whether there are defects in the video frames. In particular, by introducing the defect map, it is equivalent to introducing an attention mechanism to pay more attention to the defective areas, which can effectively improve the efficiency and accuracy of the defect detection results, thereby helping to reduce the false detection rate of defect detection.

[0118] Using the video defect detection method of the embodiments of the present disclosure, it is possible to support video source media with any resolution as input, support precise detection of defect areas and analysis of defect severity, use the terminal display resolution (such as 4K) as the evaluation standard, and be able to output multi-dimensional detection reports for each type of defect, including single-frame defect confidence, single-frame target defect map, single-frame defect severity, overall video defect severity, obvious defect segments of the video, etc. The introduction of the video defect detection method of the embodiments of the present disclosure into the media production quality control system can provide a reference for subsequent quality control and restoration of old films, greatly improve work efficiency, and effectively support operations such as video media delivery, selection of ultra-high-definition videos, and restoration of old films.

[0119] Using the video defect detection method of the embodiments of the present disclosure, for the problem that the original old film restoration plan affects the normal area of the video frame when repairing dirt and scratches, resulting in poor restoration effects, it is possible to output multi-dimensional information such as defect category, defect area map, defect feature parameters, single-frame / overall video defect severity, and defective video segments, providing an effective reference for manual judgment, greatly improving work efficiency, and while ensuring precise detection of the defect area, it can also provide accurate guidance for subsequent old film restoration work.

[0120] In practical applications, the above-mentioned defect detection method proposed in the embodiments of the present disclosure can also be implemented by designing a defect detection model. Among them, the defect detection model includes a feature extractor, a defect locator, and a defect classifier. The feature extractor is used to extract the shallow feature map and the deep feature map of the video frame. The defect locator is used to determine the defect map corresponding to the video frame according to the shallow feature map and the deep feature map. The defect classifier is used to determine the defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map, and the defect map.

[0121] In a possible implementation manner, the embodiments of the present disclosure provide Figure 3 a defect detection model shown as Figure 3 shown. In the feature extractor, multiple converter units (Transformer Block1, Transformer Block2, Transformer Block3, Transformer Block4) are used to extract multi-scale feature maps of multiple input image blocks. The first-scale feature map F extracted by the first-layer Transformer Block1 QAs the shallow feature map, three MLP layers, two upsampling layers (Upsample), and skip connections (Concat) are then used to fuse the second-scale to fourth-scale feature maps proposed by Transformer Block2, Transformer Block3, and Transformer Block4 to obtain the first fused feature map. Then, the MLP layer is used to perform dimensionality reduction on the first fused feature map to obtain the second fused feature map F2; then, three window attention mechanism layers (WinAttn) are used to perform self-attention extraction processing on the second fused feature map respectively to obtain three self-attention feature maps (F 31 , F 32 , F 33 ). The image pooling layer (Image Pooling) is used to perform pooling processing on the second fused feature map to obtain the pooled feature map F p ; then, the second fused feature map F2, the three self-attention feature maps (F 31 , F 32 , F 33 ) and the pooled feature map F p are concatenated in channels to obtain the first concatenated feature map F C1 ; then, an MLP layer is used to perform dimensionality reduction on the first concatenated feature map F C1 to obtain the deep feature map F S ;

[0122] Among them, in the defect locator, the MLP layer is used to perform dimensionality reduction on the shallow feature map F Q to obtain the first shallow feature map F Q1 . The upsampling layer (Upsample) is used to perform upsampling on the deep feature map F S to obtain the first deep feature map F Q1 with the same resolution as the first shallow feature map F S1 . Then, the first shallow feature map F Q1 and the first deep feature map F S1 are concatenated in channels to obtain the second concatenated feature map F C2 . Then, the MLP layer is used to perform dimensionality reduction on the second concatenated feature map F C2 to obtain the defect map F t1 . Using the upsampling layer (Upsample) based on the original resolution of the video frame, the defect map F t1 is upsampled to obtain the target defect map corresponding to the video frame;

[0123] Among them, in the defect classifier, stacked multiple Transformer Blocks and MLP are used for the shallow feature map F QFeature extraction is performed to obtain a second deep feature map, and dimensionality reduction processing is performed on the second deep feature map to obtain a first intermediate feature map F a1 ; The first intermediate feature map F a1 is upsampled using an upsampling layer (Upsample) to obtain a second intermediate feature map F Q with the same resolution as the shallow feature map F a2 ; Three MLP layers are respectively used to perform dimensionality reduction processing on the shallow feature map F Q , the first deep feature map F S1 and the defect map F t1 to obtain a second shallow feature map F Q2 , a third deep feature map F S3 and a first defect map F t2 ; Then, channel concatenation (Concat) is performed on the second intermediate feature map, the second shallow feature map, the third deep feature map, and the first defect map to obtain a third concatenated feature map F C3 ; An MLP layer is further used to perform dimensionality reduction processing (reduce to one dimension) on the third concatenated feature map F C3 to obtain the defect confidence level.

[0124] As Figure 3 shown, a post-processing module can be used to determine the defect severity, defect video segments, representative defect frames, and generate a defect detection report, etc. based on the target defect map and the defect confidence level.

[0125] According to Figure 3The shown defect detection model, in the feature extractor part, compared with the traditional DeepLabV3+ model, selects a more powerful Transformer architecture as the backbone network of the feature extractor, which makes the model more sensitive to the feature details of the input image and improves the feature extraction ability of the model for the input video frames. By dividing the video frames into multiple image patches as the input of the Transformer Block, finer-grained image information can be captured and rich multi-scale feature extraction can be achieved. By fusing between multi-scale feature maps, refined feature representations can be obtained, further promoting the in-depth fusion of features and context information. Multiple window attention mechanism layers and image pooling layers can be used to integrate context information from different scales, effectively capturing local and global features of various granularities. In the defect locator, the deep feature map output by the feature extractor is upsampled to the same scale as the shallow feature map, and then fused with the shallow feature map through a linear layer (i.e., MLP layer) and feature concatenation. This process enhances the shallow feature information of the image for the deep feature map, facilitating more accurate pixel-level defect area prediction. In the defect classifier, by implementing feature interaction and fusion between different feature maps, not only the learning efficiency of the model for single classification tasks is improved, but also the robustness of the model in complex scenarios is enhanced, while reducing the false detection rate.

[0126] In practical applications, when the defects in the video frames are small (such as scratches, dirt spots), the detection effect of the defect classifier may be poor. Therefore, it is also possible to determine whether there are defects in the video frames only based on the presence or absence of defect areas in the target defect map. Thus, as Figure 4 shown, only the target defect maps output by the feature extractor and the defect locator can be used to implement defect detection, and the embodiments of the present disclosure are not limited thereto.

[0127] As described above, the encoder in DeepLabV3+ can also be used as the feature extractor in the defect detection model, and the decoder in DeepLabV3+ can be used as the defect locator in the defect detection model. Thus, the embodiments of the present disclosure also provide Figure 5 a shown defect detection model. As Figure 5 shown, the encoder in DeepLabV3+ is used as the feature extractor, and the backbone network (i.e., the Deep Convolution Neural Network (DCNN)) outputs the shallow feature map F Qand a backbone feature map. Among them, the lightweight network MobileNetV2 can be selected as the backbone network in the feature extractor. Using the lightweight network MobileNetV2 can ensure the model performance. Then, the ASPP (Atrous Spatial Pyramid Pooling) module is used to generate the depth feature map F based on the backbone feature map S , where in ASPP, multiple atrous convolution blocks with different dilation rates rate (such as 3×3Conv rate 6 represents an atrous convolution block with a convolution kernel size of 3×3 and a dilation rate of 6) and an image pooling layer (ImagePooling) are used to perform atrous convolution processing and pooling processing on the backbone feature map extracted by the DCNN. Then, the feature maps output by each atrous convolution layer and the image pooling layer are concatenated, and then a 1×1 convolutional layer (1×1Conv) is used to perform dimensionality reduction processing on the concatenated feature map to obtain the depth feature map Fs. Furthermore, the decoder in DeepLabV3+ is used as the defect locator to perform convolution, upsampling (Upsample by 4 represents 4-fold upsampling), channel connection, etc. to realize based on the shallow feature map F Q and the depth feature map F S to output the defect map and the target defect map. In this way, while introducing multi-scale information, the shallow features and deep features are further fused, effectively improving the accuracy of the model in detecting the defect area; the defect classifier mainly uses multiple residual convolutional blocks (Residual Conv Blocks) to extract the second depth feature map of the shallow feature map, and mainly uses convolutional blocks for dimensionality reduction processing to realize the channel concatenation of multiple feature maps; because Figure 5 the feature extractor in mainly uses a convolutional structure, so when using the Figure 5 shown defect detection model for defect detection, there is no need to divide the video frame into multiple image blocks. Directly input the video frame into the defect detection model, and the target defect map and defect confidence can be output.

[0128] According to the Figure 5 shown defect detection model, compared with DeepLabV3+, by introducing a defect classifier to fuse the feature maps output by the feature extractor and the defect locator respectively, the defect classifier can output the defect detection result more accurately. It is equivalent to first roughly screening the candidate defect frames through the defect locator, and then filtering out the video frames with low confidence in the candidate defect frames through the defect locator, thus significantly reducing the defect misdetection rate and improving the defect detection accuracy.

[0129] In practical applications, the above defect detection model can be obtained by training with a sample data set; among them, the sample data set is obtained through the following process:

[0130] Step S61: Obtain multiple original video samples;

[0131] Step S62: Perform random damage processing on the multiple original video samples respectively to obtain damaged video samples corresponding to the respective original video samples. Each damaged video sample contains at least one damaged video frame with defects.

[0132] Step S63: According to the differences between each original video sample and its corresponding damaged video sample, obtain the damaged video frames in each damaged video sample and the mask graph corresponding to the damaged video frames. The mask graph is used to indicate the defective area where the defect is located in the damaged video frame.

[0133] Among them, the sample data set includes multiple damaged video samples corresponding to multiple original video samples, the mask graphs of the damaged video frames in each damaged video sample, and the class labels of each video frame in each damaged video sample. The class labels are used to indicate whether the video frame is a damaged video frame or an original video frame that has not been damaged.

[0134] In step S61, original video samples with different resolutions can be collected to ensure that different scenarios, motion patterns, and texture complexities can be covered. Original video samples with different video coding standards can also be collected. For example, original video samples with video coding standards such as H.264, H.265, and MPEG-2 can be collected. When collecting original video samples with different video coding standards, attention can be paid to the processing methods of frame data during the compression process. Because the damage of frame data has a particularly important impact on video quality during decoding. Whether it is a video encoded by H.264, H.265, or MPEG-2, due to the dependence of the inter-frame compression technology, once the key frame data is damaged, all subsequent frames depending on this frame may have defects during decoding, resulting in a significant decline in video quality. Therefore, by collecting original video samples with multiple video coding standards and conducting detailed analysis and preprocessing on them, a solid foundation can be laid for the training of subsequent defect detection models.

[0135] As mentioned above, the defects in the video frame include mosaic defects, ghosting defects, streaking defects, and physical defects (scratches, dirt spots). Therefore, different sample data sets can be constructed for different types of defects to train defect detection models for detecting different types of defects. For example, a defect detection model for detecting mosaic defects can be trained with a sample data set with mosaic defects, and a defect detection model for detecting ghosting defects can be trained with a sample data set with ghosting defects, and so on.

[0136] Thus, in a possible implementation, when the defect detection model is used to detect the screen freeze defect, in the above step S62, randomly damaging the multiple original video samples respectively to obtain the damaged video samples corresponding to the respective original video samples includes:

[0137] Performing data damage processing on the original video streams of the multiple original video samples based on a preset damage model to obtain the damaged video streams corresponding to the respective original video samples;

[0138] Using a video parser to parse the damaged video streams corresponding to the respective original video samples to obtain the damaged video samples corresponding to the respective original video samples;

[0139] Wherein, the damage model is used to indicate the damage ratio, damage position and damage length. The damage ratio represents the proportion of damaged video frames in the original video sample. The damage position represents the position of the damaged video frames in the original video sample. The damage length represents the length of the defect area in the damaged video frames.

[0140] Wherein, data damage can include data loss or errors (such as out-of-order), etc. In order to efficiently simulate various data damage situations that video content may encounter during transmission, storage and processing in the real world, a three-parameter damage model (P, L, S) is designed. Among them, P represents the damage ratio, and P can be randomly selected from 0 (no damage) to 1 (complete damage); L represents the damage position, and the damage position can be located based on the coding unit size (GOP, Group of Pictures) and the specific bitstream type. For example, for the H264 bitstream, data loss or errors are simulated near the key frames to produce a screen freeze effect; S represents the damage length, which represents the length of the defect area, and the upper and lower limits of S can be set according to the actual situation. It should be understood that the screen freeze defect is caused by data damage such as data loss or errors in the video stream. The more data damage there is in the video stream, the longer (i.e., larger) the defect area is. Therefore, the length of the data to be damaged in the video stream (i.e., the number of serialized data to be damaged) can be selected through S to control the length of the defect area. Then, according to this damage model, random packet loss or out-of-order, etc. damage processing can be performed on the original video streams of each original video sample to obtain the damaged video streams corresponding to the respective original video samples. Then, for the damaged video streams with different video coding standards, a bitstream parser (i.e., the video parser) that is compatible with various video coding standards such as H.264, H.265 and MPEG-2 can be used to parse the damaged video streams to generate damaged video samples with screen freeze defects.

[0141] It should be understood that the defective areas where the mosaic defects are located in the damaged video samples generated using the above damage model are unknown. Therefore, through the above step S63, based on the differences between each original video sample and its corresponding damaged video sample, the damaged video frames in each damaged video sample and the mask images corresponding to the damaged video frames can be obtained. Or rather, through the comparison between the damaged video sample and the original video sample, a damaged video sequence (i.e., multiple frames of damaged video frames) and a mask image (i.e., a mask sequence) for indicating the defective areas can be generated. These mask images can accurately indicate the defective areas where the mosaic defects are located in each frame of the damaged video frame. Furthermore, by combining the damaged video sequence and the corresponding mask sequence, a sample data set corresponding to the mosaic defect can be constructed to train a defect detection model for detecting the mosaic defect.

[0142] In a possible implementation manner, when the defect detection model is used to detect ghosting defects, in the above step S62, randomly damaging each of the multiple original video samples to obtain the damaged video samples corresponding to the respective original video samples includes:

[0143] For any original video sample, estimate the motion areas of the video frames in the original video sample. The objects in the motion areas have a motion state, and based on the sizes of the motion areas of the video frames in the original video sample, determine at least one candidate ghosting frame;

[0144] Overlap each candidate ghosting frame in the original video sample with the video frames adjacent to each candidate ghosting frame to obtain the damaged video sample.

[0145] In practical applications, known motion estimation techniques in the art, such as the optical flow estimation method, etc., can be used to estimate the motion areas of the video frames in the original video sample; then, based on the sizes of the motion areas of the video frames in the original video sample, select the video frames with larger motion areas (i.e., greater motion degrees) as candidate ghosting frames. For example, the video frames can be sorted in descending order according to the sizes of the motion areas and then select the top m video frames as candidate ghosting frames, or at least one video frame can also be sampled based on the descending order result as a candidate ghosting frame. The embodiments of the present disclosure do not limit this. Through this method, by selecting the video frames with larger motion areas, damaged video samples with ghosting defects can be effectively constructed.

[0146] Among them, referring to the generation principle of ghosting defects, each candidate ghosting frame in the original video sample can be overlapped with the video frames adjacent to each candidate ghosting frame to obtain a damaged video sample. Then, according to the differences between each original video sample and its corresponding damaged video sample through the above step S63, the damaged video frames in each damaged video sample and the mask map corresponding to the damaged video frames can be obtained. At this time, the mask map can indicate the defect area where the ghosting defect is located. Among them, the pixel-level label (i.e., the mask map) of the defect area and the class label of the damaged video frame can be obtained by using the difference between the damaged video frame and the original video frame. In addition, valid damaged video samples can be screened out through manual annotation to construct a sample data set corresponding to the ghosting defect. For example, a sample data set of 100,000 levels can be constructed and annotated. By constructing a sample data set corresponding to the ghosting defect, a defect detection model for detecting ghosting defects can be trained. It should be understood that those skilled in the art can design and develop relevant program codes according to the generation principle of ghosting defects to simulate ghosting defects in the original video sample. The embodiments of the present disclosure do not limit the generation method of ghosting defects.

[0147] In a possible implementation manner, when the defect detection model is used to detect streaking defects, in the above step S62, randomly damaging multiple original video samples respectively to obtain the damaged video samples corresponding to each original video sample includes: referring to the generation principle of streaking defects, randomly introducing streaking defects into at least one video frame in each original video sample to obtain the damaged video samples corresponding to each original video sample. As described above, streaking defects are caused by putting different frames of the original frame rate video into odd and even lines and interpolating them into the same frame of the new frame rate video when performing frame rate conversion on interlaced video frames. Therefore, those skilled in the art can design and develop relevant program codes according to the generation principle of streaking defects to randomly introduce streaking defects into at least one video frame in the original video sample to construct damaged video samples with streaking defects. Then, according to the differences between each original video sample and its corresponding damaged video sample through the above step S63, the damaged video frames in each damaged video sample and the mask map corresponding to the damaged video frames can be obtained. At this time, the mask map can indicate the defect area where the streaking defect is located. By constructing a sample data set corresponding to the streaking defect, a defect detection model for detecting streaking defects can be trained.

[0148] As described above, the defects in the video frames also include physical defects caused by physical damage to the video film. The physical defects include linear traces (i.e., scratches) caused by physical scratches in the video film, and / or irregular spots (i.e., dirt spots) caused by stains or impurities of non-image structures on the chemical coating of the video film. In a possible implementation manner, when the defect detection model is used to detect physical defects, in the above step S62, randomly damaging each of the multiple original video samples to obtain damaged video samples corresponding to each of the original video samples includes:

[0149] Obtain multiple film scan images, and mark the region boundaries of physical defects in each film scan image. The film scan image is an image obtained by scanning the video film with the physical damage;

[0150] Segment each film scan image into multiple pixel regions of a specified size, and based on the region boundaries of physical defects marked in each film scan image, respectively count the number of defects of each physical defect that appear in each pixel region;

[0151] By performing maximum likelihood estimation on the number of defects of each physical defect that appear in multiple pixel regions, fit a gamma distribution model corresponding to each physical defect. The gamma distribution model is used to characterize the quantity distribution characteristics of each physical defect;

[0152] For any original video sample, based on the gamma distribution model corresponding to each physical defect, sample the target number of defects of each physical defect to be added to the video frames in the original video sample;

[0153] Based on the target number of defects of each physical defect, extract position information matching the target number of defects from a preset spatial noise distribution for each physical defect, and extract a rotation angle matching the target number of defects from a preset rotation angle range for each physical defect. The position information is used to indicate the pixel position of the physical defect, and the rotation angle is used to indicate the direction of the physical defect;

[0154] Based on the position information and rotation angle matching the target number of defects of each physical defect, fuse the physical defects matching the target number of defects of each physical defect into at least one video frame of the original video sample to obtain a damaged video sample.

[0155] In practical applications, a batch of severely damaged 4K resolution film scan images can be collected, and then the polygon boundaries of each physical defect in the film scan images are marked. Then, each physical defect is extracted, and the polygon boundary is converted into a square region boundary by filling with zeros, creating a sample library containing physical defects.

[0156] Among them, in order to accurately record the number of dirty points and scratches, each scanned film image can be segmented into multiple pixel regions of size 256x256, and the number of two types of defects (scratches and dirty points) can be counted for each pixel region respectively, that is, the number of defects of each physical defect in each scanned film image is counted. Then, by performing maximum likelihood estimation on the number of defects of each physical defect counted, a gamma distribution model that best represents the defect data characteristics of each physical defect is fitted. Then, for each original video sample, sampling of the target defect number can be performed based on the gamma distribution model corresponding to each physical defect, that is, sampling using the gamma distribution for each physical defect, to obtain the target defect number of each physical defect to be added to the video frames in the original video sample.

[0157] Among them, Perlin noise can be used to approximately simulate a preset spatial noise distribution (that is, spatial defect density, spatial density noise distribution, etc., that is, the distribution of noise in three-dimensional space). In this way, each physical defect can independently extract its position information from the simulated spatial density noise distribution (that is, using the noise value extracted from the spatial density noise distribution as the position information of the physical defect). A rotation angle can also be extracted from a uniform distribution within a preset rotation angle range (such as [-180°, 180°]) to simulate any direction of the physical defect in two-dimensional space. The physical defect of the same shape can be rotated through the rotation angle, so that more distribution forms of physical defects can be generated.

[0158] Among them, physical defects matching the target defect number of each physical defect can be generated first based on the position information and rotation angle matching the target defect number of each physical defect. Then, image fusion techniques known in the art, such as Alpha Blending, can be used to superimpose the generated physical defects onto the normal video frames of the original video sample. Then, by adjusting the pixel transparency of each physical defect, the physical defect can be naturally fused with the video frame background, which is beneficial to making the physical defects in the damaged video sample have a more realistic fusion and superimposition effect.

[0159] Then, through the above step S63, based on the differences between each original video sample and its corresponding damaged video sample, the damaged video frames in each damaged video sample and the mask map corresponding to the damaged video frames can be obtained. At this time, the mask map can indicate the defect area where the physical defect is located. By constructing a sample data set corresponding to the physical defect, a defect detection model for detecting the physical defect can be trained.

[0160] It should be understood that those skilled in the art can adopt the model training methods known in the art to train the defect detection model using the sample data set, and the embodiments of the present disclosure do not limit this. Those skilled in the art can train defect detection models for detecting different defects using sample data sets corresponding to different types of defects, and then can use the trained defect detection models to implement video defect detection.

[0161] Exemplarily, Figure 6 A schematic diagram showing a process for detecting a screen freeze defect is shown, as Figure 6 shown. A screen freeze data set (i.e., the sample data set corresponding to the screen freeze defect) can be constructed using a three-parameter damage model, and the defect detection model can be trained using the screen freeze data set. Then, the screen freeze defect detection model obtained through training can be used to output a screen freeze defect map (i.e., the target defect map for characterizing the defect area of the predicted screen freeze defect) and a screen freeze defect confidence level (i.e., the probability that there is a screen freeze defect in the video frame); based on the screen freeze defect map and the screen freeze defect confidence level, the severity of the screen freeze defect in the video frame can be evaluated, that is, the screen freeze defect in the video frame can be classified as unobvious / slight, severe, and critical.

[0162] Exemplarily, Figure 7 A schematic diagram showing a process for detecting a global defect (ghosting defect or streaking defect) is shown, as Figure 7 shown. The video can be frame-sampled based on an adaptive sampling interval to obtain a set of video frames, and then the extracted video frames are input into the video global defect detection model (i.e., the defect detection model for detecting ghosting defects or streaking defects) frame by frame to output a defect confidence level and a target defect map; based on the defect confidence level and the target defect map, the severity of the global defect in the video frame can be evaluated; finally, a detection report can also be output, and the detection report can include information such as the severity statistics result "a total of 56 ghosting defects are detected, including 9 critical and 47 severe", production suggestions "the video ghosting defect is severe, it is recommended not to produce ultra-high-definition videos", and a repair plan "the current repair link does not support the repair of ghosting defects, it is recommended to replace the video source", etc., and the embodiments of the present disclosure do not limit this.

[0163] Exemplarily, Figure 8 A schematic diagram showing a process for detecting a physical defect (scratch, dirt spot) is shown, as Figure 8 shown. A defect simulation data set (i.e., constructing a sample data set with physical defects) is constructed, and a defect detection model including a feature extractor and a defect locator is trained. The trained defect detection model is used to accurately locate and identify physical defects in the video frame, that is, to output the target defect map of the video frame, and then the severity of the defect can be analyzed based on the target defect map, that is, the physical defects in the video frame are classified as slight, severe, and critical. For example, the more the number of physical defects and the more centered their positions, the greater the severity.

[0164] According to an embodiment of the present disclosure, on the one hand, starting from the data perspective, a three-parameter damage model is innovatively designed based on the generation principle of screen distortion, thereby creating a screen distortion dataset with high simulation degree, automatically constructing a sample dataset containing screen distortion defects and a corresponding damaged area mask sequence, effectively solving the problem of data scarcity, and providing a large-scale and high-quality training sample for model training; on the other hand, at the model level, a multi-task defect detection model based on deep learning is proposed, which can accurately detect screen distortion in video frames and precisely locate the defective area. The target defect map of the screen distortion area at the pixel level is output with high precision by the defect locator, realizing the accurate detection and identification of screen distortion defects; and, based on the feature extractor and the defect locator, a screen distortion classifier is introduced, and the detection accuracy of screen distortion is effectively improved through feature interaction and fusion. At the same time, false alarms with low confidence are excluded, further improving the detection accuracy and robustness of the model. It can perform pixel-level recognition on video frames containing screen distortion, and output the target defect map, defect confidence, and defect severity. It can also provide efficient and accurate video screen distortion defect detection and evaluation for the video quality control system, provide precise guidance for subsequent video repair solutions, and effectively support services such as ultra-high-definition video selection.

[0165] According to an embodiment of the present disclosure, on the data side, by formulating detailed defect definition and annotation standards, the formation principles of global defects such as ghosting and streaking are deeply studied. Then, based on the formation principles of ghosting and streaking defects, a set of automated paired defect frame / target defect map generation processes are developed by introducing corresponding distortions into normal video frames, that is, a set of defect dataset automatic generation processes are designed, which can construct a large-scale global defect dataset with pixel-level labels (target defect maps) and whole-image defect category labels, making it possible to train a global defect detection model driven by big data; based on the defect dataset with pixel-level target defect map labels built on the data side, on the algorithm side, an end-to-end defect detection model based on a multi-task learning architecture is proposed to accurately identify ghosting and streaking defects in videos. Among them, the fine target defect map output is realized through the defect locator, and the defective area at the pixel level can be output. Through the defect classifier, the features in all aspects are interacted and fused, which is beneficial to accurately filtering out video frames with low confidence, significantly reducing the defect misdetection rate, and being able to accurately detect multi-dimensional information such as defect confidence, pixel-level target defect map, single-frame / image overall defect severity, and defective video segments, realizing business implementation, and effectively supporting services such as video medium delivery, ultra-high-definition video selection, and old film restoration.

[0166] According to an embodiment of the present disclosure, a data construction scheme for simulating real film damage is designed based on the principle of dirt spots and scratch defects on the data side, and a sample data set for defect simulation is constructed. On the model side, a feature extractor based on the Transformer module + multi-scale feature fusion is used to accurately segment the defect area. By training a deep learning model, the model gradually masters various defect features during the learning process and can perform pixel-level accurate recognition on frames or images containing defects. It can provide efficient and accurate dirt spot and scratch defect detection and evaluation for the video quality control system, and at the same time support the old film restoration service to assist related algorithms in improving the video restoration effect.

[0167] According to an embodiment of the present disclosure, an automated construction process for a defect data set simulating real film damage is proposed to create a sample data set containing video frames with dirt spot and scratch defects for simulating real old film damage, to generate defect frames or target defect maps containing dirt spot and scratch defects, successfully introducing corresponding defects into normal video frames, greatly enriching the source of training data and solving the problem of dirt spot and scratch detection of old films in the industry; on the model side, a defect detection model based on the Transformer module + multi-scale feature fusion is proposed, which can perform pixel-level recognition on video frames containing defects to accurately segment the defect area. By training the model using the constructed sample data set, the model gradually masters various defect features during the learning process and can perform pixel-level analysis and recognition on frames or images containing defects, realizing pixel-level recognition of video frames containing dirt spot and scratch defects; a set of defect severity evaluation criteria is also proposed according to actual business requirements. For each detected dirt spot or scratch defect, its characteristic parameters such as position, quantity, area, length, width, and gray level difference can be calculated, and the number of defects is combined as a standard for measuring the severity of damage. The defects are classified into levels such as minor, severe, and severe according to the degree of affecting the visual effect. Finally, the defect characteristic parameters and the defect severity can be output to provide accurate guidance for subsequent old film restoration algorithms, effectively supporting services such as ultra-high-definition video film selection and old film restoration.

[0168] Figure 9 The block diagram of a video defect detection device according to an embodiment of the present disclosure is shown, as Figure 9 shown, the device includes:

[0169] An acquisition module 901, configured to acquire video frames of a video to be detected;

[0170] A feature extraction module 902, configured to extract a shallow feature map and a deep feature map of the video frame;

[0171] A defect map determination module 903, configured to determine a defect map corresponding to the video frame according to the shallow feature map and the deep feature map, where the defect map represents a defect area where a defect is located in the predicted video frame;

[0172] A result determination module 904, configured to determine a defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map, and the defect map, where the defect detection result is used to indicate whether there is a defect in the video frame.

[0173] In a possible implementation manner, the extracting the shallow feature map and the deep feature map of the video frame includes: dividing the video frame into a plurality of image blocks; inputting the plurality of image blocks into a converter module to obtain multi-scale feature maps extracted by N converter units in the converter module, where a first-scale feature map extracted by a first converter unit is used as the shallow feature map, and N is a positive integer; performing a fusion process on second-scale feature maps to N-scale feature maps output by a second converter unit to an Nth converter unit in the converter module to obtain a first fusion feature map; and determining the deep feature map according to the first fusion feature map.

[0174] In a possible implementation manner, the determining the deep feature map according to the first fusion feature map includes: performing a dimensionality reduction process on the first fusion feature map to obtain a second fusion feature map; respectively performing self-attention extraction processes on the second fusion feature map by using a plurality of window attention mechanisms to obtain a plurality of self-attention feature maps, where window sizes used for extracting self-attention in different window attention mechanisms are different; performing a pooling process on the second fusion feature map to obtain a pooled feature map; performing channel concatenation on the second fusion feature map, the plurality of self-attention feature maps, and the pooled feature map to obtain a first concatenated feature map; and performing a dimensionality reduction process on the first concatenated feature map to obtain the deep feature map.

[0175] In a possible implementation manner, the determining the defect map corresponding to the video frame according to the shallow feature map and the deep feature map includes: performing a dimensionality reduction process on the shallow feature map to obtain a first shallow feature map; performing an upsampling process on the deep feature map based on the resolution of the first shallow feature map to obtain a first deep feature map, where the first deep feature map has the same resolution as the first shallow feature map; performing channel concatenation on the first shallow feature map and the first deep feature map to obtain a second concatenated feature map; performing a dimensionality reduction process on the second concatenated feature map to obtain a defect map, where the defect map represents a defect area where a defect in the video frame is predicted; and performing an upsampling process on the defect map based on the original resolution of the video frame to obtain a target defect map that matches the original resolution of the video frame.

[0176] In a possible implementation, determining the defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map, and the defect map includes: extracting features from the shallow feature map to obtain a second deep feature map, and performing dimensionality reduction processing on the second deep feature map to obtain a first intermediate feature map; based on the resolution of the shallow feature map, performing upsampling processing on the first intermediate feature map to obtain a second intermediate feature map, where the resolution of the second intermediate feature map is the same as that of the shallow feature map; respectively performing dimensionality reduction processing on the shallow feature map, the first deep feature map, and the defect map to obtain a second shallow feature map, a third deep feature map, and a first defect map; performing channel splicing on the second intermediate feature map, the second shallow feature map, the third deep feature map, and the first defect map to obtain a third spliced feature map; performing dimensionality reduction processing on the third spliced feature map to obtain a defect confidence level, where the defect detection result includes the defect confidence level, and the defect confidence level represents the probability that the video frame has a defect.

[0177] In a possible implementation, the apparatus further includes: a severity assessment module, configured to determine the defect severity of the video frame according to at least one of the position, size, and quantity of the defect regions represented by the target defect map when the defect detection result indicates that the video frame has a defect.

[0178] In a possible implementation, the apparatus further includes: a defect segment determination module, configured to determine a video segment in which defects continuously exist in the video according to the defect detection results of multiple video frames in the video.

[0179] In a possible implementation, the apparatus further includes: a repair module, configured to repair the defective video frame according to the target defect map of the defective video frame to obtain a repaired video frame.

[0180] In a possible implementation, for the apparatus using the defect detection model, the defect detection model includes a feature extractor, a defect locator, and a defect classifier; wherein, the feature extraction module is configured to extract the shallow feature map and the deep feature map of the video frame by using the feature extractor, the defect map determination module is configured to determine the defect map corresponding to the video frame according to the shallow feature map and the deep feature map by using the defect locator, and the result determination module is configured to determine the defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map, and the defect map by using the defect classifier.

[0181] In a possible implementation, the defect detection model is trained using a sample data set; wherein, the sample data set is obtained through the following process: obtaining a plurality of original video samples; respectively performing random damage processing on the plurality of original video samples to obtain damaged video samples corresponding to the respective original video samples, wherein each damaged video sample includes at least one damaged video frame with a defect; according to the differences between the respective original video samples and the damaged video samples corresponding thereto, obtaining the damaged video frames in each damaged video sample and the mask graphs corresponding to the damaged video frames, where the mask graphs are used to indicate the defect regions where the defects are located in the damaged video frames; wherein, the sample data set includes a plurality of damaged video samples corresponding to the plurality of original video samples, the mask graphs of the damaged video frames in each damaged video sample, and the class labels of each video frame in each damaged video sample, and the class labels are used to indicate whether the video frame is a damaged video frame or an original video frame that has not been damaged.

[0182] In a possible implementation, the defects in the video frames include mosaic defects caused by partial data damage during video transmission; wherein, when the defect detection model is used to detect the mosaic defects, the process of respectively performing random damage processing on the plurality of original video samples to obtain damaged video samples corresponding to the respective original video samples includes: based on a preset damage model, performing data damage processing on the original video streams of the plurality of original video samples to obtain damaged video streams corresponding to the respective original video samples; using a video parser to parse the damaged video streams corresponding to the respective original video samples to obtain damaged video samples corresponding to the respective original video samples; wherein, the damage model is used to indicate the damage ratio, the damage position, and the damage length, the damage ratio represents the ratio of the damaged video frames to the original video sample, the damage position represents the position of the damaged video frame in the original video sample, and the damage length represents the length of the defect region in the damaged video frame.

[0183] In a possible implementation, the defects in the video frames include ghosting defects caused by adjacent frame overlap during video format conversion; wherein, when the defect detection model is used to detect the ghosting defects, the process of respectively performing random damage processing on the plurality of original video samples to obtain damaged video samples corresponding to the respective original video samples includes: for any original video sample, estimating the motion regions of each video frame in the original video sample, where the objects in the motion regions have a motion state, and determining at least one candidate ghosting frame according to the sizes of the motion regions of each video frame in the original video sample; overlapping each candidate ghosting frame in the original video sample with the video frames adjacent to each candidate ghosting frame to obtain a damaged video sample.

[0184] In a possible implementation, the defects in the video frame include the moiré defects caused by interpolating the odd and even lines of different frames of the original frame rate video into the same frame of the new frame rate video during the frame rate conversion of the interlaced video frame. The moiré defects have a horizontal comb shape. Wherein, when the defect detection model is used to detect the moiré defects, the randomly damaging the multiple original video samples respectively to obtain the damaged video samples corresponding to the original video samples includes: referring to the generation principle of the moiré defects, randomly introducing the moiré defects into at least one video frame of each original video sample to obtain the damaged video samples corresponding to the original video samples.

[0185] In a possible implementation, the defects in the video frame include physical defects caused by physical damage to the video film; the physical defects include linear traces caused by physical scratches in the video film, and / or irregular spots caused by stains or impurities in non-image structures on the chemical coating of the video film. Wherein, when the defect detection model is used to detect the physical defects, the randomly damaging the multiple original video samples respectively to obtain the damaged video samples corresponding to the original video samples includes: acquiring multiple film scan images and marking the region boundaries of the physical defects in each film scan image, where the film scan image is an image obtained by scanning the video film with the physical damage; dividing each film scan image into multiple pixel regions of a specified size, and respectively counting the number of defects of each physical defect appearing in each pixel region based on the region boundaries of the physical defects marked in each film scan image; fitting a gamma distribution model corresponding to each physical defect by performing maximum likelihood estimation on the number of defects of each physical defect appearing in the multiple pixel regions, where the gamma distribution model is used to characterize the quantity distribution characteristics of each physical defect; for any original video sample, sampling the target number of defects of each physical defect to be added to the video frames in the original video sample based on the gamma distribution model corresponding to each physical defect; based on the target number of defects of each physical defect, extracting the position information matching the target number of defects from a preset spatial noise distribution for each physical defect, and extracting the rotation angle matching the target number of defects from a preset rotation angle range for each physical defect, where the position information is used to indicate the pixel position of the physical defect, and the rotation angle is used to indicate the direction of the physical defect; based on the position information and rotation angle matching the target number of defects of each physical defect, fusing the physical defects matching the target number of defects of each physical defect into at least one video frame of the original video sample to obtain the damaged video sample.

[0186] According to the device of an embodiment of the present disclosure, a pixel-level defect area can be predicted based on the shallow feature map and the deep feature map of a video frame. This process is equivalent to a preliminary screening of the video frames with defects. Then, by using the shallow feature map, the deep feature map and the defect map, the defect detection result can be determined more precisely, that is, it can be more precisely judged whether there are defects in the video frame. In particular, by introducing the defect map, it is equivalent to introducing an attention mechanism to pay more attention to the areas with defects, which can effectively improve the accuracy of the defect detection result, thereby helping to reduce the false detection rate of defect detection.

[0187] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0188] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0189] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above methods when executing the instructions stored in the memory.

[0190] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.

[0191] Figure 10 The block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 10 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules corresponding to a set of instructions each. In addition, the processing component 1922 is configured to execute instructions to execute the above methods.

[0192] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0193] In an exemplary embodiment, there is also provided a non-transitory computer-readable storage medium, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.

[0194] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0195] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punch card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0196] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0197] The computer program instructions for carrying out operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0198] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0199] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions which implement various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0200] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, whereby the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0201] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.

[0202] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A video defect detection method, characterized in that, Including: Obtain video frames of the video to be detected; Extract the shallow feature map and the deep feature map of the video frame; Determine a defect map corresponding to the video frame according to the shallow feature map and the deep feature map, where the defect map characterizes the defect area where the defect is located in the predicted video frame; Determine a defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map and the defect map, where the defect detection result is used to indicate whether there is a defect in the video frame; Wherein, the extracting the shallow feature map and the deep feature map of the video frame includes: Divide the video frame into multiple image blocks; Input the multiple image blocks into a converter module to obtain multi-scale feature maps extracted by N converter units in the converter module, where the first-scale feature map extracted by the first converter unit is used as the shallow feature map, and N is a positive integer; Perform a fusion process on the second-scale feature map to the N-scale feature map output by the second converter unit to the Nth converter unit in the converter module to obtain a first fusion feature map; Determine the deep feature map according to the first fusion feature map.

2. The method according to claim 1, characterized in that, The determining the deep feature map according to the first fusion feature map includes: Perform a dimensionality reduction process on the first fusion feature map to obtain a second fusion feature map; Use multiple window attention mechanisms to perform self-attention extraction processes on the second fusion feature map respectively to obtain multiple self-attention feature maps, where the window sizes used for extracting self-attention in different window attention mechanisms are different; Perform a pooling process on the second fusion feature map to obtain a pooled feature map; Perform channel splicing on the second fusion feature map, the multiple self-attention feature maps and the pooled feature map to obtain a first spliced feature map; Perform a dimensionality reduction process on the first spliced feature map to obtain the deep feature map.

3. The method according to claim 1, wherein The determining the defect map corresponding to the video frame according to the shallow feature map and the deep feature map includes: Perform a dimensionality reduction process on the shallow feature map to obtain a first shallow feature map; Perform an upsampling process on the deep feature map based on the resolution of the first shallow feature map to obtain a first deep feature map, where the resolution of the first deep feature map is the same as that of the first shallow feature map; Perform channel splicing on the first shallow feature map and the first deep feature map to obtain a second spliced feature map; Perform a dimensionality reduction process on the second spliced feature map to obtain the defect map; Perform an upsampling process on the defect map based on the original resolution of the video frame to obtain a target defect map that matches the original resolution of the video frame.

4. The method according to claim 3, wherein The determining the defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map and the defect map includes: Extract features from the shallow feature map to obtain a second deep feature map, and perform a dimensionality reduction process on the second deep feature map to obtain a first intermediate feature map; Based on the resolution of the shallow feature map, perform upsampling on the first intermediate feature map to obtain a second intermediate feature map, where the resolution of the second intermediate feature map is the same as that of the shallow feature map; Perform dimensionality reduction processing on the shallow feature map, the first deep feature map, and the defect map respectively to obtain a second shallow feature map, a third deep feature map, and a first defect map; Perform channel concatenation on the second intermediate feature map, the second shallow feature map, the third deep feature map, and the first defect map to obtain a third concatenated feature map; Perform dimensionality reduction processing on the third concatenated feature map to obtain a defect confidence level. The defect detection result includes the defect confidence level, and the defect confidence level represents the probability that the video frame has a defect.

5. The method according to claim 3, characterized in that, The method further includes: When the defect detection result indicates that the video frame has a defect, determine the severity of the defect of the video frame according to at least one of the position, size, and quantity of the defect area represented by the target defect map.

6. The method according to claim 1, characterized in that The method further includes: Determine the video segments with continuously existing defects in the video according to the defect detection results of multiple video frames in the video.

7. The method according to claim 3, characterized in that, The method further includes: Repair the video frame with a defect according to the target defect map of the video frame with a defect to obtain a repaired video frame.

8. The method according to any one of claims 1 to 4, characterized in that Implement the method according to any one of claims 1 to 4 by using a defect detection model, where the defect detection model includes a feature extractor, a defect locator, and a defect classifier; Among them, the feature extractor is used to extract the shallow feature map and the deep feature map of the video frame, the defect locator is used to determine the defect map corresponding to the video frame according to the shallow feature map and the deep feature map, and the defect classifier is used to determine the defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map, and the defect map.

9. The method according to claim 8, wherein The defect detection model is trained by using a sample data set; among them, the sample data set is obtained through the following process: Obtain a plurality of original video samples; Perform random damage processing on the plurality of original video samples respectively to obtain damaged video samples corresponding to the respective original video samples, where each damaged video sample contains at least one damaged video frame with a defect; According to the difference between each original video sample and its corresponding damaged video sample, obtain the damaged video frames in each damaged video sample and the mask map corresponding to the damaged video frame, where the mask map is used to indicate the defect area where the defect is located in the damaged video frame; Among them, the sample data set includes the damaged video samples corresponding to the plurality of original video samples, the mask maps of the damaged video frames in each damaged video sample, and the class labels of each video frame in each damaged video sample, where the class labels are used to indicate whether the video frame is a damaged video frame or an original video frame that has not been damaged.

10. The method according to claim 9, wherein The defect in the video frame includes a mosaic defect caused by partial data damage during the video transmission process; Wherein, when the defect detection model is used to detect the screen freeze defect, the random damage processing of the multiple original video samples respectively to obtain the damaged video samples corresponding to the respective original video samples includes: Performing data damage processing on the original video streams of the multiple original video samples based on a preset damage model to obtain the damaged video streams corresponding to the respective original video samples; Using a video parser to parse the damaged video streams corresponding to the respective original video samples to obtain the damaged video samples corresponding to the respective original video samples; Wherein, the damage model is used to indicate a damage ratio, a damage position, and a damage length. The damage ratio represents the proportion of damaged video frames in the original video sample. The damage position represents the position of the damaged video frame in the original video sample. The damage length represents the length of the defect area in the damaged video frame.

11. The method according to claim 9, wherein The defects in the video frames include ghosting defects caused by overlapping adjacent frames during the video format conversion process; Wherein, when the defect detection model is used to detect the ghosting defect, the random damage processing of the multiple original video samples respectively to obtain the damaged video samples corresponding to the respective original video samples includes: For any original video sample, estimating the motion areas of the video frames in the original video sample, where the objects in the motion areas have a motion state, and determining at least one candidate ghosting frame according to the sizes of the motion areas of the video frames in the original video sample; Overlapping each candidate ghosting frame in the original video sample with the video frames adjacent to each candidate ghosting frame to obtain the damaged video sample.

12. The method according to claim 9, wherein The defects in the video frames include streaking defects caused by interpolating the odd and even lines of different frames of the original frame rate video into the same frame of the new frame rate video during the frame rate conversion process of interlaced video frames; Wherein, when the defect detection model is used to detect the streaking defect, the random damage processing of the multiple original video samples respectively to obtain the damaged video samples corresponding to the respective original video samples includes: Referring to the generation principle of the streaking defect, randomly introducing the streaking defect into at least one video frame in each of the original video samples to obtain the damaged video samples corresponding to the respective original video samples.

13. The method according to claim 9, characterized in that, The defects in the video frames include physical defects caused by physical damage to the video film; the physical defects include linear traces caused by physical scratches in the video film, and / or irregular spots caused by stains or impurities in non-image structures on the chemical coating of the video film; Wherein, when the defect detection model is used to detect the physical defect, the random damage processing of the multiple original video samples respectively to obtain the damaged video samples corresponding to the respective original video samples includes: Obtaining multiple film scan images and marking the region boundaries of the physical defects in each film scan image, where the film scan images are images obtained by scanning the video film with the physical damage. Segment the scanned image of each film into multiple pixel regions of a specified size, and based on the region boundaries of the physical defects marked in the scanned image of each film, respectively count the number of defects of each physical defect that appear in each pixel region; By performing maximum likelihood estimation on the number of defects of each physical defect that appear in multiple pixel regions, fit a gamma distribution model corresponding to each physical defect, and the gamma distribution model is used to characterize the quantity distribution characteristics of each physical defect; For any original video sample, based on the gamma distribution model corresponding to each physical defect, sample the target number of defects of each physical defect to be added to the video frames in the original video sample; Based on the target number of defects of each physical defect, extract position information matching the target number of defects from a preset spatial noise distribution for each physical defect, and extract a rotation angle matching the target number of defects from a preset range of rotation angles for each physical defect. The position information is used to indicate the pixel position of the physical defect, and the rotation angle is used to indicate the direction of the physical defect; Based on the position information and rotation angle matching the target number of defects of each physical defect, fuse the physical defects matching the target number of defects of each physical defect into at least one video frame of the original video sample to obtain a damaged video sample.

14. A video defect detection device, characterized in that, Including: An acquisition module, configured to acquire video frames of a video to be detected; A feature extraction module, configured to extract a shallow feature map and a deep feature map of the video frame; A defect map determination module, configured to determine a defect map corresponding to the video frame according to the shallow feature map and the deep feature map, and the defect map characterizes the defect region where the defect is predicted to be in the video frame; A result determination module, configured to determine a defect detection result corresponding to the video frame according to the shallow feature map, the deep feature map and the defect map, and the defect detection result is used to indicate whether there is a defect in the video frame; Wherein, the extraction of the shallow feature map and the deep feature map of the video frame includes: Divide the video frame into multiple image blocks; Input the multiple image blocks into a converter module to obtain multi-scale feature maps extracted by N converter units in the converter module. Among them, the first-scale feature map extracted by the first converter unit is used as the shallow feature map, and N is a positive integer; Perform fusion processing on the second-scale feature map to the N-scale feature map output by the second converter unit to the Nth converter unit in the converter module to obtain a first fusion feature map; Determine the deep feature map according to the first fusion feature map.

15. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the method according to any one of claims 1 to 13 when executing the instructions stored in the memory.

16. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by the processor, implement the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Method for identifying blurred screen, device thereof and equipment and storage medium

    CN113177529A

  • Video restoration method and system

    CN117455812A