Automatic restoration method and device for film picture flicker based on artificial intelligence

An AI-driven method for film flicker removal improves film processing accuracy by identifying and correcting exposure anomalies in film frames, enhancing the viewing experience.

CN120321349AActive Publication Date: 2025-07-15CFGDC (BEIJING) TECHNOLOGY CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510804770.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The flashing phenomenon caused by rapid and irregular fluctuations in the video screen seriously affects the user's viewing experience.

Method used

By obtaining video clips, the local spatial information capacity of the video frame is determined, the exposure abnormal frame is identified, and the pixel values of adjacent frames are used for bias correction processing, and the flickering of the video screen is automatically fixed.

Benefits of technology

It improves the accuracy of video processing and user viewing experience, ensuring the stability of video picture quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321349A_ABST
    Figure CN120321349A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an automatic restoration method and device for film picture flicker based on artificial intelligence. The method comprises the following steps: acquiring a to-be-processed film clip; wherein the to-be-processed film clips are film clips in the same film scene; the to-be-processed film clip comprises a plurality of film frames; determining the local space information capacity of the film frames, and determining exposure abnormal frames from the plurality of film frames according to the local space information capacity; wherein the local space information capacity is used for representing regional energy distribution of the film frame; and correcting the exposure abnormal frame according to the pixel value of the adjacent frame of the exposure abnormal frame to obtain the to-be-processed film clip after the flicker removal processing. The method is used for achieving the effect of improving the film watching experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to an automatic repair method and device for film frame flicker based on artificial intelligence. Background Art

[0002] Film flicker refers to the phenomenon where the brightness or color in a film frame fluctuates rapidly and irregularly, resulting in visual alternation of light and dark or color distortion. This phenomenon seriously affects the user's film viewing experience.

[0003] Therefore, there is an urgent need for a method to remove flicker from films to improve the user's film viewing experience. Summary of the Invention

[0004] Embodiments of this application provide an automatic repair method and device for film frame flicker based on artificial intelligence, so as to achieve the effect of improving the user's film viewing experience.

[0005] In a first aspect, embodiments of this application provide an automatic repair method for film frame flicker based on artificial intelligence, including:

[0006] Obtain a film clip to be processed; wherein, the film clip to be processed is a clip of film frames in the same film scene; the film clip to be processed includes multiple film frames;

[0007] Determine the local spatial information capacity of the film frames, and determine exposure abnormal frames from the multiple film frames according to the local spatial information capacity; wherein, the local spatial information capacity is used to characterize the regional energy distribution of the film frames;

[0008] Perform deviation correction processing on the exposure abnormal frames according to the pixel values of the adjacent frames of the exposure abnormal frames to obtain the film clip to be processed after flicker removal processing.

[0009] In a possible implementation, the local space information capacity includes at least one local space information sub-capacity; wherein, the local space information sub-capacity is used to characterize the regional energy distribution of pixel blocks in the video frame; determining an exposure abnormal frame from the multiple video frames according to the local space information capacity includes: calculating the probability density of the local space information sub-capacity of the video frame according to at least one of the local space information sub-capacities included in the local space information capacity; wherein, the probability density of the local space information sub-capacity indicates the proportion of the number of pixel blocks corresponding to each local space information sub-capacity to the total number of pixel blocks in the video frame; constructing a heat map histogram of the video frame according to at least one of the local space information sub-capacities included in the local space information capacity; determining whether the video frame is an exposure abnormal frame according to the proportion of the number of pixel blocks corresponding to each local space information sub-capacity in the video frame to the total number of pixel blocks in the video frame, and / or, the heat map histogram of the video frame.

[0010] In a possible implementation, determining whether the video frame is an exposure abnormal frame according to the proportion of the number of pixel blocks corresponding to each local space information sub-capacity in the video frame to the total number of pixel blocks in the video frame, and / or, the heat map histogram of the video frame includes: if the proportion of the number of pixel blocks corresponding to the local space information sub-capacity lower than the local space information sub-capacity threshold in the video frame to the total number of pixel blocks in the video frame exceeds the proportion threshold, and the pixel difference value between the video frame and the adjacent frame exceeds the difference value threshold; and / or, the slope mutation value of the heat map histogram of the video frame exceeds the mutation threshold, and the pixel difference value between the video frame and the adjacent frame exceeds the difference value threshold, then it is determined that the video frame is an exposure abnormal frame.

[0011] In a possible implementation, performing deviation correction processing on the exposure abnormal frame according to the pixel values of the adjacent frames of the exposure abnormal frame to obtain a to-be-processed video segment after de-flickering processing includes: performing regression calculation processing on the pixel values of the adjacent frames of the exposure abnormal frame to obtain the target pixel value of the exposure abnormal frame; performing deviation correction processing on the pixel value of the exposure abnormal frame according to the target pixel value of the exposure abnormal frame to obtain the exposure abnormal frame after deviation correction processing; if the structural similarity index and peak signal-to-noise ratio between the exposure abnormal frame after deviation correction processing and the exposure abnormal frame meet the preset conditions, then replacing the exposure abnormal frame with the exposure abnormal frame after deviation correction processing; otherwise, increasing the number of adjacent frames of the exposure abnormal frame, and repeating the process of regression calculation processing and deviation correction processing.

[0012] In a possible implementation manner, determining the local spatial information capacity of the video frame includes: obtaining the noise variance of the video frame, and obtaining the energy information of each pixel block in the video frame; determining the local spatial information capacity according to the noise variance and the energy information of each pixel block.

[0013] In a possible implementation manner, obtaining a video clip to be processed includes: obtaining a complete video; where the complete video is a video picture including at least one video clip to be processed; based on deep feature extraction technology and shallow feature extraction technology, determining the similarity between each adjacent video frame in the complete video; dividing the complete video according to the similarity between each adjacent video frame in the complete video to obtain at least one video clip to be processed.

[0014] In a possible implementation manner, dividing the complete video according to the similarity between each adjacent video frame in the complete video to obtain at least one video clip to be processed includes: if the similarity between two adjacent video frames in the complete video is less than a first similarity threshold, then dividing the two video frames in the adjacent frames into video clips under different video scenes to obtain the video clip to be processed; if the similarity between multiple consecutive adjacent video frames in the complete video is greater than or equal to the first similarity threshold and less than a second similarity threshold, then performing scene division processing on the video frames in the multiple consecutive adjacent video frames based on dynamic programming technology to obtain the video clip to be processed.

[0015] In a second aspect, an embodiment of the present application provides an automatic repair device for video picture flickering based on artificial intelligence, including:

[0016] An acquisition module, configured to acquire a video clip to be processed; where the video clip to be processed is a clip of a video picture under the same video scene; the video clip to be processed includes multiple video frames;

[0017] A determination module, configured to determine the local spatial information capacity of the video frame, and determine an exposure abnormal frame from the multiple video frames according to the local spatial information capacity; where the local spatial information capacity is used to characterize the regional energy distribution of the video frame;

[0018] A processing module, configured to perform deviation correction processing on the exposure abnormal frame according to the pixel values of the adjacent frames of the exposure abnormal frame to obtain a video clip to be processed after de-flickering processing.

[0019] In a possible implementation, the local space information capacity includes at least one local space information sub-capacity; wherein, the local space information sub-capacity is used to characterize the regional energy distribution of pixel blocks in the video frame; the determining module is specifically configured to calculate the probability density of the local space information sub-capacity of the video frame according to at least one of the local space information sub-capacities included in the local space information capacity; wherein, the probability density of the local space information sub-capacity indicates the proportion of the number of pixel blocks corresponding to each local space information sub-capacity to the number of all pixel blocks in the video frame; construct a heat histogram of the video frame according to at least one of the local space information sub-capacities included in the local space information capacity; determine whether the video frame is an exposure abnormal frame according to the proportion of the number of pixel blocks corresponding to each local space information sub-capacity in the video frame to the number of all pixel blocks in the video frame, and / or the heat histogram of the video frame.

[0020] In a possible implementation, the determining module is further specifically configured to determine that the video frame is an exposure abnormal frame if the proportion of the number of pixel blocks corresponding to the local space information sub-capacity lower than the local space information sub-capacity threshold in the video frame to the number of all pixel blocks in the video frame exceeds the proportion threshold, and the pixel difference value between the video frame and the adjacent frame exceeds the difference value threshold; and / or, the slope mutation value of the heat histogram of the video frame exceeds the mutation threshold, and the pixel difference value between the video frame and the adjacent frame exceeds the difference value threshold.

[0021] In a possible implementation, the processing module is specifically configured to perform regression calculation processing on the pixel values of the adjacent frames of the exposure abnormal frame to obtain the target pixel value of the exposure abnormal frame; perform deviation correction processing on the pixel values of the exposure abnormal frame according to the target pixel value of the exposure abnormal frame to obtain the exposure abnormal frame after deviation correction processing; if the structural similarity index and peak signal-to-noise ratio between the exposure abnormal frame after deviation correction processing and the exposure abnormal frame meet the preset conditions, replace the exposure abnormal frame with the exposure abnormal frame after deviation correction processing; otherwise, increase the number of adjacent frames of the exposure abnormal frame, and repeat the process of regression calculation processing and deviation correction processing.

[0022] In a possible implementation, the determining module is further specifically configured to obtain the noise variance of the video frame, and obtain the energy information of each pixel block in the video frame; determine the local space information capacity according to the noise variance and the energy information of each pixel block.

[0023] In a possible implementation manner, the obtaining module is specifically configured to obtain a complete film, where the complete film is a film screen including at least one of the to-be-processed film segments; determine the similarity between each adjacent film frame in the complete film based on a deep feature extraction technique and a shallow feature extraction technique; and perform a partitioning process on the complete film according to the similarity between each adjacent film frame in the complete film to obtain at least one of the to-be-processed film segments.

[0024] In a possible implementation manner, the obtaining module is further specifically configured to, if the similarity between an adjacent film frame in the complete film is less than a first similarity threshold, divide the two film frames in the adjacent frames into film segments in different film scenes to obtain the to-be-processed film segment; if the similarity between multiple consecutive adjacent film frames in the complete film is greater than or equal to the first similarity threshold and less than a second similarity threshold, perform a scene partitioning process on the film frames in the multiple consecutive adjacent film frames based on a dynamic programming technique to obtain the to-be-processed film segment.

[0025] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0026] The memory stores computer-executable instructions;

[0027] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0028] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0029] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above first aspect and / or various possible implementation manners of the first aspect.

[0030] The automatic repair method and device for film screen flicker based on artificial intelligence provided by the embodiments of the present application obtain to-be-processed film segments, determine the local spatial information capacity of film frames, and determine exposure-abnormal frames from multiple film frames according to the local spatial information capacity, and perform a rectification process on the exposure-abnormal frames according to the pixel values of the adjacent frames of the exposure-abnormal frames to obtain the to-be-processed film segments after flicker removal processing. Among them, through the local spatial information capacity, the exposure-abnormal frames can be accurately determined from the to-be-processed film segments, further improving the accuracy of film processing, thereby improving the user's film viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings herein are incorporated into and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0032] Figure 1 Schematic flowchart of the automatic repair method for film frame flickering based on artificial intelligence provided by the present application Figure 1 ;

[0033] Figure 2 Schematic flowchart of the automatic repair method for film frame flickering based on artificial intelligence provided by the present application Figure 2 ;

[0034] Figure 3 Schematic flowchart of the automatic repair method for film frame flickering based on artificial intelligence provided by the present application Figure 3 ;

[0035] Figure 4 Schematic flowchart of the automatic repair method for film frame flickering based on artificial intelligence provided by the present application Figure 4 ;

[0036] Figure 5 Schematic flowchart of the automatic repair method for film frame flickering based on artificial intelligence provided by the present application Figure 5 ;

[0037] Figure 6 Schematic structural diagram of the automatic repair device for film frame flickering based on artificial intelligence provided by the present application;

[0038] Figure 7 Schematic structural diagram of the electronic device provided by the present application.

[0039] Through the above drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0041] In the prior art, there is a technical problem that the user's film viewing experience is poor.

[0042] The automatic repair method for film frame flickering based on artificial intelligence provided by this application obtains the film clip to be processed, determines the local spatial information capacity of the film frames, and determines the frames with abnormal exposure from multiple film frames according to the local spatial information capacity. The frames with abnormal exposure are corrected according to the pixel values of the adjacent frames of the frames with abnormal exposure, and the film clip to be processed after de-flickering is obtained. Among them, through the local spatial information capacity, the frames with abnormal exposure can be accurately determined from the film clip to be processed, further improving the accuracy of film processing, thereby improving the user's film viewing experience.

[0043] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0044] Figure 1 Flow schematic of the automatic repair method for film frame flickering based on artificial intelligence provided by this application Figure 1 , as Figure 1 shown, this method includes:

[0045] Step S101, obtain the film clip to be processed.

[0046] Specifically, during the process of de-flickering the film, if the film to be processed is a clip of film frames in different film scenes, there will be a problem of inaccurate determination of abnormal frames due to different film scenes to which each film frame belongs, further resulting in inaccurate de-flickering processing of the film. Therefore, the film clip to be processed can be obtained. Among them, the film clip to be processed is a clip of film frames in the same film scene. The film clip to be processed includes multiple film frames.

[0047] Specifically, this application does not limit the process of obtaining the film clip to be processed. Optionally, a complete film can be obtained; among them, the complete film is a film image including at least one film clip to be processed; based on deep feature extraction technology and shallow feature extraction technology, the similarity between each adjacent film frame in the complete film is determined; according to the similarity between each adjacent film frame in the complete film, the complete film is divided to obtain at least one film clip to be processed.

[0048] Step S102, determine the local spatial information capacity of the film frames, and determine the frames with abnormal exposure from multiple film frames according to the local spatial information capacity.

[0049] Specifically, after obtaining the video clip to be processed, the local spatial information capacity of each video frame in the video clip to be processed can be determined.

[0050] In an optical imaging system, objects outside the depth of field range will produce defocus blur. When an object is far from the focal plane, its spatial high-frequency components, that is, detail information, will decay non-linearly with the increase of the defocus distance, showing a change in the band-limited characteristic of the spatial frequency. A more accurate description should be "Local Spatial Information Capacity" (LSIC for short). Specifically, for the near-view region, each pixel corresponds to a smaller physical scale and has a higher spatial sampling density, that is, high local spatial information capacity. For the far-view region, each pixel covers a larger physical scale, resulting in spatial aliasing, that is, low local spatial information capacity.

[0051] Among them, the local spatial information capacity is used to characterize the regional energy distribution of the video frame. Specifically, based on the above description of the local spatial information capacity, the stronger the regional energy distribution of the video frame characterized by the local spatial information capacity, the higher the local spatial information capacity. On the contrary, the weaker the regional energy distribution of the video frame characterized by the local spatial information capacity, the lower the local spatial information capacity.

[0052] Specifically, the present application does not limit the process of determining the local spatial information capacity of the video frame. Optionally, the noise variance of the video frame can be obtained, and the energy information of each pixel block in the video frame can be obtained; according to the noise variance and the energy information of each pixel block, the local spatial information capacity can be determined.

[0053] Specifically, after determining the local spatial information capacity of each video frame in the video clip to be processed, the exposure abnormal frame can be determined from multiple video frames according to the local spatial information capacity.

[0054] Among them, the exposure abnormal frame is the video frame to be corrected among multiple video frames of the video clip to be processed during the de-flickering process. Among them, the present application does not limit the number of determined exposure abnormal frames.

[0055] Specifically, this application does not limit the process of determining exposure anomaly frames from multiple video frames according to the local space information capacity. Optionally, the local space information capacity includes at least one local space information sub-capacity; wherein, the local space information sub-capacity is used to characterize the regional energy distribution of pixel blocks in the video frame; according to at least one local space information sub-capacity included in the local space information capacity, calculate the probability density of the local space information sub-capacity of the video frame; wherein, the probability density of the local space information sub-capacity indicates the proportion of the number of pixel blocks corresponding to each local space information sub-capacity to the total number of pixel blocks in the video frame; according to at least one local space information sub-capacity included in the local space information capacity, construct a heat map histogram of the video frame; determine whether the video frame is an exposure anomaly frame according to the proportion of the number of pixel blocks corresponding to each local space information sub-capacity in the video frame to the total number of pixel blocks in the video frame, and / or, the heat map histogram of the video frame.

[0056] Step S103: Perform deviation correction processing on the exposure anomaly frame according to the pixel values of the adjacent frames of the exposure anomaly frame to obtain the video clip to be processed after de-flickering processing.

[0057] Specifically, after determining the exposure anomaly frame, deviation correction processing can be performed on the exposure anomaly frame according to the pixel values of the adjacent frames of the exposure anomaly frame to obtain the video clip to be processed after de-flickering processing.

[0058] Specifically, this application does not limit the process of performing deviation correction processing on the exposure anomaly frame according to the pixel values of the adjacent frames of the exposure anomaly frame to obtain the video clip to be processed after de-flickering processing. Optionally, regression calculation processing can be performed on the pixel values of the adjacent frames of the exposure anomaly frame to obtain the target pixel value of the exposure anomaly frame; perform deviation correction processing on the pixel values of the exposure anomaly frame according to the target pixel value of the exposure anomaly frame to obtain the exposure anomaly frame after deviation correction processing; if both the structural similarity index and the peak signal-to-noise ratio between the exposure anomaly frame after deviation correction processing and the exposure anomaly frame meet the preset conditions, then replace the exposure anomaly frame with the exposure anomaly frame after deviation correction processing; otherwise, increase the number of adjacent frames of the exposure anomaly frame and repeat the process of regression calculation processing and deviation correction processing.

[0059] Among them, during the automatic repair of video frame flickering, exposure abnormal frames are determined based on an artificial intelligence approach. Optionally, a data-driven anomaly detection model can be used to automatically identify exposure problems, demonstrating the pattern classification ability of artificial intelligence. Specifically, the data-driven anomaly detection model mainly identifies exposure abnormal frames through methods such as feature quantization and modeling and an anomaly decision mechanism. Among them, during the process of feature quantization and modeling, video frames can be divided into pixel blocks (such as 8×8), the LSIC value (based on noise variance and spectral energy) of each pixel block is calculated, and a heat histogram is constructed. Among them, during the implementation of the anomaly decision mechanism, it is automatically determined whether a video frame is abnormally exposed by analyzing the LSIC probability density distribution. For example, when the proportion of pixel blocks with low LSIC values exceeds a threshold, an anomaly flag is triggered, similar to the "outlier detection" algorithm in machine learning.

[0060] The automatic repair method for video frame flickering based on artificial intelligence provided by the embodiments of this application obtains a video segment to be processed, determines the local spatial information capacity of video frames, determines exposure abnormal frames from multiple video frames according to the local spatial information capacity, and corrects the exposure abnormal frames according to the pixel values of adjacent frames of the exposure abnormal frames to obtain the video segment to be processed after de-flickering processing. Among them, through the local spatial information capacity, exposure abnormal frames can be accurately determined from the video segment to be processed, further improving the accuracy of video processing, thereby improving the user's video viewing experience.

[0061] Figure 2 It is a schematic flow of the automatic repair method for video frame flickering based on artificial intelligence provided by this application Figure 2 , such as Figure 2 shown. Based on the Figure 1 embodiment, the process of determining exposure abnormal frames from multiple video frames according to the local spatial information capacity is described in detail. The method includes:

[0062] Step S201: Calculate the probability density of the local spatial information sub-capacity of the video frame according to at least one local spatial information sub-capacity included in the local spatial information capacity.

[0063] Specifically, according to at least one local spatial information sub-capacity included in the local spatial information capacity, the probability density of the local spatial information sub-capacity of the video frame can be calculated.

[0064] Among them, the local spatial information capacity includes at least one local spatial information sub-capacity. Among them, the local spatial information sub-capacity is used to characterize the regional energy distribution of pixel blocks in the video frame.

[0065] Specifically, the number of local space information sub-capacities included in the local space information capacity corresponds to the number of pixel blocks included in the video frame. Specifically, a pixel block is a pixel block centered on a pixel point (i, j) in the video frame with a preset size. Herein, the present application does not limit the preset size of the pixel blocks included in the video frame. Optionally, the preset size may be 8 pixels × 8 pixels.

[0066] Specifically, the present application does not limit the process of calculating the probability density of the local space information sub-capacity of the video frame based on at least one local space information sub-capacity included in the local space information capacity. Optionally, the probability density of the local space information sub-capacity of the video frame may be calculated through a probability density function. Among them, the probability density function (abbreviated as PDF) is a core concept in probability theory and mathematical statistics, used to describe the distribution of data, that is, the occurrence probability of data within different value ranges. Among them, by calculating the probability density of the local space information sub-capacity, the distribution of the local space information sub-capacity in the video frame can be understood. The calculation formula of the probability density of the local space information sub-capacity is as follows:

[0067]

[0068] Among them, represents the probability density of the local space information sub-capacity, and represent the weights of the overexposed abnormal area and the non-overexposed abnormal area respectively, and represent the means of the Gaussian distributions of the overexposed abnormal area and the non-overexposed abnormal area respectively, and represent the variances of the overexposed abnormal area and the non-overexposed abnormal area respectively, represents the probability density of the space information sub-capacity of the overexposed abnormal area, represents the probability density of the space information sub-capacity of the non-overexposed abnormal area.

[0069] Among them, the probability density of the local space information sub-capacity indicates the proportion of the number of pixel blocks corresponding to each local space information sub-capacity in the total number of pixel blocks of the video frame.

[0070] Among them, the expression formula of the proportion of the number of pixel blocks corresponding to each local space information sub-capacity indicated by the probability density of the local space information sub-capacity in the total number of pixel blocks of the video frame is: , where is the probability density of the i-th local space information sub-capacity, where n is the number of probability densities of the local space information sub-capacities included in the video frame.

[0071] Step S202: Construct a heat histogram of the video frame according to at least one local space information sub-capacity included in the local space information capacity.

[0072] Specifically, a heat histogram of the video frame can be constructed according to at least one local space information sub-capacity included in the local space information capacity.

[0073] Among them, the heat value is a quantitative index used to intuitively display the distribution of dynamic elements, movement trajectories or interaction heat in the video.

[0074] Specifically, the present application does not limit the process of constructing a heat histogram of the video frame according to at least one local space information sub-capacity included in the local space information capacity. Optionally, first, according to at least one local space information sub-capacity included in the local space information capacity, based on the inverse normalization algorithm, calculate the heat value corresponding to each local space information sub-capacity. The calculation formula of the heat value is as follows:

[0075]

[0076] Among them, the heat value (i,j) represents the heat value corresponding to the local space information sub-capacity of the pixel block centered on the pixel point (i, j), and LSIC(i,j) represents the local space information sub-capacity of the pixel block centered on the pixel point (i, j). represents the smoothing constant, for example, 10 -3 . Among them, in the calculation process of the inverse normalization described in the above calculation formula of the heat value, take the reciprocal of the local space information sub-capacity and normalize it to [0, 255], so that the high-exposure area shows a high heat value.

[0077] Optionally, after obtaining the heat value corresponding to each local space information sub-capacity, a heat histogram of the video frame can be constructed according to the calculated heat value corresponding to each local space information sub-capacity.

[0078] Among them, the heat histogram of the video frame, that is, the heat map frequency band histogram of the video frame, constructs histograms for the energy distributions of the low-frequency, medium-frequency, and high-frequency bands in the heat value respectively. Among them, the heat histogram of the video frame uses the frequency band energy value as the horizontal axis and counts the frequency of the corresponding energy value appearance as the vertical axis.

[0079] Specifically, the relationship between the local spatial information sub-capacity and the thermal histogram of the video frame is as follows: The local spatial information sub-capacity reflects the information of the local structure and intensity contrast of the image. By calculating the local spatial information sub-capacity of each video frame, a thermal map can be constructed, where the energy distribution in different frequency bands reflects different characteristics of the image. For example, the high-frequency band energy reflects the detail distribution of the image, and the exposure abnormal area will cause the high-frequency energy to be concentrated and deviate from the normal range. Therefore, the local spatial information sub-capacity is the basis for constructing the thermal histogram. By analyzing the distribution and mutation points of the thermal histogram, the exposure abnormal area in the image can be further determined.

[0080] Step S203: Determine whether the video frame is an exposure abnormal frame according to the ratio of the number of pixel blocks corresponding to the local spatial information sub-capacities in the video frame to the total number of pixel blocks in the video frame, and / or the thermal histogram of the video frame.

[0081] Specifically, according to the ratio of the number of pixel blocks corresponding to the local spatial information sub-capacities in the video frame to the total number of pixel blocks in the video frame, and / or the thermal histogram of the video frame, it can be determined whether the video frame is an exposure abnormal frame.

[0082] Specifically, this application does not limit the process of determining whether the video frame is an exposure abnormal frame according to the ratio of the number of pixel blocks corresponding to the local spatial information sub-capacities in the video frame to the total number of pixel blocks in the video frame, and / or the thermal histogram of the video frame. Optionally, determining whether the video frame is an exposure abnormal frame according to the ratio of the number of pixel blocks corresponding to the local spatial information sub-capacities in the video frame to the total number of pixel blocks in the video frame, and / or the thermal histogram of the video frame includes:

[0083] If the ratio of the number of pixel blocks corresponding to the local spatial information sub-capacities lower than the local spatial information sub-capacity threshold in the video frame to the total number of pixel blocks in the video frame exceeds the ratio threshold, and the pixel difference value between the video frame and the adjacent frame exceeds the difference value threshold.

[0084] And / or, the slope mutation value of the thermal histogram of the video frame exceeds the mutation threshold, and the pixel difference value between the video frame and the adjacent frame exceeds the difference value threshold, then it is determined that the video frame is an exposure abnormal frame.

[0085] Among them, the local spatial information sub-capacity threshold is a preset threshold of the local spatial information sub-capacity used to determine whether there is an exposure abnormality in the area corresponding to the local spatial information sub-capacity. Optionally, the expression of the local spatial information sub-capacity threshold is as follows:

[0086]

[0087] Among them, T lowrepresents the threshold of the local space information sub - capacity, and μ1 represents the mean of the Gaussian distribution of the exposure - abnormal area. represents the standard deviation of the exposure - abnormal area, where k is a regulation coefficient, usually determined by cross - validation. Optionally, k = 1.5.

[0088] Among them, the ratio threshold is the threshold of a preset ratio used to determine whether a video frame is an exposure - abnormal frame. In this application, the ratio threshold is not limited. Optionally, it can be set to 30%.

[0089] Among them, an adjacent frame refers to a video frame adjacent to the video frame. Specifically, in this application, the number of adjacent frames of the video frame is not limited. Optionally, it can be 3 - 5 consecutive frames adjacent to the current frame.

[0090] Among them, the pixel difference value between the video frame and the adjacent frame is obtained by performing a difference operation on the video frame and the adjacent frame. Optionally, the calculation formula of the pixel difference value between the video frame and the adjacent frame is as follows:

[0091]

[0092] Among them, D i represents the pixel difference value between the i - th video frame and the adjacent frame, F i represents the pixel value of the i - th video frame, F i-1 represents the pixel value of the (i - 1) - th video frame, F i+1 represents the pixel value of the (i + 1) - th video frame.

[0093] Among them, the difference - value threshold is the threshold of the pixel difference value preset for determining whether the corresponding video frame is an exposure - abnormal frame. Specifically, in this application, the difference - value threshold is not limited. Optionally, it can be to first calculate the average value of the pixel difference values between non - exposure - abnormal frames and adjacent frames, and then determine the sum value of the average value and twice the normal fluctuation value as the difference - value threshold.

[0094] Among them, through the difference - value threshold, the distinguishability of the exposure - abnormal frame can be enhanced, and false judgments caused by noise or short - term interference can be excluded.

[0095] Among them, the slope mutation value of the heat - map histogram of the video frame is obtained by performing a cumulative - distribution - function calculation on the heat - map histogram of the video frame and then further calculating the obtained cumulative - distribution - function curve.

[0096] Among them, the mutation threshold is the threshold of the slope mutation value preset for determining whether there is an exposure - abnormal area in the corresponding video frame.

[0097] Among them, in the process of determining whether a video frame is an overexposed abnormal frame, by determining whether the pixel difference value between the video frame and its adjacent frames exceeds the difference value threshold, misjudgment caused by noise or short-term interference can be excluded, the accuracy of determining overexposed abnormal frames can be improved, thereby improving the accuracy of video processing and further enhancing the user's video viewing experience.

[0098] In the process of determining overexposed abnormal frames from multiple video frames according to the local space information capacity provided by the embodiments of the present application, by calculating the probability density of the local space information sub-capacity of the video frame according to at least one local space information sub-capacity included in the local space information capacity, constructing a heat map histogram of the video frame according to at least one local space information sub-capacity included in the local space information capacity, and determining whether the video frame is an overexposed abnormal frame based on the ratio of the number of pixel blocks corresponding to each local space information sub-capacity in the video frame to the total number of pixel blocks in the video frame, and / or the heat map histogram of the video frame. Among them, in the process of determining whether the video frame is an overexposed abnormal frame, it can be determined by at least one method, which can improve the accuracy of determining overexposed abnormal frames, thereby improving the accuracy of video processing and further enhancing the user's video viewing experience.

[0099] Figure 3 Schematic flow of the automatic repair method for video frame flickering based on artificial intelligence provided by the present application Figure 3 , such as Figure 3 shown, based on the Figure 1 or Figure 2 embodiment, the method for correcting the overexposed abnormal frame according to the pixel values of the adjacent frames of the overexposed abnormal frame to obtain the video clip to be processed after de-flickering processing is described in detail. The method includes:

[0100] Step S301: Perform regression calculation processing on the pixel values of the adjacent frames of the overexposed abnormal frame to obtain the target pixel value of the overexposed abnormal frame.

[0101] Specifically, regression calculation processing can be performed on the pixel values of the adjacent frames of the overexposed abnormal frame to obtain the target pixel value of the overexposed abnormal frame.

[0102] Among them, if the overexposed abnormal frame is a non-first / last frame of the video clip to be processed, the adjacent frames of the overexposed abnormal frame include the previous adjacent frame of the overexposed abnormal frame, that is, the adjacent video frame before the overexposed abnormal frame, and the subsequent adjacent frame of the overexposed abnormal frame, that is, the adjacent video frame after the overexposed abnormal frame. Among them, if the overexposed abnormal frame is the first / last frame of the video clip to be processed, the adjacent frames of the overexposed abnormal frame include the subsequent adjacent frame of the overexposed abnormal frame, that is, the adjacent video frame after the overexposed abnormal frame, or the previous adjacent frame of the overexposed abnormal frame, that is, the adjacent video frame before the overexposed abnormal frame.

[0103] Specifically, this application does not limit the number of adjacent video frames. Among them, the adjacent frames for regression calculation in this step are adjacent video frames, that is, one front adjacent frame and / or one back adjacent frame.

[0104] Among them, if the exposure abnormal frame is a non-head and non-tail frame of the video segment to be processed, the calculation formula for the target pixel value of the exposure abnormal frame is as follows:

[0105]

[0106] Among them, F corrected (x,y) represents the target pixel value corresponding to the pixel point (x,y) of the exposure abnormal frame, 、 、and represent regression coefficients, F prev (x,y) represents the pixel value of the pixel point (x,y) of the front adjacent frame of the exposure abnormal frame, F next (x,y) represents the pixel value of the pixel point (x,y) of the back adjacent frame of the exposure abnormal frame.

[0107] Among them, if the exposure abnormal frame is a head or tail frame of the video segment to be processed, the calculation formula for the target pixel value of the exposure abnormal frame is as follows:

[0108]

[0109] Among them, F corrected (x,y) represents the target pixel value corresponding to the pixel point (x,y) of the exposure abnormal frame, 、 represent regression coefficients, F next (x,y) represents the pixel value of the pixel point (x,y) of the back adjacent frame of the exposure abnormal frame.

[0110] Among them, if the exposure abnormal frame is the tail frame of the video segment to be processed, the calculation formula for the target pixel value of the exposure abnormal frame is as follows:

[0111]

[0112] Among them, F corrected (x,y) represents the target pixel value corresponding to the pixel point (x,y) of the exposure abnormal frame, 、 represent regression coefficients, F prev (x,y) represents the pixel value of the pixel point (x,y) of the front adjacent frame of the exposure abnormal frame.

[0113] Step S302: Correct the pixel value of the exposure abnormal frame according to the target pixel value of the exposure abnormal frame to obtain the exposure abnormal frame after correction processing.

[0114] Specifically, according to the target pixel values of the exposure abnormal frame, the pixel values of the exposure abnormal frame can be corrected to obtain the exposure abnormal frame after the correction process.

[0115] Specifically, this application does not limit the process of correcting the pixel values of the exposure abnormal frame according to the target pixel values of the exposure abnormal frame to obtain the exposure abnormal frame after the correction process. Optionally, it can be solved by the target optimization function of the least squares method. The formula of the target optimization function is as follows:

[0116]

[0117] Among them, F abnormal is the pixel value of the exposure abnormal frame, and F abnormal (x, y) is the pixel value corresponding to the pixel point (x, y) of the exposure abnormal frame.

[0118] Specifically, based on the above-described target optimization function, pixel-by-pixel correction can be performed, that is, pixel-by-pixel correction process. Specifically, all pixels of the exposure abnormal frame can be traversed, and regression calculation is performed using the corresponding pixel values of adjacent frames to generate the corrected pixel values. This method can suppress local brightness mutations by constraining the pixel information of the exposure abnormal frame with the pixel information of adjacent frames.

[0119] Step S303: If the structural similarity index and peak signal-to-noise ratio between the exposure abnormal frame after the correction process and the exposure abnormal frame both meet the preset conditions, then replace the exposure abnormal frame with the exposure abnormal frame after the correction process; otherwise, increase the number of adjacent frames of the exposure abnormal frame, and repeat the regression calculation process and the correction process.

[0120] Specifically, if the structural similarity index and peak signal-to-noise ratio between the exposure abnormal frame after the correction process and the exposure abnormal frame both meet the preset conditions, it can be determined that the exposure abnormal frame after the correction process is not distorted. At this time, the exposure abnormal frame after the correction process can be used to replace the exposure abnormal frame.

[0121] Among them, this application does not limit the preset conditions. Optionally, it can be that the structural similarity index between the exposure abnormal frame after the correction process and the exposure abnormal frame is greater than or equal to the structural similarity index threshold, and the peak signal-to-noise ratio is greater than or equal to the peak signal-to-noise ratio threshold.

[0122] Among them, the structural similarity index SSIM simulates the distortion degree of the exposure abnormal frame after the correction process perceived by the human eye through three dimensions: luminance, contrast, and structure. The structural similarity index threshold is a preset threshold for determining whether the exposure abnormal frame is distorted. Optionally, the structural similarity index threshold can be set to 0.85.

[0123] Among them, the peak signal-to-noise ratio (PSNR) measures the distortion degree of the abnormally exposed frame after rectification processing by comparing the numerical differences of pixel values. The peak signal-to-noise threshold is a preset threshold of the peak signal-to-noise used to determine whether the abnormally exposed frame is distorted. Optionally, the peak signal-to-noise threshold can be set to 25 dB.

[0124] Specifically, if the structural similarity index and the peak signal-to-noise ratio between the abnormally exposed frame after rectification processing and the abnormally exposed frame do not meet the preset conditions, it can be determined that the abnormally exposed frame after rectification processing is distorted. At this time, the number of adjacent frames of the abnormally exposed frame can be increased, and the processes of regression calculation processing and rectification processing can be repeated until it is determined that the abnormally exposed frame after rectification processing is not distorted.

[0125] Among them, compared with the process described in step S301, in the process of increasing the number of adjacent frames of the abnormally exposed frame and repeating the processes of regression calculation processing and rectification processing, the adjacent frames for regression calculation processing are at least two adjacent film frames, that is, at least two previous adjacent frames and / or at least two subsequent adjacent frames.

[0126] Among them, in the process of repeating the regression calculation processing, if the abnormally exposed frame is a non-first and non-last frame of the film segment to be processed, the calculation formula of the target pixel value of the abnormally exposed frame is as follows:

[0127]

[0128] Among them, F corrected (x,y) represents the pixel point of the abnormally exposed frame corresponding target pixel value, , , , , , and represent regression coefficients, represents the pixel value of the pixel point (x,y) of the nth previous adjacent frame of the abnormally exposed frame, represents the pixel value of the pixel point (x,y) of the mth subsequent adjacent frame of the abnormally exposed frame.

[0129] Among them, in the process of repeating the regression calculation processing, if the abnormally exposed frame is the first or last frame of the film segment to be processed, the calculation formula of the target pixel value of the abnormally exposed frame is as follows:

[0130]

[0131] Among them, F corrected (x,y) represents the target pixel value corresponding to the pixel point (x,y) of the abnormally exposed frame, , , , and represents the regression coefficient, represents the pixel value of the pixel point (x, y) of the m-th subsequent adjacent frame of the exposure abnormal frame.

[0132] Among them, during the process of repeated regression calculation processing, if the exposure abnormal frame is the last frame of the film segment to be processed, the calculation formula for the target pixel value of the exposure abnormal frame is as follows:

[0133]

[0134] Among them, F corrected (x, y) represents the target pixel value corresponding to the pixel point (x, y) of the exposure abnormal frame, , , , and represents the regression coefficient, represents the pixel value of the pixel point (x, y) of the n-th previous adjacent frame of the exposure abnormal frame.

[0135] In the process of correcting the exposure abnormal frame according to the pixel values of the adjacent frames of the exposure abnormal frame provided in the embodiments of the present application to obtain the film segment to be processed after de-flickering processing, by performing regression calculation processing on the pixel values of the adjacent frames of the exposure abnormal frame, the target pixel value of the exposure abnormal frame is obtained, and the pixel value of the exposure abnormal frame is corrected according to the target pixel value of the exposure abnormal frame to obtain the exposure abnormal frame after correction processing. Among them, through the regression calculation processing, the target pixel value of the exposure abnormal frame can be accurately and efficiently obtained, so as to improve the accuracy of the exposure abnormal frame after correction processing, improve the accuracy of film processing, and further improve the user's film viewing experience.

[0136] Figure 4 is a schematic flow of the automatic repair method for film frame flicker provided by the present application Figure 4 , as Figure 4 shown, based on the embodiment of Figure 1 or Figure 2 or Figure 3 On the basis of the embodiment, the local space information capacity of the film frame is determined and described in detail. The method includes:

[0137] Step S401, obtain the noise variance of the film frame and obtain the energy information of each pixel block in the film frame.

[0138] Specifically, the noise variance of the film frame can be obtained.

[0139] Among them, the noise variance is a core concept in signal processing, statistics, and machine learning, and is used to quantify the degree of dispersion or fluctuation intensity of random noise.

[0140] Specifically, this application does not limit the process of obtaining the noise variance of the video frames. Optionally, the noise variance of the video frames can be obtained based on the local adaptive noise calibration method.

[0141] Among them, the formula for the noise variance is as follows:

[0142]

[0143] Among them, S represents the set of flat regions, |S| represents the number of pixels in the set S, where the flat regions include but are not limited to regions such as the sky and walls. I(x, y) represents the pixel value of the image at the position (x, y), and μ S represents the average value of the pixel values in the set S. Among them, calculating the average value of the pixels in the flat region is to obtain the average level of the overall brightness or color of the region, which helps to eliminate the overall brightness change or color shift in the image and make the subsequent noise variance calculation more accurate.

[0144] Among them, through the above-described formula for calculating the noise variance, we can effectively estimate the noise level in the video frames, providing an important reference for subsequent image processing, and it is applicable to local noise analysis, especially in the case where there are obvious flat regions in the image.

[0145] For example, assume that the pixel values of the following flat region S in the video frame are: [100, 102, 98, 101, 99, 103, 100, 97, 101], and the average value of the pixel values in the calculated set S: μ S =(100 + 102 + 98 + 101 + 99 + 103 + 100 + 97 + 101) / 9 = 100, calculate the square of the difference between each pixel in the set S and the average value: [0, 4, 4, 1, 1, 9, 0, 9, 1], and calculate the variance =(0 + 4 + 4 + 1 + 1 + 9 + 0 + 9 + 1) / 9 = 3.11.

[0146] Specifically, this application does not limit the process of obtaining the noise variance of the video frames. Optionally, the noise variance of the video frames can be obtained based on the non-local means noise algorithm.

[0147] Among them, the formula for the noise variance is as follows:

[0148]

[0149] Among them, is the pixel value of the k-th similar image patch. Similar image patches refer to regions in an image that have similar structural and textural features. These patches usually repeat in the image or have similar patterns at different locations. By finding similar image patches, the statistical information of these patches can be used to estimate the noise variance and improve the accuracy of noise estimation. In addition, similar image patches can also be used for image inpainting and enhancement by referring to the information of similar patches to restore or improve the image quality.

[0150] Among them, is the calculation of the mean local energy of the k-th similar image patch. N represents the number of similar image patches

[0151] Specifically, the non-local means noise algorithm is an effective noise estimation and denoising method. It estimates the noise variance by finding similar image patches in the image and using the statistical information of these similar patches. The non-local means algorithm can more accurately estimate the noise level in the image because similar image patches usually have similar structural and textural features. By analyzing these similar patches, the noise and signal can be better separated, thus improving the accuracy of noise estimation. It is suitable for global noise analysis, especially when there are a large number of repeated textures or structures in the image.

[0152] For example, assume that there are two similar image patches in a movie frame: Similar patch 1: [100, 102, 98, 101], the mean local energy of the first similar image patch = 100.25, Similar patch 2: [99, 101, 97, 100], the mean local energy of the second similar image patch = 99.25. Among them, the process of calculating the variance of each similar image patch is as follows:

[0153] (100 - 100.25)^2+(102 - 100.25)^2+(98 - 100.25)^2+(101 - 100.25)^2 = 5.5

[0154] (99 - 99.25)^2+(101 - 99.25)^2+(97 - 99.25)^2+(100 - 99.25)^2 = 5.5

[0155] Noise variance =(5.5 + 5.5) / 2 = 5.5

[0156] Specifically, the energy information of each pixel block in the movie frame can be obtained.

[0157] Among them, the energy information is the information characterizing the energy distribution of the pixel block. Optionally, the energy information includes but is not limited to frequency domain energy information and structural energy information.

[0158] Specifically, the frequency-domain energy information, that is, the local spectral energy, focuses on the frequency-domain information. The local spectral energy mainly concerns the distribution of the image in the frequency domain. By performing a windowed Fourier transform on the image block, the image is transformed from the spatial domain to the frequency domain, and the energy distribution of different frequency components in the image is analyzed. This helps to understand the frequency characteristics of features such as textures and edges contained in the image block. For example, high-frequency components are usually related to details and edges in the image, while low-frequency components are related to the overall brightness and slowly changing regions of the image. The calculation of the local frequency-domain energy is to transform the image block from the spatial domain to the frequency domain through the Fourier transform and then calculate its energy distribution. This is like decomposing a painting into waves of different frequencies to analyze the detailed information contained therein. The energy in the frequency domain reflects the intensity of different frequency components in the image block. High-frequency components correspond to details and edges in the image, and low-frequency components correspond to the overall brightness and color changes.

[0159] Optionally, the calculation process of the frequency-domain energy information is as follows:

[0160] First, select an image block: Select a pixel block of size Δ×Δ from the image, i.e., the video frame.

[0161] Second, perform windowing: Apply the Hann window function to the selected image block to reduce spectral leakage. This is similar to adding a gradually dimming border to the image block, making the edges of the image block transition smoothly and avoiding incorrect frequency components caused by mutations.

[0162] Third, perform Fourier transform: Perform a two-dimensional Fourier transform on the windowed image block to transform it into the frequency domain and obtain a complex representation in the frequency domain.

[0163] Then, calculate the energy spectrum: Calculate the square of the magnitude of each frequency point in the frequency domain to obtain the energy spectrum.

[0164] Next, apply the frequency filter H(u,v) to select the energy in a specific frequency range. Usually, the mid-low frequency energy is retained and the high-frequency noise is removed.

[0165] Finally, accumulate the energy: Accumulate the energy in the selected frequency range to obtain the local frequency-domain energy.

[0166] Specifically, the structural energy information, that is, the local structural energy, focuses on the gradient information in the spatial domain. The local structural energy mainly concerns the structural features of the image in the spatial domain, especially the edge and texture information in the image. By calculating the gradient change of the pixel values within the image block in the spatial domain, the local structural features of the image block are reflected. The larger the gradient magnitude, the richer the structural information in this area, such as obvious edges or textures; the smaller the gradient magnitude, the smoother this area is and the less structural information it has. The calculation of the local structural energy is to measure the structural information of the image block by calculating the sum of the squared gradient magnitudes of each pixel point in the image block. The gradient reflects the edge and texture information of the image, and the areas with large gradients usually correspond to the areas with rich edges or textures in the image. The higher the local structural energy, the richer the structural information in the image block.

[0167] Optionally, the calculation process of the structural energy information is as follows:

[0168] First, select an image block: Select a pixel block of size 8×8 from the image, that is, the video frame.

[0169] Second, calculate the gradient: Use the Sobel operator to calculate the gradients of each pixel point in the x and y directions in the image block respectively. The gradient reflects the rate of change of the image in this direction.

[0170] Then, calculate the sum of the squared gradient magnitudes: Square the gradients in the x and y directions of each pixel point and add the two to obtain the squared gradient magnitude of this pixel point.

[0171] Finally, accumulate the squared gradient magnitudes of all pixel points in the entire image block to obtain the local structural energy.

[0172] Step S402: Determine the local spatial information capacity according to the noise variance and the energy information of each pixel block.

[0173] Specifically, according to the noise variance calculated in step S401 and the energy information of each pixel block, the local spatial information capacity can be determined. Among them, the formula for the local spatial information capacity is as follows:

[0174]

[0175] Among them, LSIC(i,j) represents the local spatial information sub-capacity of the pixel block centered on the pixel point (i, j), represents the energy information of the pixel block centered on the pixel point (i, j), represents the noise variance, represents a very small constant to prevent the denominator from being zero, usually taking 10 -6 .

[0176] Among them, the result of the local spatial information capacity calculated through the formula of the local spatial information capacity described above is a two-dimensional matrix, where each element corresponds to the local spatial information sub-capacity of a pixel block in the video frame. This matrix can be used to generate a heat map. In the heat map, the area with a high local spatial information sub-capacity corresponds to the area with significant structural features in the image, while the area with a low local spatial information sub-capacity may correspond to the area with high noise or weak structural features. By analyzing this heat map, the high-exposure area and non-exposure abnormal area in the image can be effectively identified.

[0177] In the process of determining the local spatial information capacity of a video frame provided by an embodiment of the present application, by obtaining the noise variance of the video frame, obtaining the energy information of each pixel block in the video frame, and determining the local spatial information capacity according to the noise variance and the energy information of each pixel block. Among them, by according to the noise variance and the energy information of each pixel block, the local spatial information capacity of the video frame can be accurately determined, thereby improving the accuracy of video processing and further improving the user's video viewing experience.

[0178] Figure 5 Schematic flow of the automatic repair method for video frame flickering based on artificial intelligence provided by the present application Figure 5 , such as Figure 5 shown, in this embodiment, on the basis of Figure 1 or Figure 2 or Figure 3 or Figure 4 embodiment, the acquisition of the video segment to be processed is described in detail. The method includes:

[0179] Step S701, obtain the complete video.

[0180] Specifically, the complete video can be obtained. Among them, the complete video is a video picture including at least one video segment to be processed. Among them, the description of the video segment to be processed can refer to the description in the above-mentioned embodiment and will not be elaborated here.

[0181] Step S702, determine the similarity between each adjacent video frame in the complete video based on the deep feature extraction technology and the shallow feature extraction technology.

[0182] Specifically, based on the deep feature extraction technology and the shallow feature extraction technology, the similarity between each adjacent video frame in the complete video can be determined.

[0183] Specifically, this application does not limit the process of determining the similarity between each adjacent video frame in a complete video based on deep feature extraction technology and shallow feature extraction technology. Optionally, the deep features of each video frame in the complete video can be obtained based on deep feature extraction technology, and the shallow features of each video frame in the complete video can be obtained based on shallow feature extraction technology; the deep features and shallow features are subjected to feature fusion processing to obtain the fusion feature vector of each video frame; according to the fusion feature vector of each video frame, the similarity between each adjacent video frame in the complete video is determined.

[0184] Specifically, this application does not limit the deep feature extraction technology. Optionally, a pre-trained convolutional neural network VGG16 can be used to extract deep features. Specifically, each frame image of the complete video can be input into the VGG16 network, and a 4096-dimensional feature vector of its fully connected layer (FC7 layer) can be extracted. Among them, these feature vectors can reflect the high-level semantic information of the image, such as object category, scene layout, etc., and are highly sensitive to changes in the overall content of the picture, providing a semantic basis for subsequent scene change detection.

[0185] Among them, the VGG network is trained with a large amount of image data and has a powerful feature expression ability, which can effectively capture the semantic differences between different scenes in the video. For example, when switching from an indoor scene to an outdoor scene, its deep features will change significantly, thus laying a foundation for accurate scene segmentation.

[0186] Specifically, this application does not limit the shallow feature extraction technology. Optionally, the shallow features can be extracted based on the grayscale histogram analysis method and the filtering texture analysis method.

[0187] Among them, in the process of extracting shallow features based on the grayscale histogram analysis method, the RGB image of each video frame in the complete video can be first converted into a grayscale image, and then the grayscale histogram corresponding to the grayscale image can be generated. Among them, the formula used in the process of converting the RGB image into a grayscale image is as follows:

[0188]

[0189] Among them, R, G, and B are the pixel values of the red, green, and blue channels in the image (pixel value range: 0 - 255), and I qray is the converted grayscale value (pixel value range: 0 - 255). This conversion process is based on the sensitivity difference of the human eye to different colors, with the green weight being the largest, the red weight being the second, and the blue weight being the smallest, which conforms to the human eye's visual perception characteristics, can effectively reduce the impact of color changes on subsequent processing, and highlight the brightness change information.

[0190] Among them, the formula in the process of generating the grayscale histogram corresponding to the grayscale image is as follows:

[0191]

[0192] Among them, in the process of generating the grayscale histogram, the pixel grayscale range of 0 - 255 is divided into 16 intervals, that is, 16 levels per interval, and the number of pixels in each interval is counted to obtain a 16-bin grayscale histogram; then the normalized histogram is calculated.

[0193] Among them, count(k) is the number of pixels in the k-th grayscale interval, and N pixels = width × height is the total number of pixels in the image, where width is the number of pixels horizontally and height is the number of pixels vertically. Among them, the normalization process makes the histogram unaffected by the image resolution, facilitating comparison between different video frames, and the probability value range is fixed in [0,1], which is beneficial for subsequent feature fusion.

[0194] Among them, the grayscale histogram can reflect the overall light and dark distribution of the picture and has a certain anti-interference ability against light changes, scene brightness differences, etc. For example, in a scene where it gradually gets darker from a bright area, the grayscale histogram will show a corresponding change trend, but will not produce abrupt differences due to local brightness fluctuations, thus providing a basis for stable segmentation.

[0195] Among them, in the process of extracting shallow features based on the filtering texture analysis method, the shallow features can be extracted based on the Gabor filter. Among them, the Gabor filter is a frequency domain analysis tool designed based on the biological visual perception mechanism. By simulating the sensitivity of the human visual system to spatial frequency and direction, it can efficiently extract the texture features of the image. It is widely used in fields such as texture classification, face recognition, and medical image analysis and is one of the classic feature extraction methods in computer vision. The formula of the filter is as follows:

[0196]

[0197] Among them, x' = xcosθ + ysinθ, y' = -xsinθ + ycosθ. Among them, the parameter λ is the wavelength (used to control the fluctuation period of the filter), θ is the direction angle (indicating the texture detection direction of the filter), is the standard deviation (indicating the width of the Gaussian envelope, affecting locality), and σ is the aspect ratio (used to control the ellipticity of the filter).

[0198] Among them, in the process of extracting shallow features based on the above filters, the conjugate matrix energy of the filtering result can be calculated for at least one filter direction, that is, the texture detection direction of the filter. Optionally, the filter directions can be (0°, 45°, 90°, 135°). The formula for calculating the conjugate matrix energy of the filtering result is as follows:

[0199]

[0200] Among them, is the convolution result of the video frame I and the Gabor filter. Among them, the convolution result is a complex matrix), is the real part of the filtering result, which is used to reflect the edge and texture intensity), is the imaginary part of the filtering result, which contains phase information.

[0201] Specifically, the energy values in four directions form a 4D vector .

[0202] Among them, the Gabor filter can capture the texture features of images in multiple scales and multiple directions, and is suitable for distinguishing scenes with different textures. For example, when switching from a grassland scene to a brick wall scene, there will be obvious differences in texture features. By calculating the energy response, the texture information is quantified into numerical values, which is convenient for subsequent processing, and the energy value directly reflects the texture complexity. For example, a high energy value corresponds to a complex texture, and a low energy value corresponds to a simple texture or a solid color area.

[0203] Specifically, in the process of performing feature fusion processing on deep features and shallow features to obtain the fusion feature vector of each video frame, the features extracted by the VGG network can be first subjected to feature normalization processing to obtain the normalized deep features; then the gray histogram and energy features in the shallow features can be subjected to weighted splicing processing to obtain the spliced shallow features; and then the normalized deep features and the spliced shallow features can be subjected to global feature fusion processing to obtain the fusion feature vector of each video frame.

[0204] Among them, in the process of performing feature normalization processing on the features extracted by the VGG network to obtain the normalized deep features, the deep feature vector can be subjected to L2 normalization to obtain , where is the L2 norm of the VGG feature vector, that is, the vector modulus. After normalization, the modulus of the feature vector is 1, which eliminates the dimensional difference and makes the deep features of different video frames comparable on the same scale, avoiding the influence of feature amplitude fluctuations caused by non-semantic factors such as brightness on similarity calculation.

[0205] Among them, in the process of weighted splicing of the grayscale histogram and energy features in the shallow features to obtain the shallow features after splicing processing, the normalized grayscale histogram is multiplied by the weight of 0.3, and the Gabor energy feature is multiplied by the weight of 0.7, and then spliced into a 20-dimensional vector. Among them, the weight assignment is optimized based on experience. Considering that the human eye is more sensitive to texture than to brightness changes, a higher weight is given to the Gabor feature to highlight the role of texture features in segmentation, while retaining the ability of the grayscale histogram to capture overall brightness changes.

[0206] Among them, in the process of global feature fusion of the normalized depth features and the shallow features after splicing processing to obtain the fusion feature vector of each video frame, the normalized deep features are multiplied by the weight of 0.6, and the shallow fusion features are multiplied by the weight of 0.4, and then vector splicing is performed to obtain the final fusion feature vector. Among them, the fusion weight is set according to the sensitivity difference of the deep features and the shallow features to scene changes. The deep features are more adept at capturing semantic mutations (such as scene switching), and the shallow features are more adept at capturing gradual changes (such as lighting changes). By reasonably allocating weights, the advantages of both are complementary, improving the segmentation accuracy.

[0207] Specifically, in the process of determining the similarity between each adjacent video frame in the complete video according to the fusion feature vector of each video frame, the initial similarity between adjacent video frames can be calculated first based on the cosine similarity algorithm, and then mean smoothing processing is performed based on a sliding window to obtain the similarity between adjacent video frames.

[0208] Among them, in the process of calculating the initial similarity between adjacent video frames based on the cosine similarity algorithm, the calculation formula of the cosine similarity is as follows:

[0209]

[0210] Among them, is the fusion feature vector of the t-th video frame, and is the fusion feature vector of the (t + 1)-th video frame. Among them, S(t, t + 1) is the initial similarity between the t-th video frame and the (t + 1)-th video frame. The value range of the initial similarity is [-1, 1]. The larger the value, the more similar the content of the two frames of the picture is, and the smaller the value, the greater the difference. This formula quantifies the similarity degree of adjacent frames in the fusion feature space by calculating the ratio of the dot product of two vectors to the product of the modulus lengths, providing a basis for subsequent tangent point determination.

[0211] Among them, in the process of obtaining the similarity between adjacent video frames through mean smoothing based on a sliding window, the size of the sliding window is not limited. Optionally, the sliding window W can be 5 frames, that is, two frames before and after the current video frame. Among them, in the process of obtaining the similarity between adjacent video frames through mean smoothing based on a sliding window, the calculation formula for the local average similarity is as follows:

[0212]

[0213] Among them, W = 5 is the window size. The smoothing process can suppress misjudgments caused by single-frame mutations (such as flashlights and camera jitters), enhance the sensitivity to low similarity in consecutive multiple frames, and help detect gradual changes such as fades.

[0214] Step S703: Divide the complete video according to the similarity between each pair of adjacent video frames in the complete video to obtain at least one video segment to be processed.

[0215] Specifically, according to the similarity between each pair of adjacent video frames in the complete video, the complete video can be divided to obtain at least one video segment to be processed.

[0216] Optionally, dividing the complete video according to the similarity between each pair of adjacent video frames in the complete video to obtain at least one video segment to be processed includes:

[0217] If there is a similarity between a pair of adjacent video frames in the complete video that is less than the first similarity threshold, then the two video frames in the adjacent frames are divided into video segments under different video scenes to obtain the video segments to be processed.

[0218] Specifically, this application does not limit the first similarity threshold. Optionally, it can be 0.6.

[0219] Among them, if there is a similarity between a pair of adjacent video frames in the complete video that is less than the first similarity threshold, then this scene detection process is a hard cut detection process. Among them, a hard cut usually corresponds to an obvious scene switch, such as directly switching from an indoor scene to an outdoor scene, where the picture content changes drastically and the similarity between two adjacent video frames drops sharply. By setting a lower threshold, that is, the first similarity threshold, such cut points can be captured quickly and accurately.

[0220] If there are multiple consecutive adjacent video frames in the complete video whose similarity is greater than or equal to the first similarity threshold and less than the second similarity threshold, then based on dynamic programming technology, the video frames in the multiple consecutive adjacent video frames are subjected to scene division processing to obtain the video segments to be processed.

[0221] Specifically, the second similarity threshold of the present application is not limited. Any value greater than the first similarity threshold and less than 1 can be used as the second similarity threshold provided by the present application. Optionally, it can be 0.8.

[0222] Among them, if the similarity between multiple consecutive adjacent video frames in the complete video is greater than or equal to the first similarity threshold and less than the second similarity threshold, the scene detection process is a fade detection process. Among them, the scene corresponding to the fade detection process is a fade scene. For example, assume we are watching a movie and the picture gradually transitions from day to night. The change of this scene is not sudden but gradual. Specifically, to detect such a fade scene, we use the fade detection method to determine whether the picture is changing gradually.

[0223] Specifically, if there is a similarity between multiple consecutive adjacent video frames. For example, the similarity between every two adjacent frames among 10 consecutive frames is within the interval [the first similarity threshold, the second similarity threshold]. For example, , where k ∈ [0, 9], then based on the dynamic programming technique, the video frames in multiple consecutive adjacent video frames are subjected to scene partitioning processing to obtain the video segments to be processed. Among them, dynamic programming (DP) is a mathematical method that can help us find the global optimal segmentation point within the candidate interval, rather than just finding the local optimal cut point like the greedy algorithm. The role of dynamic segmentation is that when a fade scene is detected, it can help find the best cut point for the picture change.

[0224] Among them, the formula for performing scene partitioning processing on the video frames in multiple consecutive adjacent video frames based on the dynamic programming technique is as follows:

[0225]

[0226] Among them, DP(i) represents the position of the best cut point from the first frame to the i-th frame. j represents the possible cut point position, and all possible j are traversed (from 1 to i - 1). represents the weight coefficient, which is used to adjust the influence of similarity on the cut point selection.

[0227] Among them, S(k, k + 1) represents the similarity between the k-th frame and the (k + 1)-th frame. The higher the similarity, the closer the value is to 1. 1 - S(k, k + 1) is to convert the similarity into a kind of "difference degree". The higher the similarity, the lower the difference degree, where the difference degree approaches 0. The lower the similarity, the higher the difference degree, and the difference degree approaches 1. Specifically, the higher the difference degree, the more obvious the change between two frames, and the more likely there is a cut point between the two frames.

[0228] Among them, It is to calculate the sum of the differences between all adjacent frames from the tangent point j to the current frame i. Among them, this accumulated value represents the total change amount of the gradual change process from j to i. The smaller the accumulated value, the smoother the gradual change process from j to i, and the better the visual effect at the tangent point. We hope to find a tangent point j such that the gradual change process from j to i is as smooth as possible.

[0229] Among them, α is a weight coefficient used to adjust the influence degree of the difference on the selection of the tangent point. It can help us balance the selection of the tangent point position and the smoothness of the gradual change process.

[0230] Optionally, the process of determining the tangent point based on the above formula is as follows:

[0231] First, perform initialization processing. For example, set DP(1)=0, indicating that there is no tangent point in the first frame.

[0232] Secondly, perform iterative calculation processing. For example, for each i (from 2 to the total number of frames of the video), calculate the value of DP(i). Specifically, all possible j (from 1 to i - 1) can be traversed to calculate , and select the minimum value among them as DP(i).

[0233] Then, determine the final result. Among them, the final DP(i) represents the optimal tangent point position of the entire video.

[0234] Finally, perform the backtracking path processing. Specifically, starting from the last frame, trace back frame by frame to find each tangent point position. The specific steps are as follows: Starting from i = n, find the j that makes the smallest. Record this j as the tangent point position. Update i to j and continue backtracking until i = 1.

[0235] Optionally, an example of the process of determining the tangent point based on the above description is as follows:

[0236] Suppose there is a complete video containing 4 frames. Among them, the similarity of adjacent frames is as follows: S(1,2)=0.8; S(2,3)=0.5; S(3,4)=0.8; weight coefficient =1.

[0237] Specifically, during the initialization processing, set DP(1)=0.

[0238] Specifically, during the iterative calculation processing:

[0239] Among them, first calculate DP(2), traverse j = 1, and the result is:

[0240] DP(2)=DP(1)+1⋅(1−S(1,2))=0+1⋅(1−0.8)=0.2.

[0241] Among them, calculate DP(3) again, traverse j = 1, j = 2, and the results are:

[0242] When j = 1, the result is:

[0243]

[0244] When j = 2, the result is:

[0245]

[0246]

[0247] Among them, then calculate DP(4), traverse j = 1, j = 2, j = 3, and the results are:

[0248] When j = 1, the result is:

[0249]

[0250] When j = 2, the result is:

[0251]

[0252] When j = 3, the result is:

[0253]

[0254]

[0255] Specifically, determine that the film frames corresponding to DP(3) = 0.9 and DP(4) = 1.1 are the optimal cut point positions, that is, between the first frame and the second frame.

[0256] Specifically, during the process of backtracking the path, starting from i = 4, find the j such that DP(4) = 1.1. According to the calculation process, this j can be 1 or 2. For i = 3: find the j such that DP(3) = 0.9. According to the calculation process, this j is 1 or 2. For i = 2: find the j such that DP(2) = 0.2. According to the calculation process, this j is 1. For i = 1: reach the initial position and stop backtracking.

[0257] Through backtracking, we obtain that the cut point positions are the first frame and the second frame, which means cutting between the 2nd frame and the 1st frame.

[0258] Optionally, to improve the detection accuracy, we introduce a frame skipping detection method, such as checking whether the similarity between the first frame and the fifth frame is between 0.6 and 0.8. Optionally, if the frame rate of the video with a fade scene is high, the number of skipped frames can be appropriately increased; if the frame rate of the video is low, the number of skipped frames can be reduced.

[0259] Among them, in the process of dividing the complete video to obtain at least one video segment to be processed, determining the cutting points in different scenarios based on two methods of hard cut detection and fade detection can improve the efficiency and accuracy of the division process, thereby improving the efficiency and accuracy of video processing and further enhancing the user's video viewing experience.

[0260] Among them, in the process of dividing the complete video described above, if the time interval between two adjacent cut points obtained is less than the preset time interval, for example, Δt = 0.5 seconds (15 frames for a 30 - frame - rate video), then the corresponding time interval is merged into a single interval. This rule avoids the fragmentation of cut points caused by short - term video fluctuations, ensures that the output scene segments are meaningful, and is compatible with video sources of different frame rates. By dynamically calculating the time interval threshold, that is, the frame number threshold, it adapts to the segmentation requirements of various videos.

[0261] Among them, in the process of automatically repairing video frame flickering, deep features and shallow features of video frames are extracted based on artificial intelligence. Specifically, the similarity between adjacent video frames is determined through multi - dimensional feature extraction technology, which reflects the typical application of artificial intelligence in the field of visual analysis. Specifically, in the process of deep feature extraction, a convolutional neural network (CNN) is used to automatically learn the complex semantic features of video frames. For example, through multi - layer convolution and pooling operations, object contours, texture details, and color distribution patterns are captured to form a high - order local spatial information sub - capacity feature vector. Specifically, in the process of shallow feature extraction, combined with traditional image processing algorithms (such as Gabor filtering, grayscale histogram analysis), basic texture and brightness features are extracted and fused with deep features to form an "integrated information feature vector". Specifically, after obtaining the above - described feature vectors based on the artificial intelligence process, intelligent similarity evaluation can be performed. Based on the above - mentioned feature vectors, the system can automatically calculate the cosine similarity or Euclidean distance between frames, providing a quantitative basis for subsequent dynamic programming to segment video segments.

[0262] The process of obtaining a film clip to be processed provided by the embodiment of the present application is as follows: by obtaining a complete film, based on deep feature extraction technology and shallow feature extraction technology, determining the similarity between each adjacent film frame in the complete film, and dividing the complete film according to the similarity between each adjacent film frame in the complete film to obtain at least one film clip to be processed. Among them, through the similarity between adjacent film frames, the complete film can be efficiently and accurately divided to obtain the film clip to be processed, thereby improving the accuracy of film processing and further improving the user's film viewing experience. Among them, in the process of determining the similarity between adjacent film frames, based on the extracted deep features and shallow features, the accuracy of the similarity can be improved, thereby improving the accuracy of film processing and further improving the user's film viewing experience. Based on the above description, the process of obtaining the film clip to be processed provided by the embodiment of the present application can improve the user's film viewing experience.

[0263] Figure 6 FIG. is a schematic structural diagram of an automatic repair device for film picture flickering based on artificial intelligence provided by the present application, as Figure 4 shown, the automatic repair device 60 for film picture flickering based on artificial intelligence provided in this embodiment includes:

[0264] An acquisition module 601, configured to acquire a film clip to be processed; wherein, the film clip to be processed is a clip of film pictures in the same film scene; the film clip to be processed includes multiple film frames;

[0265] A determination module 602, configured to determine the local space information capacity of the film frame, and determine an exposure abnormal frame from multiple film frames according to the local space information capacity; wherein, the local space information capacity is used to characterize the regional energy distribution of the film frame;

[0266] A processing module 603, configured to correct the exposure abnormal frame according to the pixel values of the adjacent frames of the exposure abnormal frame to obtain a film clip to be processed after de-flickering processing.

[0267] In a possible embodiment, the local space information capacity includes at least one local space information sub-capacity; wherein, the local space information sub-capacity is used to characterize the regional energy distribution of pixel blocks in a video frame; the determination module 602 is specifically configured to calculate the probability density of the local space information sub-capacity of the video frame according to at least one local space information sub-capacity included in the local space information capacity; wherein, the probability density of the local space information sub-capacity indicates the proportion of the number of pixel blocks corresponding to each local space information sub-capacity to the total number of pixel blocks in the video frame; construct a heat histogram of the video frame according to at least one local space information sub-capacity included in the local space information capacity; and determine whether the video frame is an exposure abnormal frame according to the proportion of the number of pixel blocks corresponding to each local space information sub-capacity in the video frame to the total number of pixel blocks in the video frame, and / or the heat histogram of the video frame.

[0268] In a possible embodiment, the determination module 602 is further specifically configured to determine that the video frame is an exposure abnormal frame if the proportion of the number of pixel blocks corresponding to the local space information sub-capacity lower than the local space information sub-capacity threshold in the video frame to the total number of pixel blocks in the video frame exceeds the proportion threshold, and the pixel difference value between the video frame and the adjacent frame exceeds the difference value threshold; and / or, the slope mutation value of the heat histogram of the video frame exceeds the mutation threshold, and the pixel difference value between the video frame and the adjacent frame exceeds the difference value threshold.

[0269] In a possible embodiment, the processing module 603 is specifically configured to perform regression calculation processing on the pixel values of the adjacent frames of the exposure abnormal frame to obtain the target pixel value of the exposure abnormal frame; perform deviation correction processing on the pixel values of the exposure abnormal frame according to the target pixel value of the exposure abnormal frame to obtain the exposure abnormal frame after deviation correction processing; if the structural similarity index and peak signal-to-noise ratio between the exposure abnormal frame after deviation correction processing and the exposure abnormal frame meet the preset conditions, then replace the exposure abnormal frame with the exposure abnormal frame after deviation correction processing; otherwise, increase the number of adjacent frames of the exposure abnormal frame, and repeat the process of regression calculation processing and deviation correction processing.

[0270] In a possible embodiment, the determination module 602 is further specifically configured to obtain the noise variance of the video frame, and obtain the energy information of each pixel block in the video frame; determine the local space information capacity according to the noise variance and the energy information of each pixel block.

[0271] In a possible embodiment, the acquisition module 601 is specifically configured to acquire a complete video; wherein, the complete video is a video picture including at least one video segment to be processed; determine the similarity between each adjacent video frame in the complete video based on deep feature extraction technology and shallow feature extraction technology; and perform division processing on the complete video according to the similarity between each adjacent video frame in the complete video to obtain at least one video segment to be processed.

[0272] In a possible embodiment, the obtaining module 601 is further specifically configured to, if the similarity between two adjacent video frames in a complete video is less than a first similarity threshold, divide the two video frames in the adjacent frames into video segments under different video scenes to obtain a video segment to be processed; if the similarity between multiple consecutive adjacent video frames in a complete video is greater than or equal to the first similarity threshold and less than a second similarity threshold, perform scene division processing on the video frames in the multiple consecutive adjacent video frames based on dynamic programming technology to obtain a video segment to be processed.

[0273] The automatic repair device for video frame flicker based on artificial intelligence provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0274] Figure 7 It is a schematic structural diagram of an electronic device provided in this application. As Figure 7 shown, the electronic device 70 provided in this embodiment includes: at least one processor 701 and a memory 702. Optionally, the electronic device 70 further includes a communication component 703. Among them, the processor 701, the memory 702, and the communication component 703 are connected through a bus 704.

[0275] In a specific implementation process, at least one processor 701 executes computer-executable instructions stored in the memory 702, so that at least one processor 701 executes the above method.

[0276] For the specific implementation process of the processor 701, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0277] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated: CPU), and may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated: DSP), application-specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0278] The memory may include a Random Access Memory (RAM), and may also include a Non-volatile Memory (NVM), such as at least one disk memory.

[0279] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0280] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0281] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0282] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a disk, or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0283] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0284] The division of units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed among each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0285] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0286] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0287] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs and other various media that can store program codes.

[0288] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disks or optical discs and other various media that can store program codes.

[0289] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. An automatic repair method for film frame flickering based on artificial intelligence, characterized in that, Including: Obtaining a film clip to be processed; wherein, the film clip to be processed is a clip of film images in the same film scene; the film clip to be processed includes a plurality of film frames; Determining the local spatial information capacity of the film frames, and determining exposure-abnormal frames from the plurality of film frames according to the local spatial information capacity; wherein, the local spatial information capacity is used to characterize the regional energy distribution of the film frames; Performing a rectification process on the exposure-abnormal frames according to the pixel values of the adjacent frames of the exposure-abnormal frames, to obtain the film clip to be processed after de-flickering processing.

2. The method according to claim 1, wherein The local spatial information capacity includes at least one local spatial information sub-capacity; wherein, the local spatial information sub-capacity is used to characterize the regional energy distribution of the pixel blocks in the film frames; Determining exposure-abnormal frames from the plurality of film frames according to the local spatial information capacity includes: Calculating the probability density of the local spatial information sub-capacity of the film frames according to at least one of the local spatial information sub-capacities included in the local spatial information capacity; wherein, the probability density of the local spatial information sub-capacity indicates the proportion of the number of pixel blocks corresponding to each local spatial information sub-capacity to the total number of pixel blocks in the film frame; Constructing a heat histogram of the film frames according to at least one of the local spatial information sub-capacities included in the local spatial information capacity; Determining whether the film frame is an exposure-abnormal frame according to the proportion of the number of pixel blocks corresponding to each local spatial information sub-capacity in the film frame to the total number of pixel blocks in the film frame, and / or, the heat histogram of the film frame.

3. The method according to claim 2, characterized in that, Determining whether the film frame is an exposure-abnormal frame according to the proportion of the number of pixel blocks corresponding to each local spatial information sub-capacity in the film frame to the total number of pixel blocks in the film frame, and / or, the heat histogram of the film frame includes: If the proportion of the number of pixel blocks corresponding to the local spatial information sub-capacities lower than the local spatial information sub-capacity threshold in the film frame to the total number of pixel blocks in the film frame exceeds the proportion threshold, and the pixel difference value between the film frame and the adjacent frame exceeds the difference value threshold; And / or, the slope mutation value of the heat histogram of the film frame exceeds the mutation threshold, and the pixel difference value between the film frame and the adjacent frame exceeds the difference value threshold, then determining that the film frame is an exposure-abnormal frame.

4. The method according to claim 1, wherein Performing a rectification process on the exposure-abnormal frames according to the pixel values of the adjacent frames of the exposure-abnormal frames, to obtain the film clip to be processed after de-flickering processing includes: Performing a regression calculation process on the pixel values of the adjacent frames of the exposure-abnormal frames to obtain the target pixel values of the exposure-abnormal frames; Performing a rectification process on the pixel values of the exposure-abnormal frames according to the target pixel values of the exposure-abnormal frames to obtain the exposure-abnormal frames after rectification processing; If the structural similarity index and peak signal-to-noise ratio between the corrected exposure abnormal frame and the exposure abnormal frame meet the preset conditions, then replace the exposure abnormal frame with the corrected exposure abnormal frame; otherwise, increase the number of adjacent frames of the exposure abnormal frame, and repeat the process of regression calculation processing and correction processing.

5. The method according to claim 1, wherein Determine the local spatial information capacity of the video frame, including: Obtain the noise variance of the video frame, and obtain the energy information of each pixel block in the video frame; Determine the local spatial information capacity according to the noise variance and the energy information of each pixel block.

6. The method according to any one of claims 1-5, characterized in that, Obtain the video segment to be processed, including: Obtain the complete video; wherein, the complete video is a video picture including at least one of the video segments to be processed; Based on the deep feature extraction technology and the shallow feature extraction technology, determine the similarity between each adjacent video frame in the complete video; According to the similarity between each adjacent video frame in the complete video, perform a division process on the complete video to obtain at least one of the video segments to be processed.

7. The method according to claim 6, wherein According to the similarity between each adjacent video frame in the complete video, perform a division process on the complete video to obtain at least one of the video segments to be processed, including: If the similarity between an adjacent video frame in the complete video is less than the first similarity threshold, then divide the two video frames in the adjacent frames into video segments under different video scenes to obtain the video segment to be processed; If there are multiple consecutive adjacent video frames in the complete video whose similarity is greater than or equal to the first similarity threshold and less than the second similarity threshold, then perform a scene division process on the video frames in the multiple consecutive adjacent video frames based on the dynamic programming technology to obtain the video segment to be processed.

8. An automatic repair device for film picture flicker based on artificial intelligence, characterized in that, Including: An acquisition module, configured to acquire a video segment to be processed; wherein, the video segment to be processed is a segment of a video picture under the same video scene; the video segment to be processed includes multiple video frames; A determination module, configured to determine the local spatial information capacity of the video frame, and determine an exposure abnormal frame from the multiple video frames according to the local spatial information capacity; wherein, the local spatial information capacity is used to characterize the regional energy distribution of the video frame; A processing module, configured to perform a correction process on the exposure abnormal frame according to the pixel values of the adjacent frames of the exposure abnormal frame to obtain a processed video segment to be processed after de-flickering.

9. An electronic device, characterized in that, Including: A memory, a processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1-7.

11. A computer program product, characterized in that, Including a computer program, which when executed by a processor implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for removing glitter of image

    CN101257571A

  • Self-adaptive ultrasonic imaging method for inhibiting tissue flicker, and device thereof

    CN101524284A

  • Method and system for restoring flash pulse information

    CN104656119A

  • Method for performing relay selection and sending power distribution by energy transmission full-duplex relay

    CN109302250A

  • Method for reducing flicker

    CN110012234A