A video optimization method, device, equipment and medium in a low-light environment
By acquiring video motion vectors and continuous frames in a low-light environment, converting them into optical flow diagrams and performing iterative optimization, the problem of poor generalization of video denoising optimization methods in the prior art is solved, and efficient video clarity is achieved.
Patent Information
- Application Number
- CN202411383819.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-09-30
AI Technical Summary
The prior art video denoising optimization method is poor in low-light environments and it is difficult to effectively deal with noise from different types of videos.
By obtaining the video motion vector and continuous frames in a low-light environment, converting them into optical flow diagrams, and smoothing them through the image frame context information, determining the correlation between adjacent image frames, iterative optimization, and finally filtering the video.
This method improves the clarity of the video, has high versatility and robustness, and can penetrate deep into the pixel level of the video, capturing and recording the motion details of each pixel point.
Smart Images

Figure CN119273577B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image denoising, and in particular relates to a video optimization method, device, equipment and medium in a low-light environment. Background Art
[0002] Although photographic sensors have improved significantly in recent years, video processing still requires a strong focus on noise reduction, especially when faced with challenging shooting conditions such as low light or small sensor sizes.
[0003] The selection of video denoising models includes non-local similarity model method, blind denoising method and deep learning method. The non-local similarity model method is not adaptable enough when dealing with noise distribution in actual environments, because noise does not always conform to the non-local similarity law. The blind denoising rule often ignores time domain information, resulting in low information utilization and relatively weak competitiveness, especially for noise with large variance, the effect is difficult to be satisfactory. In comparison, the deep learning method is relatively outstanding, but this type of model is more suitable for situations where the noise distribution is determined, such as Gaussian distribution. For noise with obvious changes in the time domain, the deep neural network needs to re-fit each frame of noise data, which is not only time-consuming and labor-intensive, but also has poor model generalization.
[0004] Therefore, in the existing technology, there are still high technical challenges for video denoising optimization in low-light environments. In particular, the existing methods for optimizing denoising of different types of videos are not very versatile, and it is necessary to develop more robust noise processing algorithms and efficient model compression and maintenance strategies. Summary of the invention
[0005] In order to solve the problem that the above-mentioned methods for optimizing different types of videos in the prior art are less versatile, the present invention provides a method, device, equipment and medium for optimizing videos in a low-light environment.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] The present invention provides a video optimization method in a low-light environment, the method comprising: obtaining motion vectors and continuous frames of a target video in the low-light environment; converting the motion vectors into optical flow maps, further smoothing the optical flow maps according to image frame context information in the continuous frames to obtain a rough optical flow map; determining the correlation between two adjacent image frames in the continuous frames; and iteratively optimizing the rough optical flow map according to the correlation.
[0008] Optionally, the method further includes filtering the target video after iterative optimization, and the filtering process includes: dividing each frame of the target video into non-overlapping image blocks, determining the variance mean of the image block difference map at the same position at different time points, and then determining the noise threshold through the distribution of the variance mean; dividing the types of motion vectors according to the noise threshold, and then reconstructing the motion vectors according to different types to obtain trajectory vectors; filtering the target video according to the trajectory vector.
[0009] Optionally, obtaining the motion vector and continuous frames of the target video in a low-light environment includes: extracting the initial motion vector of the target video through a motion vector initialization network; correcting the error of the initial motion vector according to minimizing the neighborhood pixel norm to obtain the motion vector of the target video; and obtaining continuous frames of the target video by decoding the target video.
[0010] Optionally, the step of correcting the error of the initial motion vector by minimizing the neighborhood pixel norm comprises: correcting the motion vector field by minimizing the L2 norm between the central image block and eight adjacent image blocks; the L2 norm comprises an amplitude and a phase, that is,
[0011]
[0012] Where A(c) and P(c) represent the amplitude and phase of the current image block c, respectively; A(n) and P(n) represent the amplitude and phase of the eight adjacent image blocks, (x, y) are spatial coordinates, t is the temporal coordinate, i is the number of the adjacent image block in the neighborhood, ω and λ are weighting coefficients, are the amplitude and phase loss functions of the central image block, respectively.
[0013] Optionally, the step of converting the motion vector into an optical flow map, and further smoothing the optical flow map according to the image frame context information in the continuous frames to obtain a rough optical flow map comprises: filling the pixels in each block with the same motion offset to convert the motion vector into a dense optical flow map; inputting the previous frame of two adjacent frames in the continuous frames into two different encoders to obtain a query matrix Q map and a key matrix K map, and the optical flow map is directly used as a value matrix V map, which is expressed as
[0014] Q=E A (I1),K=E B (I1),V=F MV
[0015] Among them, E A and E B There are two encoder blocks, each consisting of six convolutional layers and corresponding activation layers; a confidence estimation block is used to estimate the confidence of the motion prior for each pixel, which is expressed by the following formula:
[0016] CMV =CEB(I1,M MV )
[0017] Among them, M MV represents the area where the motion vector exists, I1 is the previous frame of the two adjacent frames in the continuous frames, C MV is a weight map ranging from (0,1), CEB is a convolutional neural network (CNN) block; it calculates the correlation S between the center pixel of the local window and other pixels i,j :
[0018] S i,j =softmax(Q i,j ·K i+k,j+l )
[0019] Where k,l∈[-d,d], k and l are the offsets of the local window, indicating the position relative to the center pixel, d is the radius of the window, (i,j) is the position of the center pixel of the local window, and dot· indicates the dot product operation. Therefore, the calculated correlation weight S i,j is a shape of (2d+1) 2 Tensor of; Combine the pixel's credibility and relevance to get the final weight It is expressed as:
[0020]
[0021] From C MV A (2d+1)×(2d+1) window is extracted around the center pixel (i, j), and the ⊙ mark represents the element-by-element multiplication operator; according to the weight Aggregate the motion within the local window to obtain a rough optical flow map F i,j :
[0022]
[0023] V i+k,j+l Represents the initial optical flow graph F MV The motion vector value at (i+k,j+l).
[0024] Optionally, determining the correlation between two adjacent frames in the continuous frames includes: acquiring two continuous frames as I1 and I2, extracting feature maps from the two image frames respectively using a convolutional neural network, recorded as L1 and L2; and determining a correlation matrix C of L1 and L2, the specific formula is as follows:
[0025] C(i,j)=L1(i)·L2(j)
[0026] Among them, C(i,j) is an element of C, which represents the similarity between position i in feature L1 and position j in feature L2, and · represents the dot product operation.
[0027] Optionally, taking the rough optical flow map as an initial value and performing iterative optimization according to the correlation relationship includes:
[0028] In each iteration, the optical flow estimate can be updated by the following formula:
[0029] ΔF MV =u(F MV (k),C)
[0030] F MV (k+1)=F MV (k)+ΔF MV
[0031] Among them, F MV (k) represents the optical flow estimation value of the kth iteration, ΔF MV Indicates the current optical flow estimate F MV (k) and the correlation matrix C to obtain the optical flow update amount of this iteration, u(·) a convolutional neural network (CNN) block; when the preset number of iterations is reached, or the optical flow update amount is less than or equal to the preset update value, the iteration is stopped to obtain a high-precision optical flow map F MV (*).
[0032] The present invention also provides a video optimization device in a low-light environment, characterized in that the device comprises:
[0033] An acquisition module, used to acquire motion vectors and continuous frames of a target video in a low-light environment;
[0034] The optimization module is used to convert the motion vector into an optical flow map, further smooth the optical flow map according to the image frame context information in the continuous frames to obtain a rough optical flow map; determine the correlation between two adjacent image frames in the continuous frames; and iteratively optimize the rough optical flow map according to the correlation.
[0035] The present invention also provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned video optimization method in a low-light environment are implemented.
[0036] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned video optimization method in a low-light environment when executing the program.
[0037] The video optimization method in a low-light environment provided by the present invention has the following beneficial effects:
[0038] By converting motion vectors into optical flow maps, and converting the optical flow maps into smoother rough optical flow maps through the context information of continuous frames of the image, and then iterating the rough optical flow maps as initial values according to the correlation, corresponding optical flow processing is performed for different videos, which has high versatility. Moreover, the optical flow processing can go deep into the pixel level of the target video, and then capture and record the motion details of each pixel, with high robustness. Finally, the iteration of each pixel is realized through optical flow processing, thereby completing the optimization of the target video and improving the clarity of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiment of the present invention and its design scheme, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0040] Figure 1 The figure is a flow chart of a method for optimizing video in a low-light environment according to an exemplary embodiment of the present invention.
[0041] Figure 2 A schematic diagram of a vector generator provided according to an exemplary embodiment of the present invention.
[0042] Figure 3 A schematic diagram of another vector generator provided according to an exemplary embodiment of the present invention.
[0043] Figure 4 A schematic diagram of a motion vector self-correction module provided according to an exemplary embodiment of the present invention.
[0044] Figure 5 A schematic diagram of another motion vector self-correction module provided according to an exemplary embodiment of the present invention.
[0045] Figure 6 The present invention is a schematic diagram of a process for generating a rough optical flow map according to an exemplary embodiment of the present invention.
[0046] Figure 7 The present invention is a schematic diagram of a process for generating a final optical flow map according to an exemplary embodiment of the present invention.
[0047] Figure 8 The present invention is a block diagram of a video optimization device in a low-light environment according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to enable those skilled in the art to better understand the technical solution of the present invention and implement it, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention.
[0049] In view of the complexity of noise distribution in low-light environments, it is difficult to accurately fit it through traditional probability distribution. This paper innovatively proposes a low-light surveillance video denoising algorithm based on deep learning. Specifically, the algorithm first constructs a unique denoising model, which uses a deep vector analysis network and combines the optical flow network to accurately initialize the motion vector, comprehensively considers the vector field and noise information, and efficiently realizes noise removal. Secondly, the optical flow model MVFlow is introduced, which significantly improves the accuracy and speed of video optical flow estimation by virtue of motion vectors. With the help of the optical flow motion vector analysis network, an accurate motion vector field is obtained, and through comprehensive analysis and post-processing, trajectory vectors are generated, and then filtering is implemented along the trajectory to effectively remove noise. Finally, a circuit implementation scheme for a low-light surveillance video denoising model is proposed. Through a lightweight design at the hardware level, an odd-even vector generator is used to optimize motion estimation, and the motion vector refinement filter of the hardware version is upgraded.
[0050] Specifically, it includes the following three steps:
[0051] In the first step, a surveillance video denoising model for low-light environments is proposed and the motion vector field analysis technology of the surveillance video denoising model is established. The noise data is analyzed with the help of video prior information to complete the noise removal.
[0052] In the second step, an optical flow model MVFlow is proposed, which uses motion vectors to improve the speed and accuracy of video optical flow estimation;
[0053] In the third step, a model circuit solution for surveillance video denoising in low-light environments is proposed, which realizes the lightweight of the model at the hardware level.
[0054] The technical solutions provided by various embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.
[0055] First, the present invention provides a video optimization method in a low-light environment, specifically: Figure 1 As shown, the following steps are included:
[0056] S101, obtaining motion vectors and continuous frames of a target video in a low-light environment.
[0057] In this step, the initial motion vector of the target video can be extracted through the motion vector initialization network; the error of the initial motion vector is corrected according to the minimization of the neighborhood pixel norm to obtain the motion vector of the target video.
[0058] For example, initial motion vectors may be generated by a motion vector initialization network, and these initial motion vectors may contain errors. These errors may then be corrected by a motion vector self-correction network, specifically by minimizing the norm of the neighborhood pixels to achieve precise adjustment.
[0059] Specifically, the motion vector field is corrected by minimizing the L2 norm between the central image block and the eight adjacent image blocks; the L2 norm includes amplitude and phase, that is:
[0060]
[0061] Where A(c) and P(c) represent the amplitude and phase of the current image block c, respectively; A(n) and P(n) represent the amplitude and phase of the eight adjacent image blocks, (x, y) are spatial coordinates, t is the temporal coordinate, i is the number of the adjacent image block in the neighborhood, ω and λ are weighting coefficients, are the amplitude and phase loss functions of the central image block, respectively.
[0062] Based on the above steps, the present invention proposes a motion vector initialization module, which is specifically constructed as follows: Figure 2 As shown. This module can generate slightly rough but highly accurate motion vectors in an efficient and cost-effective way. The design is inspired by the non-local similarity denoising technique and the space A is a preset constant (usually set to 16), and its value can be flexibly adjusted according to the actual resolution of the video. For each image block, the system will process it one by one in the order of raster scanning and identify two types of key image blocks: one is two spatially related blocks in the current frame, and the other is a temporally related block in the previous frame. Specifically, in the configuration of the motion vector initialization module, if the image block currently being processed (marked in red) is located at the (x, y) position of the nth frame, its temporally related block (marked in yellow) will be located at the (x, y-2N) coordinates of the n-1th frame, and the two spatially related blocks (marked in blue) are located at the (xN, yN) and (x+N, y+N) positions of the nth frame respectively.
[0063] In addition, in order to realize the lightweight model from the hardware level, the present invention designs a Figure 3The circuit solution for initializing the motion vector shown in FIG. Specifically, the solution divides each frame of the video into non-overlapping image blocks of 8×8 size, and performs parity encoding on these image blocks. At the same time, the time-domain related image blocks and the spatial-domain related image blocks of the current image block are clearly defined to adapt to different hardware processing sequences. This design is based on the following two reasons: 1. It is derived from the rules obtained from the statistical analysis of massive video data; 2. It meets the objective requirements of hardware architecture and acceleration programs. When processing vectors, two different processing orders, odd and even, are used. When the scanning order is from top to bottom, the motion vector is initialized in these two different ways.
[0064] After the motion vector is initialized, the generated motion vector often appears to be relatively rough and cannot meet the expected accuracy requirements. Therefore, the present invention further designs a motion vector self-correction module. This module subdivides the original image block into four small image blocks, where The value of is 8, such as Figure 4 As shown in the figure. Each small image block to be refined has 8 adjacent small image blocks around it. By refining the motion vectors of these adjacent small image blocks, the continuity and smoothness of the vector field in the local two-dimensional space are ensured. In order to achieve this goal, the present invention uses the least squares method to constrain the weight coefficient of each adjacent vector, that is, to minimize the error The determination of the weighting coefficient mainly depends on the distance relationship between the central small image block and the eight adjacent small image blocks. In order to fully reflect the overall continuity of the motion vector, the present invention considers the vector continuity characteristics in seven directions, including the four basic directions of south, north, east and west, and the three directions of center, upward oblique and downward oblique. For example, the vector continuity in the south, east and upward oblique directions is represented by H S , H E , H UD To express, their definitions strictly follow the mathematical expressions of formulas 1, 2, and 3.
[0065]
[0066] Correspondingly, the continuity of the vector field in the north, west, and downslope directions is represented by H N , H W , H LD These continuity features can be specifically characterized by a method similar to Formula 4. In addition, the continuity of the central area is represented by H L It means that it can be accurately described and calculated by formula 4.
[0067]
[0068] Name the refined vector VR , that is, [V1,V2,V3,V4] T The eight adjacent vectors and the central vector are named V NV =[V NW ,V N ,V NE ,V W ,V L ,V E ,V SW ,V S ,V SE ] T . It can be reconstructed according to Formula 5. In addition, the continuity equations in other directions can also be converted into matrix form, where N S and M S They represent the corresponding coefficient matrices respectively. For details, see Formula 6 and Formula 7.
[0069]
[0070] H represents the sum of the continuity of the vector field in all directions, denoted as Where N and M are coefficient matrices, N = N S +N N +N E +N W +N L +N UD +N LD , M=M S +M N +M E +M W +M L +M UD +M LD To find the optimal value of the continuity matrix, we take the derivative of H and set it to zero, that is, See Formula 8 for details. The final refined vector V R =P×V NV It is obtained by formula 9, where P is the weight matrix, which represents the array of weight coefficients of the center vector. This weighting operation helps to alleviate the local discontinuity problem caused by the error vector.
[0071]
[0072] P=(N T N) -1 N T M (9)
[0073] Improvements to the motion vector self-correction module, such as Figure 5As shown. The algorithm divides a basic block into four sub-blocks of equal size. For each sub-block, the error equation is calculated for the original block and three types of image blocks: horizontal block, vertical block and diagonal block, so as to select the best candidate vector to update the refined vector. By traversing each small image block, a refined vector field is finally obtained. For the error vector related to the small image block, the present invention uses the median filtering result of its 8 adjacent vectors as the corrected output.
[0074] By decoding the target video, continuous frames of the target video are obtained.
[0075] S102 , converting the motion vector into an optical flow map, and further smoothing the optical flow map according to the image frame context information in the continuous frames to obtain a rough optical flow map.
[0076] In one embodiment, the pixels in each block are first filled with the same motion offset to convert the motion vector into a dense optical flow map; secondly, the optical flow map is converted into a coarse optical flow map according to the continuous frames.
[0077] For example, the optical flow map may be converted into a smoother rough optical flow map according to the context information of the previous frame I1 of two adjacent frames in the continuous frames.
[0078] The optical flow map F obtained directly from the motion vector MV There is a big domain difference between F and optical flow, and it cannot be effectively used by the existing deep learning-based optical flow estimation architecture. MV Convert to the same domain as optical flow.
[0079] First, because F MV The sparsity of I1 and its lack in some areas need to be supplemented by other areas. The present invention uses the spatial correlation of I1 to complete F MV Secondly, even in the area where the motion vector offset already exists, the motion estimation may be inaccurate, which is mainly due to the coarse-grained block partitioning strategy or the matching algorithm that does not match the actual motion pattern. These erroneous areas need to be found and corrected. Here, the present invention uses the context information of I1 to solve this problem. Based on the above two points, the specific design of the present invention is as follows Figure 6 As shown, the previous frame I1 of the two adjacent frames in the continuous frame is input into two different encoders to obtain the query matrix Q map and the key matrix K map, and the optical flow map is directly used as the value matrix V map, which is expressed as:
[0080] Q=E A (I1),K=E B (I1),V=F MV
[0081] E A and E BThere are two encoder blocks, each consisting of six convolutional layers and corresponding activation layers. Then, in order to find the area that needs correction, a credibility estimation block is used to estimate the credibility of the motion prior for each pixel, which is specifically expressed by the following formula:
[0082] C MV =CEB(I1,M MV )
[0083] Among them, M MV Indicates the area where the motion vector exists, C MV is a weight map ranging from (0,1). CEB is a convolutional neural network (CNN) block that contains six convolutional layers, six dilated convolutional layers, and corresponding activation layers. The dilated convolutional layers can extract more extensive contextual information and fully utilize spatial information. The final activation function is sigmoid, which is used to limit C MV The correlation calculation is performed in the local sliding window instead of on all pixels to avoid introducing too much extra calculation. First, the correlation S between the central pixel of the local window and other pixels is calculated. i,j :
[0084] S i,j =softmax(Q i,j ·K i+k,j+l )
[0085] Where k,l∈[-d,d], k and l are the offsets of the local window, indicating the position relative to the center pixel, d is the radius of the window, (i,j) is the position of the center pixel of the local window, and dot· indicates the dot product operation. Therefore, the calculated correlation weight S i,j is a shape of (2d+1) 2 Then, the pixel credibility and relevance are combined to get the final weight, which is expressed as:
[0086]
[0087] From C MV A (2d+1)×(2d+1) window is extracted around the center pixel (i,j). The ⊙ symbol represents the element-wise multiplication operator.
[0088] Finally, according to the calculated weight Aggregate the motion within the local window to obtain a rough optical flow map F i,j :
[0089]
[0090] V i+k,j+l Represents the initial optical flow graph FMV The motion vector value at (i+k,j+l).
[0091] Among them, motion vectors and optical flow are both used to describe the movement between frames, but there are two core differences between them. The first difference lies in their level of accuracy: the motion vector is based on image blocks, that is, it focuses on the motion pattern of a larger area in the image; while the optical flow goes deep into the pixel level and can capture and record the motion details of each pixel. Secondly, in terms of processing methods, optical flow tends to use local calculation methods to estimate motion vectors during the encoding process, which means that it focuses more on analyzing the motion relationship between adjacent pixels or areas in the image. In contrast, the calculation of motion vectors usually requires extracting information from the broad context of the entire frame to construct a global motion field model. This process is more global and extensive than optical flow.
[0092] Therefore, using motion vectors as additional input can help optical flow estimation in two ways:
[0093] 1. The optical flow model can be iteratively updated based on the rough solution provided by the motion vector, making the convergence faster.
[0094] 2. Due to the distortion caused by compression, the inter-frame correspondence of some areas is destroyed, so the optical flow model relies more on the learned global priors, such as smoothness, and ignores some small objects that move independently. In contrast, the motion vector stores the best match found for each block separately, which can play an important complementary role in estimating the optical flow of the video.
[0095] S103: Determine the correlation between two adjacent image frames in the continuous frames.
[0096] Specifically, the correlation can be determined by the following steps:
[0097] (1) Feature extraction: Get two consecutive frames as I1 and I2. First, use a convolutional neural network to extract feature maps from the two image frames, denoted as L1 and L2.
[0098] (2) Correlation calculation: For the extracted feature graphs L1 and L2, calculate the correlation between them. Each element C(i,j) of the correlation matrix C represents the similarity between position i in feature L1 and position j in feature L2. The specific formula is as follows:
[0099] C(i,j)=L1(i)·L2(j)
[0100] Where · represents the dot product operation, L1(i) represents the position i in feature L1, and L2(j) represents the position j in feature L2. The calculated correlation matrix C is used to guide the subsequent optical flow estimation, ensuring that the model can fully utilize the correlation information between images in subsequent iterative optimization, thereby improving the accuracy of optical flow estimation.
[0101] S104: Iteratively optimize the rough optical flow map according to the correlation.
[0102] Specifically, the steps of obtaining the coarse-filtered optical flow map and the related relationship and then iterating the optimization are as follows: Figure 7 shown.
[0103] In each iteration, the optical flow estimate can be updated by the following formula:
[0104] ΔF MV =u(F MV (k),C)
[0105] F MV (k+1)=F MV (k)+ΔF MV
[0106] Among them, F MV (k) represents the optical flow estimation value of the kth iteration, ΔF MV Indicates the current optical flow estimate F MV (k) is combined with the correlation matrix C to obtain the optical flow update for this iteration, u(·) a convolutional neural network (CNN) block.
[0107] When the preset number of iterations is reached or the optical flow update amount is less than or equal to the preset update value, the iteration is stopped to obtain the final optical flow map F with high precision. MV (*).
[0108] For example, the conditions for the end of this iteration are as follows:
[0109] (1) Preset number of iterations: The maximum number of iterations N can be preset, and the iteration process stops when the number of iterations reaches N. For example, the maximum number of iterations used in the present invention is 12, and the iteration stops when the number of iterations reaches 12.
[0110] (2) Convergence condition: If the optical flow update amount ΔF MV The change is small in several consecutive iterations, that is, when the optical flow update amount ΔF MV If the value is less than or equal to the preset update value for several consecutive times during the iteration process, it can be considered that the optical flow estimation has converged and the iteration is terminated early.
[0111] Finally, after an iterative optimization process, the model outputs a high-precision optical flow map F MV(*). The optical flow map can accurately represent the pixel-level motion information between video frames, providing more accurate motion estimation results. High-precision optical flow estimation can not only effectively remove the noise introduced by video compression, ensuring the smoothness and accuracy of optical flow estimation, but also improve the visual quality and smoothness of the video.
[0112] The above method is adopted, by converting the motion vector into an optical flow map, and converting the optical flow map into a smoother rough optical flow map through the context information of continuous frames of the image, and then iterating the rough optical flow map as the initial value according to the correlation. In this way, corresponding optical flow processing is performed for different videos, which has high versatility, and the optical flow processing can go deep into the pixel level of the target video, thereby capturing and recording the motion details of each pixel, and has high robustness. Finally, the iteration of each pixel is realized through the optical flow processing, thereby completing the optimization of the target video, and further improving the clarity of the video.
[0113] In another embodiment, after the target video is iteratively optimized, in order to further improve the definition of the target video, the target video may be iteratively optimized. The specific process of the iterative optimization may be as follows:
[0114] S1. Divide each frame of the target video into non-overlapping image blocks, determine the variance mean of the image block difference map at the same position at different time points, and then determine the noise threshold according to the distribution of the variance mean.
[0115] There is a large amount of high-variance noise in surveillance videos under low-light conditions. This noise has obvious changes in time, space, amplitude and phase, and is random, making it difficult to accurately fit it with one or more probability distributions.
[0116] The present invention focuses on surveillance videos without organic displacement. When the average operation is performed on the stationary objects in the video, the noise reduction effect is very significant. This is because denoising can be analogized to reducing the noise variance, where the variance satisfies , where is the noise data in the restricted area and is the variance operator. If it can be understood as averaging 10 adjacent frames of images, the noise variance will be reduced to 1% of the original, and the intensity will be significantly weakened.
[0117] Based on the above data, it is concluded that no matter how the noise data is distributed, it follows two criteria: 1. The amplitude of the noise signal is much smaller than the amplitude of the noise-free signal; 2. In surveillance videos, averaging the noise on stationary objects can significantly remove it. The only disadvantage is that the averaging operation may produce artifacts and ghost images on the trajectory of moving objects.
[0118] In this step, firstly, the pixel-level difference of the image blocks at the same position at different time points can be determined to obtain a difference map; secondly, the distribution histogram of the variance mean of the difference map is determined, and the noise threshold is determined according to the distribution histogram.
[0119] Specifically, the video frame can be first subdivided into non-overlapping image blocks. For image blocks at the same position but at different time points, the pixel-level difference is calculated, the variance mean of the difference map is obtained, and the difference maps are sorted accordingly to obtain a variance mean distribution histogram.
[0120] Blocks with low and similar variance means are identified as stationary blocks, while those with significant differences are considered moving blocks. In stationary blocks, the environmental noise threshold is set by analyzing the variance mean distribution histogram.
[0121] S2. Classify the types of motion vectors according to the noise threshold, and then reconstruct the motion vectors according to different types to obtain trajectory vectors; and filter the target video according to the trajectory vectors.
[0122] In this step, based on the noise threshold, the motion vectors are subdivided into three categories: zero vectors represent static areas, regular vectors correspond to regular motion, and irregular vectors cover abnormal motion such as fast movement, drift, and occlusion.
[0123] The motion vectors of different categories are finely processed to form trajectory vectors. After filtering the noisy video along these trajectories, a clear and denoised video output can be obtained.
[0124] In this way, the noise threshold is determined by the variance mean of the image block difference map, and the types of motion vectors are divided according to the noise threshold. Trajectory vectors are generated according to different types through motion vectors, and then the video is filtered and denoised along the trajectory vector. In this way, different noise thresholds are determined according to different videos, and denoising is performed along the trajectory vector, which not only has strong applicability but also has good denoising effect.
[0125] Secondly, the present invention also provides a video optimization device in a low-light environment, such as Figure 8 As shown, including:
[0126] The acquisition module 801 is used to acquire the motion vector and continuous frames of the target video in a low-light environment.
[0127] The optimization module 802 is used to convert the motion vector into an optical flow map, further smooth the optical flow map according to the image frame context information in the continuous frames to obtain a rough optical flow map; determine the correlation between two adjacent image frames in the continuous frames; and iteratively optimize the rough optical flow map according to the correlation.
[0128] Optionally, the device also includes a denoising module 803, which is used to divide each frame of the target video into non-overlapping image blocks, determine the variance mean of the image block difference map at the same position at different time points, and then determine the noise threshold through the distribution of the variance mean; classify the types of motion vectors according to the noise threshold, and then reconstruct the motion vectors according to different types to obtain trajectory vectors; and filter the target video according to the trajectory vector.
[0129] Optionally, the acquisition module 801 is also used to extract the initial motion vector of the target video through a motion vector initialization network; correct the error of the initial motion vector according to minimizing the neighborhood pixel norm to obtain the motion vector of the target video; and obtain continuous frames of the target video by decoding the target video.
[0130] Optionally, the acquisition module 801 is further used to correct the motion vector field by minimizing the L2 norm between the central image block and the eight adjacent image blocks; the L2 norm includes amplitude and phase, that is:
[0131]
[0132] Where A(c) and P(c) represent the amplitude and phase of the current image block c, respectively; A(n) and P(n) represent the amplitude and phase of the eight adjacent image blocks, (x, y) are spatial coordinates, t is the temporal coordinate, i is the number of the adjacent image block in the neighborhood, ω and λ are weighting coefficients, are the amplitude and phase loss functions of the central image block, respectively.
[0133] Optionally, the denoising module 802 is further used to determine the pixel-level difference of image blocks at the same position at different time points to obtain a difference map; determine a distribution histogram of the variance mean of the difference map, and determine the noise threshold according to the distribution histogram.
[0134] Optionally, the optimization module 803 is further used to fill the pixels in each block with the same motion offset, converting the motion vector into a dense optical flow map; the previous frame of the two adjacent frames in the continuous frames is input into two different encoders to obtain a query matrix Q map and a key matrix K map, and the optical flow map is directly used as a value matrix V map, which is expressed as:
[0135] Q=E A (I1),K=E B (I1),V=F MV
[0136] Among them, E A and E B There are two encoder blocks, each consisting of six convolutional layers and corresponding activation layers; a confidence estimation block is used to estimate the confidence of the motion prior for each pixel, which is expressed by the following formula:
[0137] C MV =CEB(I1,M MV )
[0138] Among them, M MV Indicates the area where the motion vector exists, I1 is the previous frame of the two adjacent frames in the continuous frame, C MV is a weight map ranging from (0,1), CEB is a convolutional neural network (CNN) block; it calculates the correlation S between the center pixel of the local window and other pixels i,j :
[0139] S i,j =softmax(Q i,j ·K i+k,j+l )
[0140] Where k,l∈[-d,d], k and l are the offsets of the local window, indicating the position relative to the center pixel, d is the radius of the window, (i,j) is the position of the center pixel of the local window, and dot· indicates the dot product operation. Therefore, the calculated correlation weight S i,j is a shape of (2d+1) 2 Tensor of; Combine the pixel's credibility and relevance to get the final weight It is expressed as:
[0141]
[0142] From C MV A (2d+1)×(2d+1) window is extracted around the center pixel (i, j), and the ⊙ mark represents the element-by-element multiplication operator; according to the weight Aggregate the motion within the local window to obtain a rough optical flow map F i,j :
[0143]
[0144] V i+k,j+l Represents the initial optical flow graph F MV The motion vector value at (i+k,j+l).
[0145] Optionally, the optimization module 803 is further used to obtain two consecutive frames as I1 and I2, and use a convolutional neural network to extract feature maps from the two image frames respectively, which are recorded as L1 and L2; determine the correlation matrix C of L1 and L2, and the specific formula is as follows:
[0146] C(i,j)=L1(i)·L2(j)
[0147] Among them, C(i,j) is an element of C, which represents the similarity between position i in feature L1 and position j in feature L2, and · represents the dot product operation.
[0148] Optionally, the optimization module 803 is also used to update the optical flow estimation value in each iteration by the following formula:
[0149] ΔF MV =u(F MV (k),C)
[0150] F MV (k+1)=F MV (k)+ΔF MV
[0151] Among them, F MV (k) represents the optical flow estimation value of the kth iteration, ΔF MV Indicates the current optical flow estimate F MV (k) and the correlation matrix C to obtain the optical flow update amount of this iteration, u(·) a convolutional neural network (CNN) block; when the preset number of iterations is reached, or the optical flow update amount is less than or equal to the preset update value, the iteration is stopped to obtain a high-precision optical flow map F MV (*).
[0152] By adopting the above-mentioned device, the motion vector is converted into an optical flow map, and the optical flow map is converted into a smoother rough optical flow map through the context information of continuous frames of the image, and then the rough optical flow map is used as the initial value for iteration according to the correlation. In this way, corresponding optical flow processing is performed for different videos, which has high versatility. Moreover, the optical flow processing can go deep into the pixel level of the target video, and then capture and record the motion details of each pixel point, which has high robustness. Finally, the iteration of each pixel is realized through the optical flow processing, and then the optimization of the target video is completed, and the clarity of the video is further improved.
[0153] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 Provides a video optimization method for low-light environments.
[0154] The present invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Provides a video optimization method for low-light environments.
[0155] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0157] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0159] It should be noted that the above specific implementation method can enable those skilled in the art to more fully understand the invention, but does not limit the invention in any way. Therefore, although the present specification and embodiments have described the invention in detail, those skilled in the art should understand that the invention can still be modified or replaced by equivalents; and all technical solutions and improvements that do not deviate from the spirit and scope of the invention are included in the protection scope of the patent for the invention. Any figure mark in the claims should not be regarded as limiting the claims involved.
Claims
1. A video optimization method in a low-light environment, characterized in that: The method comprises: Get the motion vector and continuous frames of the target video in low-light environment; Converting the motion vector into an optical flow map, and further smoothing the optical flow map according to the image frame context information in the continuous frames to obtain a rough optical flow map; Determining a correlation between two adjacent image frames in the continuous frames; Iteratively optimizing the rough optical flow map according to the correlation; The converting the motion vector into an optical flow map, and further smoothing the optical flow map according to the image frame context information in the continuous frames to obtain a rough optical flow map comprises: Fill the pixels in each block with the same motion offset, converting the motion vectors into a dense optical flow map; Inputting the previous frame of two adjacent frames in the continuous frames into two different encoders to obtain a query matrix Q map and a key matrix K map, and the optical flow map is used as a value matrix V map; Use the confidence estimation block to estimate the confidence of the motion prior for each pixel; Calculate the correlation between the central pixel of the local window and other pixels; Combine the pixel’s credibility and relevance to get the final weight; The motion within the local window is aggregated according to the weights to obtain a rough optical flow map.
2. The video optimization method in a low-light environment according to claim 1, characterized in that: The method further includes filtering the target video after iterative optimization, and the filtering process includes: Divide each frame of the target video into non-overlapping image blocks, determine the variance mean of the image block difference map at the same position at different time points, and then determine the noise threshold based on the distribution of the variance mean; The motion vectors are divided into types according to the noise threshold, and then the motion vectors are reconstructed according to different types to obtain trajectory vectors; and the target video is filtered according to the trajectory vectors.
3. The video optimization method in a low-light environment according to claim 1, characterized in that: Obtaining motion vectors and consecutive frames of target video in low-light environments includes: The initial motion vector of the target video is extracted through the motion vector initialization network; the error of the initial motion vector is corrected according to the minimization of the neighborhood pixel norm to obtain the motion vector of the target video; By decoding the target video, continuous frames of the target video are obtained.
4. The video optimization method in a low-light environment according to claim 3, characterized in that: The method of correcting the error of the initial motion vector according to minimizing the neighborhood pixel norm includes: The motion vector field is corrected by minimizing the L2 norm between the central image block and the eight adjacent image blocks; the L2 norm includes the magnitude and phase, that is: Where A(c) and P(c) represent the amplitude and phase of the current image block c, respectively; A(n) and P(n) represent the amplitude and phase of the eight adjacent image blocks, (x, y) are spatial coordinates, t is the temporal coordinate, i is the number of the adjacent image block in the neighborhood, ω and λ are weighting coefficients, and L A and L P are the amplitude and phase loss functions of the central image block, respectively.
5. The video optimization method in a low-light environment according to claim 1, characterized in that: The query matrix Q graph, key matrix K graph and value matrix V graph are respectively expressed as: Q=E A (I1),K=E B (I1),V=F MV Among them, E A and E B There are two encoder blocks, each consisting of six convolutional layers and corresponding activation layers; The credibility of the motion prior of each pixel is: C MV =CEB(I1,M MV ) Among them, M MV represents the area where the motion vector exists, I1 is the previous frame of the two adjacent frames in the continuous frames, C MV is a weight map ranging from (0,1), and CEB is a convolutional neural network (CNN) block; The correlation S between the central pixel of the local window and other pixels i,j for: S i,j =softmax(Q i,j ·K i+k,j+l ) Where k,l∈[-d,d], k and l are the offsets of the local window, indicating the position relative to the center pixel, d is the radius of the window, (i,j) is the position of the center pixel of the local window, and · represents the dot product operation. Therefore, the calculated correlation weight S i,j is a shape of (2d+1) 2 Tensor of ; The final weight for: From C MV A (2d+1)×(2d+1) window is extracted around the center pixel (i,j), and the ⊙ mark represents the element-by-element multiplication operator; The rough optical flow map F i,j for: V i+k,j+l Represents the initial optical flow graph F MV The motion vector value at (i+k,j+l).
6. The video optimization method in a low-light environment according to claim 1, characterized in that: Determining the correlation between two adjacent image frames in the continuous frames includes: Take two consecutive frames as I1 and I2, use convolutional neural network to extract feature maps from the two image frames respectively, record them as L1 and L2; determine the correlation matrix C of L1 and L2, the specific formula is as follows: C(i,j)=L1(i)·L2(j) Among them, C(i,j) is an element of C, which represents the similarity between position i in feature L1 and position j in feature L2, and · represents the dot product operation.
7. The video optimization method in a low-light environment according to claim 1, characterized in that: The iterative optimization of the rough optical flow map according to the correlation includes: In each iteration, the optical flow estimate can be updated by the following formula: ΔF MV =u(F MV (k),C) F MV (k+1)=F MV (k)+ΔF MV Among them, F MV (k) represents the optical flow estimation value of the kth iteration, ΔF MV Indicates the current optical flow estimate F MV (k) is combined with the correlation matrix C to obtain the optical flow update for this iteration, u(·) a convolutional neural network (CNN) block; When the preset number of iterations is reached or the optical flow update amount is less than or equal to the preset update value, the iteration is stopped to obtain a high-precision optical flow map F MV (*).
8. A video optimization device in a low-light environment, characterized in that: The device comprises: An acquisition module, used to acquire motion vectors and continuous frames of a target video in a low-light environment; An optimization module, configured to convert the motion vector into an optical flow map, further smooth the optical flow map according to the image frame context information in the continuous frames to obtain a rough optical flow map; determine the correlation between two adjacent image frames in the continuous frames; and iteratively optimize the rough optical flow map according to the correlation; The optimization module is also used to fill the pixels in each block with the same motion offset, converting the motion vector into a dense optical flow map; inputting the previous frame of the two adjacent frames in the continuous frames into two different encoders to obtain a query matrix Q map and a key matrix K map, and the optical flow map is used as a value matrix V map; using a credibility estimation block to estimate the credibility of the motion prior of each pixel; calculating the correlation between the central pixel of the local window and other pixels; combining the credibility of the pixel with the correlation to obtain the final weight; and aggregating the motion within the local window according to the weight to obtain a rough optical flow map.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the video optimization method in a low-light environment described in any one of claims 1 to 7 are implemented.
10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the video optimization method in a low-light environment as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Image registration reliability model and reconstruction method of super-resolution image
CN102136144A
Mobile inspection video quality correction method based on significance multi-feature fusion
CN110312124A