Network model method for improving optical flow estimation
By designing a network model, the convolutional encoder and decoder extract inter-frame feature and correlation information, combined with the hollow convolutional coding module, the optical flow is estimated and inter-frame alignment is performed, which solves the problems of moving objects jumping and background penetration after video denoising, and improves image quality.
Patent Information
- Application Number
- CN202311554889.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
In the prior art, after video denoising, moving objects are prone to problems of jumping and background penetration, and optical flow estimation is mainly studied in the RGB domain and the raw domain is less researched.
By designing a network model, the convolutional encoder and decoder extract the feature and correlation information between frames, combined with the hollow convolutional coding module, the optical flow is estimated and the inter-frame alignment is performed to solve the problems of jumping and background penetration of moving objects.
It effectively reduces the noise of moving objects, solves the background penetration problem, and improves the image quality after video denoising, especially in extremely dark light conditions.
Smart Images

Figure CN120020876A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video image processing, and particularly relates to a method for improving an optical flow estimation network model. Background Art
[0002] In the prior art, in the video denoising task, especially when the ambient light is 0.01 lux, due to the position change of the moving object, after the time-domain denoising model performs denoising, there is still a lot of noise remaining on the moving object, and the moving object becomes blurred as it moves over time and penetrates the occluded background, resulting in very poor image quality. By using the optical flow estimation method, the front and rear frames can be aligned, and then the images can be fused through the fusion method to ensure the clarity of the images. Currently, there are more studies on optical flow estimation in the RGB domain, while there are fewer related studies on optical flow estimation in the RAW domain.
[0003] Currently, optical flow estimation methods are all studied for RGB images and mainly fall into two categories:
[0004] 1. Traditional optical flow estimation. For example, the reference frame and the frame to be aligned are downsampled and decomposed into feature maps of different scales, and then offsets are calculated using 8x8 or other sized blocks on the feature maps of different scales.
[0005] 2. Deep learning-based optical flow estimation methods. With the development of deep learning, training a network model using an optical flow dataset can fit the motion of small displacements and large displacements. Especially for the FlowNet network, different scales of features of different images are extracted using convolutional modules, and then the correlation between features is calculated at each scale and a rough optical flow is predicted. Finally, the predicted optical flow is upsampled, and the correlation between features is continued to be calculated at the upper scale to predict a more accurate optical flow. The recent method, the RAFT network, estimates optical flow better than the FlowNet network, that is, the optical flow is continuously iterated and updated through a recurrent network structure (GRU).
[0006] However, in the prior art, after video denoising, the moving object and the background penetrate each other. In addition, the edges of the moving object will jitter and become blurred after denoising.
[0007] In addition, the commonly used terms in the prior art include:
[0008] Optical flow estimation: Optical flow refers to the motion of each pixel in the image over time. The goal of optical flow estimation is to calculate the motion vector of each pixel between two frames based on the image information between consecutive frames. lux: The unit of illumination, used to measure the intensity of the current illumination.
[0009] Scale: The sizes of the feature maps are different. Summary of the Invention
[0010] To solve the above problems, the object of the present application is as follows: By using a network model to estimate the motion displacement of an inter-frame moving object, the previous frame image can be mapped to the current image, that is, the moving objects in the front and back frames coincide and then are fused, thereby solving the problems of jumping and penetration of the moving object. To improve the problems of jumping of moving objects and background penetration after video denoising, the present method uses a method of network-estimated optical flow to perform inter-frame alignment on moving objects that change over time. This optical flow estimation can be used in the time-domain denoising part of the raw domain. Since the positions of moving objects are different in the front and back frames, serious penetration problems will occur to the moving objects after time-domain denoising. However, this estimation method can estimate the displacement of the moving objects in the front and back frames so as to align and fuse the moving objects, which not only reduces the noise of the moving objects but also solves the penetration problem.
[0011] Specifically, the present invention provides a network model method for improving optical flow estimation, and the method includes the following steps:
[0012] S1. The input of the network is 2 frames of raw images, denoted as raw1 and raw2 respectively. First, the single-channel raw image is decomposed into a 4-channel image according to the rggb color, and the 4 channels are the r color channel, two g color channels, and one b color channel respectively. Convolutional encoders are used to extract features of two different scales, denoted as f1 and f2 respectively, as shown in formulas (1) and (2);
[0013] f1 = encoder(raw1) (1)
[0014] f2 = encoder(raw2) (2)
[0015] In the formulas, the encoder is a feature encoder composed of convolutions. Through convolution and downsampling, feature maps f1 and f2 at different scales can be obtained. f1 and f2 are the feature maps at the first and second scales respectively. The input of the model can only be two frames, so the encoder in the model can only encode two frames of images;
[0016] S2. To associate the two frames of images and extract the correlation information corr, corr is used to represent the correlation information, and this correlation information is obtained through the calculation between the pixels of the two frames of images. The result is obtained by calculating the pixels near f1 and f2 at different scales, and the formula is as shown in (3).
[0017] corr(x1, x2) = ∑(f1(x1 + o), f2(x2 + o)) (3)
[0018] In the formula, o is the pixel range around the pixel x1, which can be expressed as the pixels within the range of o ∈ [-k, k] × [-k, k];
[0019] S3. After obtaining the correlation information corr at each scale, the corresponding rough optical flow can be obtained by decoding this information through a decoder constructed by convolution; and in order to obtain a more accurate optical flow, the optical flow obtained by the decoder module is combined with the optical flow finely tuned using the dilated convolution encoding module to obtain a finer optical flow, as specifically described in Equation (4).
[0020] flow n = up(decoder n (corr)+context block (corr))+up(flow n-1 ) (4)
[0021] In the formula, n represents the optical flow calculated at the nth layer feature scale; decoder and context block respectively represent the decoder and the network module constructed by dilated convolution; and up represents the upsampling of pixels to ensure that the size of the upsampled flow is the same as the convolution scale of the previous layer.
[0022] In the step S1, the encoder structure of the feature encoder encoder further includes:
[0023] Raw image input;
[0024] Convolution, with a kernel size of 3 and a stride of 2;
[0025] Output the feature map f1 of the first feature scale as the input of the next layer;
[0026] Convolution, with a kernel size of 3 and a stride of 1;
[0027] Output the feature map f2 of the second feature scale as the input of the next layer;
[0028] Convolution, with a kernel size of 3 and a stride of 2;
[0029] Output the feature map f3 of the third feature scale as the input of the next layer;
[0030] Convolution, with a kernel size of 3 and a stride of 2;
[0031] Output the feature map f4 of the fourth feature scale as the input of the next layer;
[0032] Convolution, with a kernel size of 3 and a stride of 2;
[0033] Output the feature map f5 of the fifth feature scale as the input of the next layer;
[0034] Convolution, with a kernel size of 3 and a stride of 2;
[0035] And so on...
[0036] Output the feature map fn of the nth feature scale.
[0037] In the step S1, let n = 5, that is, only output the feature maps of 5 feature scales.
[0038] In the step S3, the network module constructed by the decoder structure decoder and the dilated convolution further includes:
[0039] S3.1, the optical flow encoded through the decoder;
[0040] Feature input;
[0041] Convolution, with the number of channels being 64, the convolution kernel size being 3, and the stride being 1;
[0042] Convolution, with the number of channels being 64, the convolution kernel size being 3, and the stride being 1;
[0043] Convolution, with the number of channels being 64, the convolution kernel size being 3, and the stride being 1;
[0044] The outputs are respectively subjected to step A and step S3.2:
[0045] A, Convolution, with the number of channels being 2, the convolution kernel size being 3, and the stride being 1; Output to obtain the rough optical flow 1;
[0046] S3.2, the optical flow encoded and finely tuned through the context module using dilated convolution;
[0047] Convolution, with the number of channels being 64, the convolution kernel size being 3, and the stride being 1;
[0048] Convolution, with the number of channels being 64, the convolution kernel size being 3, and the stride being 2;
[0049] Convolution, with the number of channels being 64, the convolution kernel size being 3, and the stride being 4;
[0050] Convolution, with the number of channels being 2, the convolution kernel size being 3, and the stride being 1;
[0051] Output to obtain the rough optical flow 2;
[0052] S3.3, according to the addition operation of the results of step S3.1 and step S3.2, obtain the finely tuned result optical flow flow.
[0053] The method further includes step S4. Through step S3, when n in this application is set to 5, the finally estimated optical flow is represented by formula (5),
[0054] flow final = decoder(corr) (5).
[0055] The method is applicable to the network structure for estimating optical flow in the raw domain.
[0056] Therefore, the advantages of this application are as follows: In order to improve the jitter and background penetration problems of moving objects after video denoising, this method uses a network to estimate optical flow. This network is a model constructed by convolutions trained with a large amount of optical flow data and is applicable to the Junzheng T41 AIISP and other chip models. It improves the problems of jitter and penetration of moving objects due to excessive noise after time-domain and spatial-domain denoising under extremely low-light conditions, thereby improving the image quality after denoising. This method performs inter-frame alignment on moving objects that change over time. Description of the Drawings
[0057] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0058] Figure 1 It is a schematic flowchart of the method of this application.
[0059] Figure 2 It is a schematic diagram of the encoder structure of the feature encoding module in the model construction part of this application.
[0060] Figure 3 It is a schematic diagram of the decoder structure of the feature decoding module in the model construction part of this application. Detailed Embodiments
[0061] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.
[0062] This method belongs to the image alignment task. By using an optical flow estimation model in the raw domain to estimate the displacement vector of a moving object during a continuous time period, the problem of moving object alignment is solved. This method is applicable to the optical flow estimation network structure after low-light enhancement in the raw domain, such as Figure 1 As shown, this application proposes a network model method for improving optical flow estimation. The specific implementation steps are as follows: S1. The input of the network is two frames of raw images, denoted as raw1 and raw2 respectively. First, the single-channel raw image is decomposed into a 4-channel image through rggb, and convolutional encoders are used to extract features of two different scales from the two frames, denoted as f1 and f2 respectively, as shown in formulas (1) and (2);
[0063] f1 = encoder(raw1) (1)
[0064] f2 = encoder(raw2) (2)
[0065] In the formula, the encoder is a feature encoder composed of convolutions. Through convolution and downsampling, feature maps at different scales can be obtained, f1 and f2. f1 and f2 are the feature maps at the first and second scales respectively. The encoder structure of the feature encoder encoder is as follows Figure 2 shown as
[0066] Raw image input;
[0067] Convolution, with a convolution kernel size of 3 and a stride of 2;
[0068] Output the feature map f1 at the first feature scale as the input for the next layer;
[0069] Convolution, with a convolution kernel size of 3 and a stride of 1;
[0070] Output the feature map f2 at the second feature scale as the input for the next layer;
[0071] Convolution, with a convolution kernel size of 3 and a stride of 2;
[0072] Output the feature map f3 at the third feature scale as the input for the next layer;
[0073] Convolution, with a convolution kernel size of 3 and a stride of 2;
[0074] Output the feature map f4 at the fourth feature scale as the input for the next layer;
[0075] Convolution, with a convolution kernel size of 3 and a stride of 2;
[0076] Output the feature map f5 at the fifth feature scale as the input for the next layer;
[0077] Convolution, with a convolution kernel size of 3 and a stride of 2;
[0078] And so on...
[0079] Output the feature map fn at the nth feature scale.
[0080] In step S1, let n = 5, that is, only output feature maps at 5 feature scales.
[0081] S2. In order to correlate two frames of images and extract the correlation information corr, the result is obtained by calculating the pixels near f1 and f2 at different scales. The formula is as shown in (3).
[0082] corr(x1, x2) = ∑(f1(x1 + o), f2(x2 + o)) (3)
[0083] In the formula, o is the pixel range around pixel x1, which can be expressed as the pixels within the range of o ∈ [-k, k] × [-k, k].
[0084] S3. After obtaining the correlation information at each scale, the corresponding rough optical flow can be obtained by decoding this information through a decoder constructed by convolution; in order to obtain a more accurate optical flow, the present application combines the encoded optical flow with the optical flow obtained by fine-tuning using dilated convolution encoding to obtain a finer optical flow, as specifically described in Equation (4).
[0085] flow n = up(decoder n (corr)+context block (corr))+up(flow n-1 ) (4)
[0086] In the formula, n represents the optical flow calculated at the nth layer feature scale; Decoder and context_block respectively represent the decoder and the network module constructed by dilated convolution. And up represents the upsampling of pixels to ensure that the size of the upsampled flow is the same as the previous layer convolution scale.
[0087] The decoder structure decoder and the network module constructed by dilated convolution are as Figure 3 shown:[[]]
[0088] S3.1, through the decoder, the encoded optical flow is obtained;
[0089] Feature input;
[0090] Convolution, with 64 channels, a convolution kernel size of 3, and a stride of 1;
[0091] Convolution, with 64 channels, a convolution kernel size of 3, and a stride of 1;
[0092] Convolution, with 64 channels, a convolution kernel size of 3, and a stride of 1;
[0093] The outputs are respectively subjected to step A and step S3.2:
[0094] A, Convolution, with 2 channels, a convolution kernel size of 3, and a stride of 1; the rough optical flow 1 is obtained as the output;
[0095] S3.2, through the context module, the optical flow obtained by fine-tuning using dilated convolution encoding is obtained;
[0096] Convolution, with 64 channels, a convolution kernel size of 3, and a stride of 1;
[0097] Convolution, with 64 channels, a convolution kernel size of 3, and a stride of 2;
[0098] Convolution, with 64 channels, a convolution kernel size of 3, and a stride of 4;
[0099] Convolution, with 2 channels, a convolution kernel size of 3, and a stride of 1;
[0100] Output to obtain rough optical flow 2;
[0101] S3.3. According to the addition operation of the results of steps S3.1 and S3.2, fine-tune the optical flow of flow.
[0102] S4. Through step S3, if n in this application is set to 5, the finally estimated optical flow is represented by formula (5),
[0103] flow final = decoder(corr) (5).
[0104] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A network model method for improving optical flow estimation, characterized in that: The method comprises the following steps: S1. The input of the network is 2 frames of raw images, denoted as raw1 and raw2. First, the single-channel raw image is decomposed into a 4-channel image according to rggb color. The 4 channels are r color channel, two g color channels and one b color channel. The convolution encoder is used to extract the features of the two frames at different scales, denoted as f1 and f2, respectively, as shown in formulas (1) and (2). f1=encoder(raw1) (1) f2=encoder(raw2) (2) The encoder in the formula is a feature encoder composed of convolution. Through convolution and downsampling, feature maps f1 and f2 at different scales can be obtained. f1 and f2 are feature maps at the first and second scales respectively. The input of the model can only be two frames, so the encoder in the model can only encode two frames of images. S2. In order to associate the two frames of images and extract the correlation information corr, corr is used to represent the correlation information, and the correlation information is obtained by calculating the pixels between the two frames of images. The results are obtained by calculating the pixels near f1 and f2 at different scales. The formula is shown in (3). corr(x1,x2)=Σ(f1(x1+o),f2(x2+o) (3) Where o is the pixel range around pixel x1, which can be expressed as pixels within the range of o∈[-k, k]×[-k, k]; S3. After obtaining the correlation information corr at each scale, the decoder constructed by convolution decodes the information to obtain the corresponding rough optical flow. In order to obtain a more accurate optical flow, the optical flow obtained by the decoder module is combined with the optical flow obtained by fine-tuning the hole convolutional coding module to obtain a more refined optical flow, as described in formula (4). flow n =up(decoder n (corr)+context block (corr))+up(flow n-1 ) (4) Where n represents the optical flow calculated at the feature scale of the nth layer; decoder and context block They represent the network modules constructed by the decoder and the dilated convolution respectively; and up represents the upsampling of pixels to ensure that the size of the upsampled stream is the same as the scale of the previous convolution layer.
2. A network model method for improving optical flow estimation according to claim 1, characterized in that: In the step S1, the encoder structure of the feature encoder encoder further includes: Raw image input; Convolution, with kernel size of 3 and stride of 2; Output the feature map f1 of the first feature scale as the input of the next layer; Convolution, the convolution kernel size is 3 and the stride is 1; Output the feature map f2 of the second feature scale as the input of the next layer; Convolution, with kernel size of 3 and stride of 2; Output the feature map f3 of the third feature scale as the input of the next layer; Convolution, with kernel size of 3 and stride of 2; Output the feature map f4 of the fourth feature scale as the input of the next layer; Convolution, with kernel size of 3 and stride of 2; Output the feature map f5 of the fifth feature scale as the input of the next layer; Convolution, with kernel size of 3 and stride of 2; And so on… Output the feature map fn of the nth feature scale.
3. A network model method for improving optical flow estimation according to claim 2, characterized in that: In the step S1, it is assumed that n=5, that is, only feature maps of 5 feature scales are output.
4. A network model method for improving optical flow estimation according to claim 1, characterized in that: In the step S3, the decoder structure decoder and the network module constructed by the dilated convolution further include: S3.1, optical flow obtained after decoding; Feature input; Convolution, the number of channels is 64, the convolution kernel size is 3, and the stride is 1; Convolution, the number of channels is 64, the convolution kernel size is 3, and the stride is 1; Convolution, the number of channels is 64, the convolution kernel size is 3, and the stride is 1; The outputs are respectively processed in step A and step S3.2: A, convolution, the number of channels is 2, the convolution kernel size is 3, and the stride is 1; the output is a rough optical flow of 1; S3.2, optical flow obtained by fine-tuning with dilated convolutional coding after the context module; Convolution, the number of channels is 64, the convolution kernel size is 3, and the stride is 1; Convolution, the number of channels is 64, the convolution kernel size is 3, and the stride is 2; Convolution, the number of channels is 64, the convolution kernel size is 3, and the stride is 4; Convolution, the number of channels is 2, the convolution kernel size is 3, and the stride is 1; The output is rough optical flow 2; S3.3, according to the addition operation of the results of step S3.1 and step S3.2, the fine-tuned result optical flow flow is obtained.
5. A network model method for improving optical flow estimation according to claim 1, characterized in that: The method further includes step S4. Through step S3, in the present application, n is set to 5, and the final estimated optical flow is expressed by formula (5), flow final =decoder(corr) (5)。 6. A network model method for improving optical flow estimation according to claim 1, characterized in that: The method is applicable to the network structure of estimating optical flow in the raw domain.