Raw domain optical flow prediction method under extremely dark light

By simulating raw data under extremely dark light and generating motion data, and improving the fastflownet network, a lightweight optical flow neural network is designed to solve the problem of inaccurate optical flow under extremely dark light, and high-precision and low-computation optical flow prediction is achieved, meeting real-time requirements and improving image quality.

CN120020875APending Publication Date: 2025-05-20HEFEI JUNZHENG TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311538961.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In extremely dark light scenarios, traditional optical flow methods are low in accuracy and are susceptible to noise. Although the neural network-based methods are high in accuracy, they are large in calculations, which are difficult to use in real time, and are difficult to directly apply to extremely dark raw domain images.

Method used

By simulating raw data under extremely dark light, using motion prospects to generate motion data, and improving the fastflownet network to improve computing efficiency, and designing a lightweight optical flow neural network to generate optical flow data online.

Benefits of technology

The problem of inaccurate optical flow under extremely dark light is solved, high-precision optical flow prediction is achieved, and the calculation amount is small, which meets the real-time requirements, improves the noise level in extremely dark light scenes, and improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020875A_ABST
    Figure CN120020875A_ABST
Patent Text Reader

Abstract

The invention provides a raw domain optical flow prediction method under extremely dark light. The method comprises the following steps: S1, preparing clean raw data and data required by noise modeling; s2, designing motion data; s3, generating optical flow training data; s4, designing an optical flow neural network; and S5, training the network. According to the method, the optical flow data can be generated online, the motion data in the extremely dark light scene is simulated, the problem that the optical flow is inaccurate in the traditional optical flow extremely dark light scene is solved, meanwhile, the light-weight optical flow neural network is designed, the calculated amount is relatively small, and the real-time requirement is met under the condition that the precision is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent monitoring video processing, and particularly relates to a method for predicting optical flow in the raw domain under extremely low light conditions. Background Art

[0002] In the existing intelligent monitoring video processing technology, due to the limited number of photons in extremely low light scenarios, image sensors face serious noise problems. Denoising in the raw domain is of great significance for image signal processing. There are many problems in denoising dynamic scenes under extremely low light, such as ghosting, penetration, loss of motion details, etc. Motion alignment is necessary, and optical flow can complete motion detection and alignment. Among them, dense optical flow can complete pixel-level alignment, which is beneficial to motion scene denoising. Currently, the representative methods of traditional dense optical flow include the Farneback method and the DIS (Dense Inverse Search-based method); currently, the representative methods of dense optical flow based on neural networks mainly include flownet, pwcnet, fastflownet, raft, etc.

[0003] However, the existing traditional optical flow methods have low accuracy and are easily affected by noise, resulting in inaccurate motion detection and alignment. At the same time, the raw domain images are darker, which is more unfavorable for optical flow calculation; while the existing neural network-based methods have relatively high accuracy, but usually have a large amount of calculation and are difficult to use in real time. Moreover, the existing data is mainly clean and bright rgb data, which is difficult to use directly.

[0004] In addition, the commonly used terms in the existing technology include:

[0005] Raw domain: The raw data output by the image sensor is called raw data. Image Signal Processing (ISP) performs signal processing on the raw image data output by the image sensor. After passing through the demosaic interpolation algorithm, it will be transferred to the rgb domain. The domain before being transferred to the rgb domain is called the raw domain.

[0006] Optical flow: Optical flow refers to the instantaneous velocity of pixel motion of a moving object in an image, which is a method for calculating the motion information of an object between adjacent frames. The task of optical flow is to find the corresponding points of the pixel points on the first frame image in the second frame image.

[0007] Sparse optical flow: Sparse optical flow refers to selecting a small number of key points in an image and then calculating the motion speed of these key points.

[0008] Dense optical flow: Dense optical flow refers to calculating the corresponding motion speed for each pixel point in an image. Summary of the Invention

[0009] To solve the above problems, the purpose of this application is as follows: Based on the existing noise modeling method, this application simulates raw data in extremely low light, simultaneously uses the moving foreground to generate motion data in extremely low light, generates optical flow data at the same time, and improves the fastflownet network to enhance the computing efficiency.

[0010] Specifically, the present invention provides a method for predicting optical flow in the raw domain under extremely low light, and the method includes the following steps:

[0011] S1. Prepare clean raw data and the data required for noise modeling. The required data includes a black frame obtained by completely covering the lens with black tape and a flat frame of a flat white paper, and calibrate the noise generation related parameter p. Noise can be randomly generated according to the related parameter, as shown in Equation (1), where I N0 is the noise data, I C is the clean data, G n is the noise generation function, p is the calibrated noise generation parameter, and s is the noise sampling related parameter:

[0012] I N0 = I C + G n (I C , p, s) Equation (1);

[0013] S2. Design motion data:

[0014] S2.1. Prepare moving foreground data. Use the open-source rgb domain human parsing data ATR to obtain the human segmentation mask, and intercept the smallest rectangle of the human segmentation mask as the foreground source data;

[0015] S2.2. Randomly select background data from the clean raw data, and perform inverse white balance on the foreground data prepared in step S2.1 randomly selected from the foreground source data, as shown in Equation (2), where U represents the uniform distribution, rgain and bgain are the gain coefficients of the r channel and the b channel respectively, and I Fr is the r channel data and the b channel data of I Fb foreground image I F :

[0016]

[0017] S2.3. Extract the rgb values at each pixel coordinate of the data after inverse white balance in step S2.2 in bayer format. The target bayer format is rggb. Only the r data in the rgb domain is taken at the target bayer r position, only the g data in the rgb domain is taken at the target bayer g position, and only the b data in the rgb domain is taken at the target bayer b position, and finally obtain the foreground raw data;

[0018] S2.4. Randomly generate the upper-left coordinates (x1, y1) of the first-frame moving rectangle, determine the lower-right coordinates (x2, y2) of the moving rectangle according to the width and height of the foreground data, and use the foreground mask to cover the foreground data at the determined rectangle position. The data of the first-frame rectangle position is as shown in the following formula (3), where I R1 is the data of the first-frame moving rectangle area, I RB1 is the data of the first-frame background rectangle area, I F1 is the data of the first-frame foreground rectangle area.

[0019] I R1 = I RB1 ·(1 - mask)+I F1 ·mask Formula (3);

[0020] S2.5. Replace the data at the corresponding background position with the data of the first-frame moving rectangle area to obtain the first-frame data I 1 ;

[0021] S3. Generate optical flow training data:

[0022] S3.1. Randomly translate the foreground raw data in step S2.3. The number of pixels translated up, down, left, and right is controlled by the following formula (4), where m represents the size of the translated pixels, and U int represents integer uniform sampling;

[0023] m ~ 2·U int (0, 10) Formula (4);

[0024] S3.2. Determine the upper-left coordinates and lower-right coordinates of the second-frame moving rectangle according to the translation size generated in step S3.1, and determine the data of the second-frame moving rectangle area I R2 ;

[0025] S3.3. In the same way as in step S2.5, replace the data at the corresponding background position with the data of the second-frame moving rectangle area to obtain the second-frame data I 2 ;

[0026] S3.4. Generate all-zero data with the width and height dimensions of the second-frame data and 2 channels as the initial optical flow. Take the corresponding position of the second-frame foreground mask in the second-frame moving rectangle, and set the mask areas of the first channel and the second channel of the initial optical flow to the left and right translation pixel sizes and the up and down translation pixel sizes in step S3.1 respectively to generate the optical flow data flow;

[0027] S3.5. Add noise to the first-frame data and the second-frame data according to formula (1) in step S1, where the noise sampling related parameter s is the same to ensure that the generated noise sizes are not very different, and generate the first-frame noise data I n1 and the second-frame noise data In2 , and form paired data with the optical flow data flow generated in step S3.4;

[0028] S4. Design an optical flow neural network:

[0029] S4.1. Design a basic feature extraction module, which uses three consecutive convolutional layers with a convolutional kernel of 3x3, a stride of 2, and a padding of 1 in series. After each convolution, a relu activation function is connected, and then two average pooling operations are used in series. Five basic feature maps of different sizes are saved after each activation function and average pooling. The first frame and the second frame respectively pass through the basic feature extraction module to obtain 10 basic feature maps in total. The basic feature extraction modules of the first frame and the second frame share weights;

[0030] S4.2. Use a cost volume module, set the matching window to a size of 5x5, and perform a dot product on each point in the channel dimension of the features of the first frame and the second frame in the matching window to represent the matching degree;

[0031] S4.3. Design five optical flow prediction modules. Each optical flow prediction module uses five consecutive convolutional layers with a convolutional kernel of 3x3, a stride of 2, and a padding of 1 in series. The first four convolutions are followed by a relu activation function; S4.4. Design an optical flow refinement module. In step 5.1, the smallest basic feature maps of the first frame and the second frame pass through the cost volume module to obtain the smallest matching features. The features of the second frame's smallest basic feature map after convolution with a convolutional kernel of 3x3, a stride of 2, and a padding of 1 are concatenated with the smallest matching features and input into the optical flow prediction module to obtain the optical flow at the smallest scale; The smallest-scale optical flow nearest neighbor upsampling result is used for forward warping of the second-smallest scale basic feature of the first frame. The warping result passes through the cost volume module to obtain the second-smallest matching features. The features of the second frame's second-smallest basic feature map after convolution with a convolutional kernel of 3x3, a stride of 2, and a padding of 1 are concatenated with the second-smallest matching features and input into the optical flow prediction module to obtain the refined result of the optical flow at the second-smallest scale, and then the nearest neighbor upsampling result of the smallest-scale optical flow is added to obtain the final result of the optical flow at the second-smallest scale; The same applies to the refinement of other scale optical flows;

[0032] S4.5. The final optical flow at the largest scale undergoes one nearest neighbor upsampling to obtain the final optical flow prediction result; S5. Use the paired data required for training generated online in steps S2 and S3, and train using the optical flow neural network designed in step S4. The optimizer uses Adam, the initial learning rate is 0.0001, the training cycle is 60, and the learning rate is reduced by 0.5 times every 10 training cycles. The loss is L1 loss.

[0033] The image sensor used in the method is sc2210.

[0034] Therefore, the advantages of this application are as follows: This method can generate optical flow data online, simulate motion data in extremely low-light scenarios, solve the problem of inaccurate optical flow in traditional optical flow for extremely low-light scenarios. At the same time, this patent designs a lightweight optical flow neural network with relatively small computational complexity, meeting the real-time requirement while ensuring accuracy. This method is applicable to, for example, the chip t41 aiisp of Beijing Junzheng Integrated Circuit Co., Ltd. (referred to as: Junzheng), in the field of image signal processing, improving the noise level in extremely low-light scenarios and enhancing the image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.

[0036] Figure 1 It is a schematic diagram of the process of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.

[0038] As Figure 1 shown, this application proposes a method for predicting optical flow in the raw domain under extremely low light. The image sensor used is sc2210. The main implementation steps of the method are as follows:

[0039] Step S1: Prepare clean raw data and data required for noise modeling, calibrate the noise generation related parameter p, and randomly generate noise according to the relevant parameters, as shown in Equation (1), where I N0 is the noise data, I C is the clean data, G n is the noise generation function, p is the calibrated noise generation parameter, and s is the noise sampling related parameter.

[0040] I N0 = I C + G n (I C , p, s) Equation (1)

[0041] Step S2: Design motion data:

[0042] S2.1, Prepare motion foreground data, use the open-source rgb domain human parsing data ATR to obtain the human shape segmentation mask, and intercept the smallest rectangle of the human shape mask as the foreground source data; where ART (Active Template Regression) is an open-source dataset, originating from the paper Deep Human Parsing with Active Template Regression;

[0043] S2.2. Randomly select background data from the clean raw data and perform inverse white balance on the foreground data (the foreground data prepared in step S2.1) randomly selected from the foreground source data. As shown in Equation (2), where U represents a uniform distribution, rgain and bgain are the gain coefficients of the r channel and the b channel respectively, and I Fr is I Fb the r channel data and the b channel data of the foreground image I F .

[0044]

[0045] S2.3. Extract the rgb values at each pixel coordinate of the data after inverse white balance in step S2.2 in bayer format. The target bayer format is rggb. Only the r data in the rgb domain is taken at the target bayer r position, only the g data in the rgb domain is taken at the target bayer g position, and only the b data in the rgb domain is taken at the target bayer b position, finally obtaining the foreground raw data.

[0046] S2.4. Randomly generate the upper left corner coordinates (x1, y1) of the first frame of the moving rectangle, determine the lower right corner coordinates (x2, y2) of the moving rectangle according to the width and height of the foreground data, and use the foreground mask to cover the foreground data at the determined rectangular position. The data of the first frame of the rectangular position is as follows in Equation (3), where I R1 is the data of the first frame of the moving rectangle area, I RB1 is the data of the first frame of the background rectangle area, and I F1 is the data of the first frame of the foreground rectangle area;

[0047] I R1 =I RB1 ·(1 - mask)+I F1 ·mask Equation (3)

[0048] S2.5. Replace the data at the corresponding background position with the data of the first frame of the moving rectangle area to obtain the first frame of data I 1 ;

[0049] Step S3 generates optical flow training data

[0050] S3.1. Randomly translate the foreground raw data in step S2.3. The translation pixels up, down, left, and right are controlled by the sampling in the following Equation (4), where m represents the translation pixel size and U int represents integer uniform sampling;

[0051] m ~ 2·U int (0, 10) Equation (3)

[0052] S3.2. Determine the upper-left and lower-right coordinates of the moving rectangle in the second frame according to the translation size generated in step S3.1, and determine the data I of the moving rectangle area in the second frame in the same way as in step S2.4 R2 ;

[0053] S3.3. In the same way as in step S2.5, replace the data at the corresponding background position with the data of the moving rectangle area in the second frame to obtain the second-frame data I 2 ;

[0054] S3.4. Generate all-zero data with the width and height dimensions of the second-frame data and 2 channels as the initial optical flow. Take the position of the second-frame foreground mask corresponding to the moving rectangle in the second frame, and set the mask areas of the first and second channels of the initial optical flow to the left and right translation pixel sizes and the up and down translation pixel sizes in step S3.1 respectively to generate the optical flow data flow;

[0055] S3.5. Add noise to the first-frame data and the second-frame data according to equation (1) in step S1, where the noise sampling related parameter s is the same to ensure that the generated noise sizes are not very different, and generate the first-frame noise data I n1 and the second-frame noise data I n2 , which together with the optical flow data flow generated in step S3.4 form paired data;

[0056] Step S4 designs an optical flow neural network

[0057] S4.1. Design a basic feature extraction module. Use three consecutive convolutional layers with a convolutional kernel size of 3x3, a stride of 2, and a padding of 1 in series. After each convolution, connect a relu activation function, and then use two average pooling operations in series. Save a total of 5 basic feature maps of different sizes after each activation function and after average pooling. The first frame and the second frame respectively pass through the basic feature extraction module to obtain 10 basic feature maps in total. The basic feature extraction modules of the first frame and the second frame share weights;

[0058] S4.2. Use a cost volume module. Set the matching window size to 5x5. The features of the first frame and the features of the second frame perform a dot product at each point in the matching window in the channel dimension to represent the matching degree;

[0059] S4.3. Design 5 optical flow prediction modules. Each optical flow prediction module consists of a series of 5 consecutive convolutional layers with a convolutional kernel size of 3x3, a stride of 2, and a padding of 1. After the first 4 convolutions, a relu activation function is applied. S4.4. Design of the optical flow refinement module. In step 5.1, the minimum basic feature maps of the first and second frames pass through the cost volume module to obtain the minimum matching features. The feature map of the second frame after convolution with a convolutional kernel size of 3x3, a stride of 2, and a padding of 1 is concatenated with the minimum matching features and input into the optical flow prediction module to obtain the optical flow at the minimum scale. The forward warping is performed on the sub-minimum scale basic feature map of the first frame using the nearest neighbor upsampling result of the minimum scale optical flow. The warping result passes through the cost volume module to obtain the sub-minimum matching features. The feature map of the second frame's sub-minimum basic feature map after convolution with a convolutional kernel size of 3x3, a stride of 2, and a padding of 1 is concatenated with the sub-minimum matching features and input into the optical flow prediction module to obtain the refined result of the optical flow at the sub-minimum scale. Then, adding the nearest neighbor upsampling result of the minimum scale optical flow to obtain the final result of the optical flow at the sub-minimum scale. The refinement of optical flow at other scales is similar.

[0060] S4.5. The final optical flow at the maximum scale undergoes one nearest neighbor upsampling to obtain the final optical flow prediction result.

[0061] In step S5, the paired data required for training is generated online using steps S2 and S3, and the optical flow neural network designed in step S4 is trained. The optimizer used is Adam, with an initial learning rate of 0.0001, a training epoch of 60, and the learning rate is reduced by 0.5 times every 10 training epochs. The loss is L1 loss.

[0062] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting optical flow in raw domain under extremely dark light, characterized in that: The method comprises the following steps: S1, prepare clean raw data and data required for noise modeling. The required data include black frames with black tape completely covering the lens and flat frames with flat white paper. Calibrate the relevant parameters for noise generation, and randomly generate noise according to the relevant parameters. S2, design motion data: S2.1, prepare motion foreground data, use the open source RGB domain human body analysis data ATR to obtain the human segmentation mask, and intercept the minimum rectangle of the human segmentation mask as the foreground source data; S2.2, randomly select background data from the clean raw data, and randomly select the foreground data prepared in step S2.1 from the foreground source data to perform inverse white balance, as shown in formula (2), where U represents uniform distribution, rgain and bgain are the gain coefficients of the r channel and b channel respectively, and I Fr For I Fb Foreground Image I F The r channel data and b channel data: S2.3, extracting the RGB value at each pixel coordinate of the data after inverse white balance in step S2.2 in Bayer format, the target Bayer format is RGB, only the R data in the RGB domain is taken at the target Bayer R position, only the G data in the RGB domain is taken at the target Bayer G position, and only the B data in the RGB domain is taken at the target Bayer B position, and finally obtaining the foreground raw data; S2.4, randomly generate the coordinates of the upper left corner of the moving rectangle of the first frame (x1, y1), determine the coordinates of the lower right corner of the moving rectangle (x2, y2) according to the width and height of the foreground data, and use the foreground mask to cover the foreground data to the determined rectangular position. The rectangular position data of the first frame is as follows (3), where I R1 is the first frame motion rectangular area data, I RB1 is the background rectangular area data of the first frame, I F1 Foreground rectangular area data of the first frame, I R1 = I RB1 ·(1 - mask)+I F1 ·mask Equation (3); S2.5, the first frame moving rectangular area data replaces the corresponding background position data to obtain the first frame data I1; S3, generate optical flow training data: S3.1, randomly shift the foreground raw data in step S2.

3. The pixels shifted up, down, left, and right are sampled and controlled by the following formula (4), where m represents the size of the shifted pixels, U int represents integer uniform sampling; m~2·U int (0,10) Formula (4); S3.2, determine the coordinates of the upper left corner and the lower right corner of the second frame motion rectangle according to the translation size generated in step S3.1, and determine the second frame motion rectangle area data I by the same method as step S2.4 R2 ; S3.3, the same method as step S2.5, the second frame moving rectangular area data replaces the corresponding background position data to obtain the second frame data I2; S3.4, generate all-0 data with width and height dimensions equal to the width and height of the second frame data and a channel number of 2 as the initial optical flow, take the foreground mask of the second frame corresponding to the moving rectangle position of the second frame, set the mask area of ​​the first channel and the second channel of the initial optical flow to the left and right translation pixel size and the up and down translation pixel size of step S3.1 respectively, and generate the optical flow data flow; S3.5, according to the formula (1) in step S1, noise is added to the first frame data and the second frame data, wherein the noise sampling related parameter s is the same to ensure that the difference in the generated noise size is not large, and the first frame noise data I is generated. n1 and the second frame noise data I n2 , and the optical flow data flow generated in step S3.4 form paired data; S4, design optical flow neural network: S4.1, design a basic feature extraction module, use three consecutive convolutions with a kernel of 3x3, a stride of 2, and a pad of 1, connect each convolution with a relu activation function, and then use two average pooling operations in series, save a total of 5 basic feature maps of different sizes after each activation function and after average pooling, the first frame and the second frame are respectively passed through the basic feature extraction module to obtain a total of 10 basic feature maps, and the basic feature extraction module of the first frame and the second frame share weights; S4.2, using the cost volume module, set the matching window size to 5x5, and do a dot product of the features of the first frame and the features of the second frame at each point in the matching window in the channel dimension to indicate the degree of matching; S4.3, design 5 optical flow prediction modules, each of which uses 5 consecutive convolutions with a kernel of 3x3, a stride of 2, and a pad of 1. The first 4 convolutions are followed by a relu activation function; S4.4, optical flow refinement module design, step 5.1 the minimum basic feature map of the first frame and the second frame passes through the cost module to obtain the minimum matching feature, the minimum basic feature map of the second frame passes through the convolution kernel of 3x3, stride of 2, pad of 1, and the concatenation feature is concat with the minimum matching feature, and input into the optical flow prediction module to obtain the minimum scale optical flow; the sub-small scale basic feature of the first frame is forward warped using the minimum scale optical flow nearest neighbor upsampling result, and the warping result passes through the cost module to obtain the sub-small matching feature, the sub-small basic feature map of the second frame passes through the convolution kernel of 3x3, stride of 2, pad of 1, and the concatenation feature is concat with the most sub-matching feature, and input into the optical flow prediction module to obtain the sub-small scale optical flow refinement result, and then adds the minimum scale optical flow nearest neighbor upsampling result to obtain the final result of the sub-small scale optical flow; the refinement of other scale optical flow is similar; S4.5, the final maximum scale optical flow is subjected to a nearest neighbor upsampling to obtain the final optical flow prediction result; S5. Use steps S2 and S3 to generate the paired data required for training online, and use the optical flow neural network designed in step S4 for training.

2. According to the method for predicting optical flow in raw domain under extremely dark light in claim 1, it is characterized in that: In step S1, the calibration noise generation related parameter p is set, and the generated noise is shown in formula (1), where I N0 is the noise data, I C For clean data, G n is the noise generation function, p is the calibrated noise generation parameter, and s is the noise sampling related parameter: I N0 = I C + G n (I C , p, s) Equation (1).

3. The method for predicting optical flow in raw domain under extremely dark light according to claim 1, characterized in that: In step S5, the optimizer uses Adam, the initial learning rate is 0.0001, the training cycle is 60, the learning rate is reduced by 0.5 times every 10 training cycles, and the loss is L1 loss.

4. The method for predicting optical flow in raw domain under extremely dark light according to claim 1, characterized in that: The image sensor used in the method is sc2210.

Citation Information

Cited By

  • Dark light graph optical flow synthesis algorithm based on data enhancement and intelligent test data set

    CN121366181A