Optical flow information prediction network training method, image enhancement method, device and equipment
By using noise-free image frames in the optical flow information prediction network for training and mapping true image frames based on the predicted optical flow information, the problem of low learning accuracy of optical flow information in the prior art is solved, and higher optical flow information accuracy and more accurate network parameter adjustment are achieved.
Patent Information
- Application Number
- CN202510108780.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the signal-to-noise ratio between the sample image frame and the reference image frame is low, resulting in a low accuracy and a greater impact on noise differences when learning optical flow information.
By acquiring the first sample image frame and the second sample image frame with noise, input it to the optical flow information prediction network of the initial structure, obtain the optical flow information, and map the true image frame based on the optical flow information, adjust the network parameters until the preset convergence conditions are met, and the trained optical flow information prediction network is obtained.
The accuracy of optical flow information is improved, and the adverse effects of noise differences on network parameter adjustment are reduced, so that the deep learning model can more realistically identify optical flow information between image frames.
Smart Images

Figure CN119991465A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an optical flow information prediction network training method, image enhancement method, device and equipment. Background Art
[0002] The camera can capture images and generate a video image containing multiple image frames. In order to improve the signal-to-noise ratio of an image frame in the video image, the optical flow (optical flow for short) information between the image frame and the adjacent image frames (also called optical flow estimation) can be used to align the adjacent image frames with the image frame, and the aligned adjacent image frames are superimposed with the image frame to obtain an enhanced image frame.
[0003] In the prior art, the sample image frame and the corresponding reference image frame can be input into the deep learning model to be trained to obtain the optical flow information output by the deep learning model; then the reference image frame is aligned using the optical flow information to obtain the aligned reference image frame; the aligned reference image frame is fused with the sample image frame to obtain the enhanced sample image frame, and the network parameters of the deep learning model are adjusted based on the difference between the enhanced sample image frame and the ground truth (GT) image frame. The signal-to-noise ratio of the ground truth image frame is greater than that of the sample image frame and the reference image frame, and the ground truth image frame and the sample image frame have the same content.
[0004] However, the signal-to-noise ratio of the sample image frame and the reference image frame is lower than that of the true image frame. The difference between the enhanced sample image frame and the true image frame also reflects the noise difference between the enhanced sample image frame and the true image frame. Therefore, the noise difference will also be regarded as the difference caused by the optical flow information to adjust the network parameters, resulting in lower accuracy of the optical flow information learned by the deep learning model. Summary of the invention
[0005] The purpose of the embodiments of the present invention is to provide an optical flow information prediction network training method, image enhancement method, device and equipment to improve the accuracy of optical flow information. The specific technical solution is as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for training an optical flow information prediction network, the method comprising:
[0007] Acquire a first sample image frame and a second sample image frame with noise; wherein the first sample image frame and the second sample image frame are two image frames in the first sample video with an interval less than a first number;
[0008] Inputting the first sample image frame and the second sample image frame into an optical flow information prediction network of an initial structure to obtain optical flow information of the second sample image frame relative to the first sample image frame;
[0009] Mapping the true value image frame corresponding to the second sample image frame based on the obtained optical flow information to obtain a mapping result of the true value image frame corresponding to the second sample image frame; wherein the true value image frame corresponding to the second sample image frame has the same image content as the second sample image frame and has no noise;
[0010] Based on the difference between the obtained mapping result and the true value image frame corresponding to the first sample image frame, the network parameters of the optical flow information prediction network of the initial structure are adjusted until the preset convergence conditions are reached, thereby obtaining a trained optical flow information prediction network; wherein the true value image frame corresponding to the first sample image frame is consistent with the image content of the first sample image frame and does not contain noise.
[0011] Optionally, the acquiring of the first sample image frame and the second sample image frame with noise includes:
[0012] Acquire a first original image frame and a second original image frame without noise; wherein the first original image frame and the second original image frame are two image frames in the second sample video with an interval less than a first number;
[0013] Performing noise adding processing on the first original image frame and the second original image frame to obtain a first sample image frame and a second sample image frame with noise;
[0014] The first original image frame is a true value image frame corresponding to the first sample image frame, and the second original image frame is a true value image frame corresponding to the second sample image frame.
[0015] Optionally, the performing noise adding processing on the first original image frame and the second original image frame to obtain the first sample image frame and the second sample image frame with noise includes:
[0016] Dividing the pixel values of the first original image frame and the second original image frame by a specified value respectively to obtain a third original image frame and a fourth original image frame;
[0017] Noise is added to the third original image frame and the fourth original image frame respectively to obtain a first sample image frame and a second sample image frame.
[0018] Optionally, before adding noise to the third original image frame and the fourth original image frame respectively to obtain the first sample image frame and the second sample image frame, the method further includes:
[0019] Acquiring a noise parameter of a sensor that acquires the second sample video;
[0020] The adding noise to the third original image frame and the fourth original image frame respectively to obtain the first sample image frame and the second sample image frame comprises:
[0021] Based on the noise parameter, noise is added to the third original image frame and the fourth original image frame respectively to obtain a first sample image frame and a second sample image frame.
[0022] Optionally, adding noise to the third original image frame and the fourth original image frame respectively based on the noise parameter to obtain a first sample image frame and a second sample image frame includes:
[0023] In a case where the noise parameter includes a mean value of Poisson noise, adding Poisson noise to the third original image frame and the fourth original image frame respectively according to the noise parameter to obtain a first sample image frame and a second sample image frame;
[0024] In a case where the noise parameters include a mean value and a standard deviation of Gaussian noise, Gaussian noise is added to the third original image frame and the fourth original image frame respectively according to the noise parameters to obtain a first sample image frame and a second sample image frame;
[0025] When the noise parameters include the mean of Poisson noise and the mean and standard deviation of Gaussian noise, Poisson noise and Gaussian noise are added to the third original image frame and the fourth original image frame respectively according to the noise parameters to obtain a first sample image frame and a second sample image frame.
[0026] Optionally, the optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork and an output subnetwork;
[0027] Inputting the first sample image frame and the second sample image frame into an optical flow information prediction network of an initial structure to obtain optical flow information of the second sample image frame relative to the first sample image frame includes:
[0028] splicing the first sample image frame and the second sample image frame to obtain a first splicing result;
[0029] Inputting the first concatenation result into the input subnetwork to obtain a first feature vector;
[0030] Inputting the first feature vector into the feature extraction subnetwork to obtain a second feature vector;
[0031] The second feature vector is input into the output sub-network to obtain the optical flow information of the second sample image frame relative to the first sample image frame.
[0032] Optionally, the output subnetwork includes: an activation layer, an amplification layer, and an upsampling layer;
[0033] Inputting the second feature vector into the output subnetwork to obtain optical flow information of the second sample image frame relative to the first sample image frame includes:
[0034] Inputting the second feature vector into the activation layer to obtain first optical flow information to be used;
[0035] Inputting the first optical flow information to be used into the amplification layer to obtain amplified first optical flow information to be used;
[0036] The amplified first optical flow information to be used is input into the upsampling layer to obtain the optical flow information of the second sample image frame relative to the first sample image frame.
[0037] In a second aspect, an embodiment of the present invention provides an image enhancement method, the method comprising:
[0038] Acquire an image frame to be enhanced and a reference image frame; wherein the image frame to be enhanced and the reference image frame are two image frames in the same video whose interval is less than a second number;
[0039] Inputting the image frame to be enhanced and the reference image frame into a pre-trained optical flow information prediction network to obtain optical flow information of the reference image frame relative to the image frame to be enhanced; wherein the optical flow information prediction network is trained based on the above optical flow information prediction network training method;
[0040] Mapping the reference image frame based on the obtained optical flow information to obtain a mapping result of the reference image frame;
[0041] The mapping result of the reference image frame is fused with the image frame to be enhanced to obtain an enhanced image frame of the image frame to be enhanced.
[0042] Optionally, the optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork and an output subnetwork;
[0043] Inputting the image frame to be enhanced and the reference image frame into a pre-trained optical flow information prediction network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced, including:
[0044] Splicing the image frame to be enhanced and the reference image frame to obtain a second splicing result;
[0045] Inputting the second concatenation result into the input subnetwork to obtain a third feature vector;
[0046] Inputting the third feature vector into the feature extraction subnetwork to obtain a fourth feature vector;
[0047] The fourth feature vector is input into the output sub-network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced.
[0048] Optionally, the output subnetwork includes: an activation layer, an amplification layer, and an upsampling layer;
[0049] Inputting the fourth feature vector into the output sub-network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced includes:
[0050] Inputting the fourth eigenvector into the activation layer to obtain second optical flow information to be used;
[0051] Inputting the second optical flow information to be used into the amplification layer to obtain amplified second optical flow information to be used;
[0052] The amplified second optical flow information to be used is input into the upsampling layer to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced.
[0053] In a third aspect, an embodiment of the present invention provides an optical flow information prediction network training device, the device comprising:
[0054] A first acquisition module is used to acquire a first sample image frame and a second sample image frame with noise; wherein the first sample image frame and the second sample image frame are two image frames in the first sample video with an interval less than a first number;
[0055] A first input module, used for inputting the first sample image frame and the second sample image frame into an optical flow information prediction network of an initial structure to obtain optical flow information of the second sample image frame relative to the first sample image frame;
[0056] a first mapping module, configured to map the true value image frame corresponding to the second sample image frame based on the obtained optical flow information, to obtain a mapping result of the true value image frame corresponding to the second sample image frame; wherein the true value image frame corresponding to the second sample image frame has the same image content as the second sample image frame and has no noise;
[0057] An adjustment module is used to adjust the network parameters of the optical flow information prediction network of the initial structure based on the difference between the obtained mapping result and the true value image frame corresponding to the first sample image frame, until a preset convergence condition is reached to obtain a trained optical flow information prediction network; wherein the true value image frame corresponding to the first sample image frame is consistent with the image content of the first sample image frame and does not contain noise.
[0058] In a fourth aspect, an embodiment of the present invention provides an image enhancement device, the device comprising:
[0059] A second acquisition module is used to acquire an image frame to be enhanced and a reference image frame; wherein the image frame to be enhanced and the reference image frame are two image frames in the same video with an interval less than a second number;
[0060] A second input module is used to input the image frame to be enhanced and the reference image frame into a pre-trained optical flow information prediction network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced; wherein the optical flow information prediction network is trained based on the above optical flow information prediction network training device;
[0061] A second mapping module, used for mapping the reference image frame based on the obtained optical flow information to obtain a mapping result of the reference image frame;
[0062] The fusion module is used to fuse the mapping result of the reference image frame with the image frame to be enhanced to obtain an enhanced image frame of the image frame to be enhanced.
[0063] An embodiment of the present invention further provides an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0064] Memory, used to store computer programs;
[0065] The processor is used to implement the above-mentioned optical flow information prediction network training method and image enhancement method when executing the program stored in the memory.
[0066] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the above-mentioned optical flow information prediction network training method and image enhancement method are implemented.
[0067] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the above-mentioned optical flow information prediction network training method and image enhancement method when executed by a processor.
[0068] Beneficial effects of the embodiments of the present invention:
[0069] In the embodiment of the present invention, the first sample image frame and the second sample image frame are noisy, and the optical flow information (predicted optical flow information) of the second sample image frame output by the optical flow information prediction network is used to map the true value image frame corresponding to the second sample image frame, that is, the image frame without noise is mapped according to the predicted optical flow information to obtain a mapping result without noise, and the mapping of the second sample image frame with noise is avoided, so that the loss value between the mapping result and the true value image frame (image frame without noise) of the first sample image frame can reflect the gap between the predicted optical flow information and the true optical flow information (optical flow information of the true value image frame corresponding to the second sample image frame relative to the true value image frame corresponding to the first sample image frame) as realistically as possible, and reduce the adverse effects of noise differences on adjusting network parameters. In other words, the trained optical flow information prediction network can identify the optical flow information between the two image frames as realistically as possible, and improve the accuracy of the learned optical flow information.
[0070] Of course, it is not necessary to achieve all of the advantages described above at the same time to implement any product or method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.
[0072] Figure 1 A schematic diagram of a flow chart of a first optical flow information prediction network training method provided by an embodiment of the present invention;
[0073] Figure 2 A schematic diagram of a flow chart of a second optical flow information prediction network training method provided by an embodiment of the present invention;
[0074] Figure 3 A schematic diagram of a flow chart of a third optical flow information prediction network training method provided by an embodiment of the present invention;
[0075] Figure 4 A schematic diagram of a flow chart of a fourth optical flow information prediction network training method provided by an embodiment of the present invention;
[0076] Figure 5 A schematic diagram of a flow chart of a first image enhancement method provided by an embodiment of the present invention;
[0077] Figure 6 A schematic diagram of a flow chart of a second image enhancement method provided by an embodiment of the present invention;
[0078] Figure 7 A schematic diagram of the principle of an optical flow information prediction network training method provided by an embodiment of the present invention;
[0079] Figure 8 A schematic diagram of the structure of an optical flow information prediction network provided by an embodiment of the present invention;
[0080] FIG9 (a), FIG9 (b), and FIG9 (c) are schematic diagrams showing the effect of an image enhancement method provided by an embodiment of the present invention;
[0081] Fig.10 A schematic diagram of the structure of an optical flow information prediction network training device provided by an embodiment of the present invention;
[0082] Fig.11 A schematic diagram of the structure of an image enhancement device provided by an embodiment of the present invention;
[0083] Fig.12 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0084] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field based on the present invention belong to the scope of protection of the present invention.
[0085] In order to more clearly understand the embodiments of the present invention, the prior art will be described in detail below.
[0086] In low-light environments at night, images captured by cameras often contain significant noise, which affects visual perception and further restricts subsequent image analysis. For example, the recognition accuracy of pedestrians or vehicles in noisy images is low. In engineering, the original image in the video is usually superimposed with adjacent images to enhance the original image in the time domain and improve the signal-to-noise ratio. However, if there are moving objects in the image, such as walking people, moving vehicles, or rotating camera images, directly superimposing the original image with the adjacent image will introduce motion blur. In order to avoid introducing motion blur, motion alignment (also known as motion compensation) technology can be used. Consider each area in the reference image frame, find similar or matching areas in the adjacent image frame (also known as alignment frame) with the reference image frame, determine the optical flow information based on the corresponding relationship between the matching areas of the reference image frame and the adjacent image frame, use the optical flow information to eliminate the motion blur between the original image and the adjacent image, and superimpose the two frames after eliminating motion blur to reduce noise, improve image details and clarity, and improve image enhancement effect.
[0087] In the existing technology, the learning method based on neural network can easily realize the deployment of the model and improve the computing efficiency by processing the data through the modern commonly used neural network processor (Neural Processing Unit, NPU).
[0088] However, in the case of strong noise, the signal-to-noise ratio of the sample image frame and the reference image frame is lower than that of the true image frame, and the difference between the enhanced sample image frame and the true image frame also reflects the noise difference between the enhanced sample image frame and the true image frame, that is, the error of the difference in optical flow information reflected by the difference is large. Therefore, adjusting the network parameters by treating the noise difference as the difference caused by the optical flow information is processing with optical flow information that has errors. If the network parameters of the deep learning model are adjusted using the gradient descent method according to the difference between the enhanced sample image frame and the true image frame, the results obtained will be different from the actual situation, resulting in low accuracy of the optical flow information learned by the deep learning model, low credibility of the optical flow information, affecting the accuracy of the deep learning model learning, and further resulting in low quality of the enhanced image generated using the optical flow information.
[0089] In order to improve the quality of images, embodiments of the present invention provide an optical flow information prediction network training method, an image enhancement method, a device and an apparatus.
[0090] The following first introduces an optical flow information prediction network training method provided by an embodiment of the present invention.
[0091] The optical flow information prediction network training method provided in the embodiment of the present invention can be applied to an electronic device. For example, the electronic device can be a remote server or a local computer.
[0092] An optical flow information prediction network training method provided by an embodiment of the present invention may include the following steps:
[0093] Acquire a first sample image frame and a second sample image frame with noise; wherein the first sample image frame and the second sample image frame are two image frames in the first sample video with an interval less than a first number;
[0094] Inputting the first sample image frame and the second sample image frame into the optical flow information prediction network of the initial structure to obtain the optical flow information of the second sample image frame relative to the first sample image frame;
[0095] Based on the obtained optical flow information, the true value original image frame corresponding to the second sample image frame is mapped to obtain a mapping result of the true value image frame corresponding to the second sample image frame; the true value image frame corresponding to the second sample image frame is consistent with the image content of the second sample image frame and has no noise;
[0096] Based on the difference between the obtained mapping result and the true value image frame corresponding to the first sample image frame, the network parameters of the optical flow information prediction network of the initial structure are adjusted until the preset convergence conditions are reached to obtain a trained optical flow information prediction network; the true value image frame corresponding to the first sample image frame is consistent with the image content of the first sample image frame and does not contain noise.
[0097] In the embodiment of the present invention, the first sample image frame and the second sample image frame are two adjacent frames with noise, and the optical flow information (predicted optical flow information) of the second sample image frame output by the optical flow information prediction network is used to map the true value image frame corresponding to the second sample image frame, that is, the second sample image frame without noise is mapped according to the predicted optical flow information to obtain a mapping result without noise, and the mapping of the second sample image frame with noise is avoided, so that the loss value between the mapping result and the true value image frame of the first sample image frame (the first sample image frame without noise) can reflect the gap between the predicted optical flow information and the true optical flow information (the optical flow information of the true value image frame corresponding to the second sample image frame relative to the true value image frame corresponding to the first sample image frame) as realistically as possible, and reduce the adverse effects of noise differences on adjusting network parameters. In other words, the trained optical flow information prediction network can identify the optical flow information between the two image frames as realistically as possible, and improve the accuracy of the learned optical flow information.
[0098] Combine the following Figure 1 An optical flow information prediction network training method provided by an embodiment of the present invention is introduced. Figure 1 As shown, the method may include steps S101-S104.
[0099] S101, obtaining a first sample image frame and a second sample image frame with noise.
[0100] The first sample image frame and the second sample image frame are two image frames in the first sample video whose interval is less than the first number.
[0101] It can be understood that if the signal-to-noise ratio of an image is less than a first preset threshold, the image is considered to be an image with noise, and if the signal-to-noise ratio of an image is greater than a second preset threshold, the image is considered to be an image without noise; wherein the first preset threshold is less than the second preset threshold, for example, the first preset threshold is 20 decibels, and the second preset threshold is 55 decibels. The first sample image frame and the second sample image frame are images with noise, and the signal-to-noise ratio of the first sample image frame and the second sample image frame may be less than the first preset threshold.
[0102] Optionally, in one implementation, the image acquisition device can acquire a video without noise under good acquisition conditions, select two image frames from the video without noise, and then perform noise addition processing on the image frames without noise to obtain a first sample image frame and a second sample image frame. This implementation will be described in detail in subsequent embodiments and will not be described in detail here.
[0103] For example, when collecting videos under good collection conditions, the staff can adjust the analog gain of the image sensor of the image collection device to the lowest analog gain in advance, and the image collection device collects sample videos under normal lighting conditions (for example, normal lighting conditions are 300-500 lux of light intensity) and at a moderate exposure time (for example, a moderate exposure time is 1 / 500th of a second to 1 / 120th of a second). When the analog gain of the image sensor is the lowest analog gain, the amplification factor of the electrical signal collected by the image sensor is the smallest, and the noise introduced is also relatively small; the noise of the image collected under normal lighting conditions is less than the noise of the image collected under dark light conditions (for example, dark light conditions are 1-100 lux of light intensity); when collecting videos, collecting at a moderate exposure time can avoid the situation where the image content is unclear due to overexposure or underexposure of the image brightness.
[0104] Optionally, in another implementation, the image acquisition device may acquire a noisy video under poor acquisition conditions, determine a first sample image frame and a second sample image frame in the video, and denoise the first sample image frame and the second sample image frame to obtain a true value image frame corresponding to the first sample image frame and a true value image frame corresponding to the second sample image frame.
[0105] The first sample video may be collected in multiple scenes, and the first sample videos in multiple scenes provide rich samples for the network to learn optical flow information, thereby enhancing the recognition ability of the trained optical flow information prediction network for multiple scenes. Exemplarily, the first sample video may be collected when the image acquisition device is fixed and the foreground being photographed is moving, such as the image acquisition device being fixed on the side of the road to photograph cars on the road; the first sample video may be collected when the image acquisition device is moving and the foreground being photographed is fixed, such as the image acquisition device being mounted on a moving drone to photograph static field scenery; the first sample video may be collected when both the image acquisition device and the foreground being photographed are moving, such as the image acquisition device being mounted on a moving drone to photograph track and field competitions.
[0106] The image frames in the first sample video may be in RAW format. The image frames in RAW format contain the original information that has been collected but not processed. For example, the original information may include original color information and detail information for image enhancement, etc. Image processing of images in RAW format can better process image frames collected under dark light conditions and improve image quality.
[0107] The electronic device can determine the first sample image frame and the second sample image frame from the first sample video. Exemplarily, the electronic device can use an image frame in the first sample video as the first sample image frame, and the first sample image frame is the image to be enhanced; the image frame in the first sample video whose interval with the first sample image frame is less than the first number can be used as the second sample image frame. Exemplarily, the first number can be a value between 2 and 5, for example, the first number can be 5, that is, the image frame of the fourth frame before the first sample image frame can be selected in the first sample video as the second sample image frame, or the image frame of the fourth frame after the first sample image frame can be selected as the second sample image frame. Since the interval between the selected second sample image frame and the first sample image frame is small, the second sample image frame has the same content as the first sample image frame, so that the content of the second sample image frame can provide a reference for the content of the first sample image frame, that is, the second sample image frame is used to provide a reference for the image to be enhanced, the first sample image frame can be called a noise image frame to be enhanced, and the second sample image frame can be called a noise reference image frame. Among them, the first sample image frame and the second sample image frame can be called noise data.
[0108] S102, inputting the first sample image frame and the second sample image frame into an optical flow information prediction network of an initial structure to obtain optical flow information of the second sample image frame relative to the first sample image frame.
[0109] It can be understood that the electronic device can input the first sample image frame and the second sample image frame into the optical flow information prediction network of the initial structure, and process the first sample image frame and the second sample image frame according to the network parameters of the optical flow information prediction network of the initial structure to obtain the predicted optical flow information of the second sample image frame relative to the first sample image frame. Wherein, the first sample image frame is a noise image frame to be enhanced, and the second sample image frame is a noise reference image frame. The obtained optical flow information is the optical flow information of the second sample image frame relative to the first sample image frame, so as to subsequently map the noise-free reference image frame to the vector space of the noise-free image frame to be enhanced. The obtained optical flow information can represent the offset direction and offset of the pixel points in the second sample image frame relative to the pixel points in the first sample image frame.
[0110] S103, mapping the true value image frame corresponding to the second sample image frame based on the obtained optical flow information to obtain a mapping result of the true value image frame corresponding to the second sample image frame.
[0111] The true value image frame corresponding to the second sample image frame has the same image content as the second sample image frame and is free of noise.
[0112] It is understandable that in order to avoid the influence of the noisy second sample image frame on the difference in optical flow information, the electronic device can map the true value image frame corresponding to the second sample image frame without noise. Exemplarily, the optical flow information can be in the form of a matrix, and each element in the matrix of the optical flow information represents the offset direction and offset of a pixel point. In the mapping process, the electronic device can transform the vector space where each pixel point in the true value image frame corresponding to the second sample image frame is located into the vector space where each pixel point in the first sample image frame is located based on the optical flow information in the matrix form. Among them, the mapping result is the image frame obtained by mapping the noise-free reference image frame according to the predicted optical flow information, that is, the predicted alignment result.
[0113] S104, based on the difference between the obtained mapping result and the true value image frame corresponding to the first sample image frame, adjusting the network parameters of the optical flow information prediction network of the initial structure until a preset convergence condition is reached to obtain a trained optical flow information prediction network.
[0114] The true value image frame corresponding to the first sample image frame has the same image content as the first sample image frame and has no noise.
[0115] It can be understood that the true value image frame corresponding to the first sample image frame is used as the target of predicting the alignment result, and the network parameters are continuously adjusted so that the predicted alignment result tends to the true result without noise.
[0116] The electronic device may compare the difference between the mapping result and the true value image frame corresponding to the first sample image frame to determine the difference between the predicted optical flow information and the real optical flow information. Exemplarily, the loss value between the mapping result and the original image frame corresponding to the first sample image frame may be calculated according to a preset loss function as the difference between the mapping result and the true value image frame corresponding to the first sample image frame.
[0117] In the prior art, in order to enable the optical flow information prediction network to predict the optical flow information of an image frame with noise, the second sample image frame is mapped, and the network parameters are adjusted based on the difference between the mapping result of the second sample image frame and the true value image frame corresponding to the first sample image frame. However, the difference between the mapping result of the second sample image frame and the image frame corresponding to the first sample image frame reflects not only the difference between the predicted optical flow information and the real optical flow information, but also the noise difference between the second sample image frame and the true value image frame corresponding to the first sample image frame. If the network parameters are adjusted based on the difference between the mapping result of the second sample image frame and the true value image frame corresponding to the first sample image frame, the noise difference will also be regarded as the difference caused by the optical flow information to adjust the network parameters. In the embodiment of the present invention, the difference between the mapping result of the true value image frame corresponding to the second sample image frame and the true value image frame corresponding to the first sample image frame can truly reflect the gap between the predicted optical flow information and the real optical flow information (optical flow information of the second original image frame relative to the first original image frame), reducing the adverse effects of the noise difference on the adjustment of the network parameters.
[0118] Optionally, the preset loss function used may be an absolute value loss (Least Absolute Deviations Loss, LAD-Loss) function, also known as an L1-Loss loss function. Specifically, the formula of the absolute value loss function is: ;in, represents the loss value between the mapping result of the true value image frame corresponding to the second sample image frame and the true value image frame corresponding to the first sample image frame, represents the pixel value of the pixel point in the mapping result of the true value image frame corresponding to the second sample image frame, represents the pixel value of the pixel point of the true value image frame corresponding to the first sample image frame, and the size of the true value image frame corresponding to the second sample image frame is the same as that of the true value image frame corresponding to the first sample image frame. It represents the absolute value of the difference between the pixel value of each pixel point in the mapping result of the true value image frame corresponding to the second sample image frame and the pixel value of the pixel point at the same position in the true value image frame corresponding to the first sample image frame.
[0119] Optionally, the preset loss function used may be a mean squared error loss (MSE-Loss) function, also known as an L2-Loss function. Specifically, the formula of the mean squared error loss function is: ;in, It represents the square value of the difference between the pixel value of each pixel point in the mapping result of the true value image frame corresponding to the second sample image frame and the pixel value of the pixel point at the same position in the true value image frame corresponding to the first sample image frame.
[0120] It can be understood that the network parameters of the optical flow information prediction network of the initial structure can be adjusted based on the obtained loss value to determine whether the preset convergence conditions are met. If the preset convergence conditions are met, the training is terminated to obtain the trained image processing network. If the preset convergence conditions are not met, the process returns to step S101 to continue training the optical flow information prediction network.
[0121] Optionally, the preset convergence condition can be pre-set as the loss value is less than the preset loss threshold. If the loss value is less than the preset loss threshold, it can be determined that the preset convergence condition is met; if the loss value is greater than the preset loss threshold, it can be determined that the preset convergence condition is met. Optionally, in another implementation, the preset convergence condition can be pre-set as the number of training times reaches the preset number of iterations. If the number of training times reaches the preset number of iterations, the training is terminated. In the scheme of the present invention, the optical flow information prediction network can be trained according to the preset convergence condition to ensure the effectiveness of the trained optical flow information prediction network.
[0122] In the embodiment of the present invention, the first sample image frame and the second sample image frame are two adjacent frames with noise, and the optical flow information (predicted optical flow information) of the second sample image frame output by the optical flow information prediction network is used to map the true value image frame corresponding to the second sample image frame, that is, the second sample image frame without noise is mapped according to the predicted optical flow information to obtain a mapping result without noise, and the mapping of the second sample image frame with noise is avoided, so that the loss value between the mapping result and the true value image frame of the first sample image frame (the first sample image frame without noise) can reflect the gap between the predicted optical flow information and the true optical flow information (the optical flow information of the true value image frame corresponding to the second sample image frame relative to the true value image frame corresponding to the first sample image frame) as realistically as possible, and reduce the adverse effects of noise differences on adjusting network parameters. In other words, the trained optical flow information prediction network can identify the optical flow information between the two image frames as realistically as possible, and improve the accuracy of the learned optical flow information.
[0123] In addition, in an embodiment of the present invention, a mapping result of a true image frame corresponding to a second sample image frame without noise is used to calculate the photometric loss with the true image frame corresponding to the first sample image frame, so as to avoid the influence of noise on the gradient descent of the network, make network learning more reasonable and accurate, avoid manual labeling of optical flow information, and do not need to provide accurate optical flow supervision information of noisy images, so as to realize self-supervised learning, thereby saving labeling costs, simplifying the process, and making the method highly practical.
[0124] Optionally, in one embodiment, Figure 1 Based on the optical flow information prediction network training method shown in Figure 2 As shown, step S101 includes steps S1011 - S1012 .
[0125] S1011, acquiring a first original image frame and a second original image frame without noise.
[0126] The first original image frame and the second original image frame are two image frames in the second sample video whose interval is less than the first number.
[0127] It can be understood that the first original image frame and the second original image frame are images with a signal-to-noise ratio greater than the second preset threshold, for example, the first original image frame and the second original image frame are images with a signal-to-noise ratio of 60 decibels. The second sample video can be a video pre-collected under good acquisition conditions. The first sample video can be obtained by adding noise to each image frame in the second sample video. If a first original image frame and a first sample image frame have the same content, the position order of the first original image frame in the second sample video is the same as the position order of the first sample image frame in the first sample video.
[0128] S1012: Perform noise adding processing on the first original image frame and the second original image frame to obtain a first sample image frame and a second sample image frame with noise.
[0129] It can be understood that the first sample image frame is obtained based on the first original image frame, and the second sample image frame is obtained based on the second original image frame. Since the first original image frame is the image frame to be enhanced and the second original image frame is the reference image frame, the first sample image frame can be used as the image frame to be enhanced, the second sample image frame can be used as the reference image frame, and the first sample image frame can be enhanced using the second sample image frame.
[0130] It should be noted that since the image frames taken under dark light conditions have the disadvantage of being noisy, in order to enable the optical flow information prediction network to output optical flow information for image frames with strong noise taken under dark light conditions, noise addition processing can be performed on the first original image frame and the second original image frame.
[0131] Optionally, in this embodiment, the first original image frame is a true value image frame corresponding to the first sample image frame, and the second original image frame is a true value image frame corresponding to the second sample image frame.
[0132] It can be understood that the first original image frame is an image without noise, and the first original image frame and its corresponding first sample image frame have the same image content, and the second original image frame is an image without noise, and the second original image frame and its corresponding second sample image frame have the same image content. Therefore, the first original image frame can be used as the true value image frame corresponding to the first sample image frame, and the second original image frame can be used as the true value image frame corresponding to the second sample image frame, so as to obtain an image without noise collected under good collection conditions, further reduce noise differences, and improve the accuracy of optical flow information.
[0133] Optionally, in one embodiment, Figure 2 Based on the optical flow information prediction network training method shown in Figure 3 As shown, step S1012 includes steps S1012a-S1012b.
[0134] S1012a, dividing the pixel values of the first original image frame and the second original image frame by a specified value respectively to obtain a third original image frame and a fourth original image frame.
[0135] It is understandable that since the reduction in brightness will cause the loss of details in the darker areas of the image frame, the clarity of the image frame can be reduced, so the first original image frame and the second original image frame can be processed to reduce the image brightness to obtain low-brightness image frames, that is, low-illuminance sample pairs.
[0136] Since the mean value of all pixels in an image frame can represent the brightness of the image frame, dividing the pixel value of the image frame by a specified value can reduce the brightness of the image frame in geometric proportion with the specified value as a ratio, so that the pixel values of the first original image frame and the second original image frame can be divided by the same specified value to obtain the third original image frame and the fourth original image frame. In an embodiment of the present invention, the pixel values of the first original image frame and the second original image frame can be divided by the same specified value to ensure that the first original image frame and the second original image frame reduce the same brightness, and then the second sample image frame and the first sample image frame with the same brightness are processed using the optical flow information prediction network to reduce the brightness variable, avoid the noise difference between the second sample image frame and the first sample image frame due to the different brightness, and avoid the additional impact caused by the different brightness of the two sample image frames when adjusting the network parameters. Among them, the specified value can be pre-set, such as 10, 50, or 100.
[0137] S1012b, adding noise to the third original image frame and the fourth original image frame respectively to obtain a first sample image frame and a second sample image frame.
[0138] It can be understood that the first original image frame and the second original image frame are image frames captured by the same image sensor, and the noise of the image frames captured based on the same image sensor under dark light conditions should be the same. The same noise can be added to the third original image frame and the fourth original image frame respectively to avoid noise differences between the second sample image frame and the first sample image frame caused by different noises.
[0139] Exemplarily, in the process of adding noise, Poisson noise and / or readout noise may be added to the third original image frame and the fourth original image frame. The specific implementation method will be described in detail in subsequent embodiments.
[0140] Generally speaking, both reducing brightness and adding noise are operations that reduce clarity. The added noise is used to simulate the noise in an image captured under low light conditions. If noise is added first and then the brightness is reduced, the brightness reduction process will disturb the noise added first, causing the noise characteristics to be destroyed and making it impossible to simulate the noise of the image captured under low light conditions. In an embodiment of the present invention, reducing brightness first and then adding noise can avoid the added noise from being disturbed, ensure the effect of adding noise, simulate the noise of the image captured under low light conditions as much as possible, and improve the accuracy of the trained optical flow information prediction network in predicting brightness information for images captured under low light conditions.
[0141] Optionally, in one implementation, the method further includes step A1, and step S1012b includes step A2.
[0142] A1, obtaining noise parameters of a sensor for collecting a second sample video.
[0143] A2, based on the noise parameter, adding noise to the third original image frame and the fourth original image frame respectively to obtain a first sample image frame and a second sample image frame.
[0144] It is understood that the noise parameters can be pre-calibrated. For example, the manufacturer of the sensor can calibrate the noise parameters of the sensor under laboratory conditions. In other words, the noise parameters of a sensor are fixed and will not change with changes in the illumination conditions during acquisition.
[0145] The electronic device can obtain the noise parameters of the sensor that captures the second sample video, and then add noise based on the obtained noise parameters. The added noise simulates the noise generated by the sensor itself under low-light conditions, so that the first sample image frame and the second sample image frame can simulate the image frames captured by the sensor under low-light conditions, thereby improving the accuracy of the trained optical flow information prediction network in predicting brightness information for images captured under low-light conditions.
[0146] Optionally, in one implementation, step S10221 includes steps B1-B3.
[0147] B1, when the noise parameter includes the mean value of Poisson noise, Poisson noise is added to the third original image frame and the fourth original image frame according to the noise parameter to obtain a first sample image frame and a second sample image frame.
[0148] It can be understood that when the noise parameters include the mean of Poisson noise, the pixel value of each pixel in the third original image frame can be randomly changed according to the Poisson distribution, and the pixel value of each pixel in the fourth original image frame can be randomly changed according to the Poisson distribution, and the mean of the pixel values of each pixel after the change is the mean of Poisson noise, simulating the shot noise generated under dark light conditions.
[0149] B2. When the noise parameters include the mean and standard deviation of Gaussian noise, Gaussian noise is added to the third original image frame and the fourth original image frame according to the noise parameters to obtain a first sample image frame and a second sample image frame.
[0150] It can be understood that when the noise parameters include the mean and standard deviation of Gaussian noise, random pixel values can be generated according to the mean and standard deviation of Gaussian noise, and for each pixel point in the third original image frame, the generated random pixel value can be superimposed with the pixel value of the pixel point; for each pixel point in the fourth original image frame, the generated random pixel value can be superimposed with the pixel value of the pixel value to simulate the Gaussian noise generated under dark light conditions.
[0151] B3, when the noise parameters include the mean of Poisson noise and the mean and standard deviation of Gaussian noise, Poisson noise and Gaussian noise are added to the third original image frame and the fourth original image frame respectively according to the noise parameters to obtain a first sample image frame and a second sample image frame.
[0152] It can be understood that, when the noise parameters include the mean of Poisson noise and the mean and standard deviation of Gaussian noise, Poisson noise and Gaussian noise can be added at the same time to simulate the shot noise and Gaussian noise generated under dark light conditions, ensuring that the added noise is the noise generated by the sensor itself under dark light conditions, so that the first sample image frame and the second sample image frame can simulate the image frames captured by the sensor under dark light conditions, thereby improving the accuracy of the trained optical flow information prediction network in predicting brightness information for images captured under low light conditions.
[0153] Optionally, in one embodiment, the optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork and an output subnetwork. Figure 1 Based on the optical flow information prediction network training method shown in Figure 4 As shown, step S102 includes steps S1021-S1024.
[0154] S1021: Splice the first sample image frame and the second sample image frame to obtain a first splicing result.
[0155] S1022: Input the first concatenation result to the input sub-network to obtain a first feature vector.
[0156] S1023, input the first feature vector into the feature extraction sub-network to obtain a second feature vector.
[0157] S1024, input the second feature vector to the output sub-network to obtain optical flow information of the second sample image frame relative to the first sample image frame.
[0158] It can be understood that the optical flow information prediction network is a straight waterfall structure in which each sub-network is unidirectionally connected. In the process of the optical flow information prediction network processing the image frame, the resolution of the feature map of the image frame will gradually decrease.
[0159] The electronic device can splice the first sample image frame with the second sample image frame to obtain a first splicing result, and then input the first splicing result into an input subnetwork, the input subnetwork includes multiple groups of convolutional layers and nonlinear activation layers, and one group of convolutional layers and nonlinear activation layers in the input subnetwork includes a convolutional layer with a stride of 1 and a nonlinear activation layer. The image frame can be converted into a feature map through the convolutional layer with a stride of 1, retaining more image features so that the subsequent feature extraction subnetwork can perform more accurate extraction; the nonlinear activation layer can enable the network to learn the complex features of nonlinear information. Optionally, the input subnetwork can include two groups of convolutional layers and nonlinear activation layers, and the two groups of convolutional layers and nonlinear activation layers can not only extract rich feature information, but also avoid overly complex network structures, which can improve the convenience and applicability of the network. The first splicing result can be pre-extracted through the input subnetwork to obtain a first feature vector, which can also be called a feature map.
[0160] Then the first feature vector is input into the feature extraction subnetwork, which includes multiple groups of convolutional layers and nonlinear activation layers. The first group of convolutional layers and nonlinear activation layers in the feature extraction subnetwork includes a convolutional layer with a step size of 2 and a nonlinear activation layer. The convolutional layer with a step size of 2 plays a role of halving the resolution; the other groups of convolutional layers and nonlinear activation layers except the first group include: a convolutional layer with a step size of 1 and a nonlinear activation layer. Exemplarily, the feature extraction subnetwork includes three groups of convolutional layers and nonlinear activation layers. The first feature vector can be subjected to three dimensionality reduction processing through the feature extraction subnetwork, that is, the first feature vector can be subjected to three resolution reduction processing, and the size of the obtained second feature vector is one eighth of the size of the first splicing result. In the embodiment of the present invention, three groups of convolutional layers and nonlinear activation layers are used to extract features from feature maps, which can not only extract feature information that can express feature maps, but also avoid overly complex network structures, and improve the convenience and applicability of the network.
[0161] The second eigenvector is then input into the output subnetwork and processed by the output subnetwork to obtain two-dimensional data. The two-dimensional data are x and y, that is, the number of channels output by the network is 2, corresponding to the offsets in two different directions in the image frame, which serve as the optical flow information of the second sample image frame relative to the first sample image frame.
[0162] Optionally, in one implementation, the output subnetwork includes: an activation layer, an amplification layer, and an upsampling layer; step S1024 includes steps C1-C3.
[0163] C1, input the second feature vector into the activation layer to obtain the first optical flow information to be used.
[0164] C2, inputting the first optical flow information to be used into the amplification layer to obtain the amplified first optical flow information to be used.
[0165] C3, inputting the amplified first optical flow information to be used into an upsampling layer to obtain optical flow information of the second sample image frame relative to the first sample image frame.
[0166] It can be understood that the activation layer includes a convolution layer whose output range is in the entire real number domain and a hyperbolic tangent (tanh) activation layer. The second feature vector can be reduced to two-dimensional data through the convolution layer of the output subnetwork; the two-dimensional data can be adjusted to the interval [-1, 1] through the hyperbolic tangent activation layer. In other words, the first optical flow information to be used is two-dimensional data between [-1, 1].
[0167] In order to subsequently map the second original image frame, the first optical flow information to be used is multiplied by a preset magnification factor through the magnification layer. The magnification factor can also be called the radius (Radius, R), and the magnification factor is a preset value of the motion range that can actually be processed. For example, if the object in the captured video moves faster, a smaller magnification factor can be set to map a small range of pixels; if the object in the captured video moves slower, a larger magnification factor can be set to map a large range of pixels. For example, if the magnification factor is 50, it means that the maximum motion offset that the optical flow information can supplement is 50 pixels.
[0168] Then, the amplified first optical flow information to be used is input to the upsampling layer, and the resolution of the predicted optical flow information is aligned with the size of the original input image. For example, if the amplified first optical flow information to be used is reduced by 8 times relative to the first splicing result, the amplified first optical flow information to be used can be multiplied by 8 (×8). Exemplarily, the nearest neighbor mode can be used to take the optical flow information of each pixel in the amplified first optical flow information to be used as the optical flow information of the pixel closest to the pixel in the final optical flow information.
[0169] In the embodiment of the present invention, the first optical flow information to be used can be restored to the original input image size through the magnification layer and the upsampling layer, thereby improving the convenience of subsequent mapping using the optical flow information.
[0170] Combine the following Figure 5 An image enhancement method provided by an embodiment of the present invention is introduced. Figure 5 As shown, the method may include steps S501-S504.
[0171] S501: Acquire an image frame to be enhanced and a reference image frame.
[0172] The image frame to be enhanced and the reference image frame are two image frames in the same video with an interval less than a second number.
[0173] It is understandable that the image enhancement method can be applied to an electronic device that can obtain a video captured by a camera and can enhance the image frames in the captured video. Exemplarily, the electronic device can be a camera or a server. Optionally, when the server can have the functions of training the network and enhancing the image at the same time, the execution subject of the optical flow information prediction method and the image enhancement method can be the same execution subject; optionally, when the server can only have the functions of training the network or enhancing the image, the execution subject of the optical flow information prediction method and the image enhancement method can be different execution subjects.
[0174] Optionally, in one implementation, the image sensor of the image acquisition device can send the acquired image frames to the image processing chip in the image acquisition device frame by frame while acquiring the image frames of the video frame by frame, and the image processing chip can process the image frames received frame by frame to process the video in real time. Optionally, the image processing chip can enhance the current image frame using the historical image frames received before the current image frame; optionally, the image processing chip can also receive the advanced image frames of a specified number of frames after the current image frame after receiving the current image frame, and use the advanced image frames to enhance the current image frame. In this implementation, the video image can be processed while it is being acquired, with fast processing speed and high real-time performance.
[0175] Optionally, in another implementation, the image sensor may first capture the video, and then send the captured complete video to the image processing chip, and the image processing chip may use the image frames before or after the image frames to be enhanced in the video to process the image frames to be enhanced. In this implementation, the image frames after the image frames to be enhanced may be captured in advance, and the image frames before or after the image frames to be enhanced may be used to process the image frames to be enhanced, thereby improving the image quality.
[0176] It is understandable that the second number may be the same as or different from the first number in the aforementioned embodiment. Since the image frame to be enhanced and the reference image frame are two image frames in the same video with an interval less than the second number, the reference image frame and the image frame to be enhanced have the same content, and the image frame to be enhanced can be enhanced using the reference image frame.
[0177] The acquired image frames to be enhanced and the reference image frames may be images with lower definition and more noise acquired under dark light conditions.
[0178] S502: Input the image frame to be enhanced and the reference image frame into a pre-trained optical flow information prediction network to obtain optical flow information of the reference image frame relative to the image frame to be enhanced.
[0179] The optical flow information prediction network is trained based on the optical flow information prediction network training method provided in the aforementioned embodiment.
[0180] It is understandable that the reference image frame may have moved relative to the content of the image frame to be enhanced, and the optical flow information of the reference image frame relative to the image frame to be enhanced can be obtained through the pre-trained optical flow information prediction network. The optical flow information prediction network trained by the method provided in the above embodiment can accurately identify the optical flow information for the image frame to be enhanced and the reference image frame with more noise collected under dark light conditions.
[0181] S503, mapping the reference image frame based on the obtained optical flow information to obtain a mapping result of the reference image frame.
[0182] It is understandable that there is no motion offset between the mapping result of the reference image frame and the image frame to be enhanced. The specific implementation of the mapping process of the image frame based on the optical flow information has been described in the above embodiment, and will not be repeated here.
[0183] S504, fusing the mapping result of the reference image frame with the image frame to be enhanced to obtain an enhanced image frame of the image frame to be enhanced.
[0184] Since there is no motion offset between the mapping result of the reference image frame and the image frame to be enhanced, the mapping result of the reference image frame is fused with the image frame to be enhanced to avoid motion blur, and the mapping result of the reference image frame is used to supplement the content of the image frame to be enhanced to obtain an enhanced image frame of the image frame to be enhanced. Exemplarily, during the fusion process, the mapping result of the reference image frame and the pixel value of each pixel point of the image frame to be enhanced can be added according to a preset ratio.
[0185] It is understandable that after the optical flow information prediction network training is completed, the above steps can be used for testing. Specifically, the second sample image frame with noise is mapped using the optical flow information of the second sample image frame relative to the first sample image frame to obtain the mapping result of the second sample image frame; the mapping result of the second sample image frame is fused with the first sample image frame to obtain the fusion result as the alignment result of the noise image.
[0186] In the embodiment of the present invention, the trained optical flow information prediction network can identify the optical flow information between two image frames as realistically as possible, improve the accuracy of the output optical flow information, and further improve the quality of the enhanced image generated using the output optical flow information.
[0187] Optionally, in one embodiment, the optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork and an output subnetwork. Figure 5 Based on the optical flow information prediction network training method shown in Figure 6 As shown, step S502 includes steps S5021-S5024.
[0188] S5021: Splice the image frame to be enhanced and the reference image frame to obtain a second splicing result.
[0189] S5022, input the second concatenation result into the input sub-network to obtain a third feature vector.
[0190] S5023, input the third feature vector into the feature extraction sub-network to obtain a fourth feature vector.
[0191] S5024, input the fourth eigenvector to the output sub-network to obtain optical flow information of the reference image frame relative to the image frame to be enhanced.
[0192] Optionally, in one implementation, the output subnetwork includes: an activation layer, an amplification layer, and an upsampling layer; step S5024 includes steps D1-D3.
[0193] D1, input the fourth eigenvector into the activation layer to obtain the second optical flow information to be used.
[0194] D2, inputting the second optical flow information to be used into the amplification layer to obtain the amplified second optical flow information to be used.
[0195] D3, inputting the amplified second optical flow information to be used into the upsampling layer to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced.
[0196] It is understandable that during the two processes of training and applying the optical flow information prediction network, the network structure of the optical flow information prediction network will not change. The specific implementation method of using the network structure in the optical flow information prediction network to process the input content can refer to the aforementioned embodiment, and will not be elaborated here.
[0197] Figure 7 Schematic diagram of the principle of the optical flow information prediction network training method provided by the embodiment of the present invention. Figure 7 As shown, the electronic device can obtain the first original image frame and the second original image frame, and then perform noise addition processing on the first original image frame and the second original image frame respectively to obtain the first sample image frame and the second sample image frame. The first sample image frame and the second sample image frame are input into the optical flow information prediction network of the initial structure to obtain the optical flow information of the second sample image frame relative to the first sample image frame; then the second original image frame is mapped using the optical flow information to obtain the mapping result of the second original image frame. Then, based on the difference (loss value) between the mapping result of the second original image frame and the first original image frame, the network parameters of the optical flow information prediction network of the initial structure are adjusted until the preset convergence condition is reached to obtain the trained optical flow information prediction network.
[0198] Figure 8 This is a schematic diagram of the structure of the optical flow information prediction network, such as Figure 8As shown, the optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork, and an output subnetwork. The input subnetwork includes multiple groups of convolutional layers and nonlinear activation layers, and one group of convolutional layers and nonlinear activation layers in the input subnetwork includes a convolutional layer with a step size of 1 and a nonlinear activation layer (a). The first sample image frame is spliced with the second sample image frame to obtain a first splicing result; the first splicing result is used as input data, and the first splicing result can be pre-extracted through the input subnetwork to obtain a first feature vector.
[0199] The feature extraction subnetwork includes three groups of convolutional layers and nonlinear activation layers. The first group of convolutional layers and nonlinear activation layers in the feature extraction subnetwork includes a convolutional layer with a stride of 2 and a nonlinear activation layer (b). Except for the first group, other groups of convolutional layers and nonlinear activation layers include a convolutional layer with a stride of 1 and a nonlinear activation layer.
[0200] The output subnetwork includes: an activation layer, an amplification layer, and an upsampling layer. The activation layer includes a convolution layer whose output range is in the entire real number domain and a hyperbolic tangent activation layer (d). The convolution layer is a convolution layer (c) with a step size of 1. The second feature vector can be reduced to two-dimensional data through the convolution layer of the output subnetwork; the two-dimensional data can be adjusted to the interval [-1, 1] through the hyperbolic tangent activation layer. In other words, the first optical flow information to be used is two-dimensional data between [-1, 1]. Through the amplification layer, the first optical flow information to be used is multiplied by a preset amplification factor. The amplified first optical flow information to be used is input into the upsampling layer, and the resolution of the predicted optical flow information is aligned with the original input image size to obtain the optical flow information of the second sample image frame relative to the first sample image frame as output data.
[0201] FIG9 (a), FIG9 (b), and FIG9 (c) are schematic diagrams of the effect of the image enhancement method provided by the embodiment of the present invention. FIG9 (a) is a target image frame in a video shot under dark light conditions. The dots represent noise, and the dotted line represents a human figure. The target image frame has a lot of noise, the human figure is blurred, and the image quality is poor. The image frame of the image frame adjacent to the image frame of 5 frames is determined in the video, and the determined image frame is directly superimposed with the target image frame. Since the target image frame moves relative to the determined image frame, there is a lot of noise in the result after direct superposition (as shown in FIG9 (b)), and the human figure also has a ghost image caused by motion blur. The quality of the result after direct superposition is poor, and the quality of the image frame cannot be improved. The target image frame is enhanced according to the image enhancement algorithm provided by the embodiment of the present invention to obtain an enhanced image frame. As shown in FIG9 (c), the noise in the enhanced image frame is significantly reduced, and the solid line human figure is clear. That is to say, the quality of the enhanced image frame is improved. The image enhancement algorithm provided by the embodiment of the present invention can improve the image quality.
[0202] The embodiment of the present invention also provides an optical flow information prediction network training device, such as Fig.10 As shown, the device comprises:
[0203] A first acquisition module 1010 is used to acquire a first sample image frame and a second sample image frame with noise; wherein the first sample image frame and the second sample image frame are two image frames in the first sample video with an interval less than a first number;
[0204] A first input module 1020, configured to input the first sample image frame and the second sample image frame into an optical flow information prediction network of an initial structure, to obtain optical flow information of the second sample image frame relative to the first sample image frame;
[0205] A first mapping module 1030 is configured to map the true value image frame corresponding to the second sample image frame based on the obtained optical flow information to obtain a mapping result of the true value image frame corresponding to the second sample image frame; wherein the true value image frame corresponding to the second sample image frame has the same image content as the second sample image frame and has no noise;
[0206] The adjustment module 1040 is used to adjust the network parameters of the optical flow information prediction network of the initial structure based on the difference between the obtained mapping result and the true value image frame corresponding to the first sample image frame, until the preset convergence condition is reached to obtain a trained optical flow information prediction network; wherein the true value image frame corresponding to the first sample image frame is consistent with the image content of the first sample image frame and does not contain noise.
[0207] Optionally, the first acquisition module 1010 includes:
[0208] A first acquisition unit is used to acquire a first original image frame and a second original image frame without noise; wherein the first original image frame and the second original image frame are two image frames in the second sample video with an interval less than a first number;
[0209] An adding unit, configured to perform noise adding processing on the first original image frame and the second original image frame to obtain a first sample image frame and a second sample image frame with noise;
[0210] The first original image frame is a true value image frame corresponding to the first sample image frame, and the second original image frame is a true value image frame corresponding to the second sample image frame.
[0211] Optionally, add units including:
[0212] a brightness reduction subunit, configured to divide the pixel values of the first original image frame and the second original image frame by a specified value, respectively, to obtain a third original image frame and a fourth original image frame;
[0213] The noise adding subunit is used to add noise to the third original image frame and the fourth original image frame respectively to obtain a first sample image frame and a second sample image frame.
[0214] Optionally, the device further comprises:
[0215] A third acquisition module, used to acquire noise parameters of a sensor that collects the second sample video;
[0216] The noise adding subunit is specifically used to add noise to the third original image frame and the fourth original image frame respectively based on the noise parameter to obtain a first sample image frame and a second sample image frame.
[0217] Optionally, a noise adding subunit is specifically configured to, when the noise parameter includes a mean value of Poisson noise, add Poisson noise to the third original image frame and the fourth original image frame according to the noise parameter, respectively, to obtain a first sample image frame and a second sample image frame;
[0218] In a case where the noise parameters include a mean value and a standard deviation of Gaussian noise, Gaussian noise is added to the third original image frame and the fourth original image frame respectively according to the noise parameters to obtain a first sample image frame and a second sample image frame;
[0219] When the noise parameters include the mean of Poisson noise and the mean and standard deviation of Gaussian noise, Poisson noise and Gaussian noise are added to the third original image frame and the fourth original image frame respectively according to the noise parameters to obtain a first sample image frame and a second sample image frame.
[0220] Optionally, the optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork and an output subnetwork;
[0221] The first input module 1020 includes:
[0222] A first splicing unit, configured to splice the first sample image frame with the second sample image frame to obtain a first splicing result;
[0223] A first input unit, used for inputting the first concatenation result into the input sub-network to obtain a first feature vector;
[0224] A second input unit, used for inputting the first feature vector into the feature extraction sub-network to obtain a second feature vector;
[0225] The third input unit is used to input the second feature vector into the output sub-network to obtain the optical flow information of the second sample image frame relative to the first sample image frame.
[0226] Optionally, the output subnetwork includes: an activation layer, an amplification layer, and an upsampling layer;
[0227] A third input unit is specifically used to input the second feature vector into the activation layer to obtain the first optical flow information to be used;
[0228] Inputting the first optical flow information to be used into the amplification layer to obtain amplified first optical flow information to be used;
[0229] The amplified first optical flow information to be used is input into the upsampling layer to obtain the optical flow information of the second sample image frame relative to the billionth sample image frame.
[0230] In the embodiment of the present invention, the first sample image frame and the second sample image frame are two adjacent frames with noise, and the optical flow information (predicted optical flow information) of the second sample image frame output by the optical flow information prediction network is used to map the true value image frame corresponding to the second sample image frame, that is, the second sample image frame without noise is mapped according to the predicted optical flow information to obtain a mapping result without noise, and the mapping of the second sample image frame with noise is avoided, so that the loss value between the mapping result and the true value image frame of the first sample image frame (the first sample image frame without noise) can reflect the gap between the predicted optical flow information and the true optical flow information (the optical flow information of the true value image frame corresponding to the second sample image frame relative to the true value image frame corresponding to the first sample image frame) as realistically as possible, and reduce the adverse effects of noise differences on adjusting network parameters. In other words, the trained optical flow information prediction network can identify the optical flow information between the two image frames as realistically as possible, and improve the accuracy of the learned optical flow information.
[0231] The embodiment of the present invention also provides an image enhancement device, such as Fig.11 As shown, the device comprises:
[0232] The second acquisition module 1110 is used to acquire an image frame to be enhanced and a reference image frame; wherein the image frame to be enhanced and the reference image frame are two image frames in the same video with an interval less than a second number;
[0233] A second input module 1120 is used to input the image frame to be enhanced and the reference image frame into a pre-trained optical flow information prediction network to obtain optical flow information of the reference image frame relative to the image frame to be enhanced; wherein the optical flow information prediction network is trained based on an optical flow information prediction network training device;
[0234] A second mapping module 1130, configured to map the reference image frame based on the obtained optical flow information to obtain a mapping result of the reference image frame;
[0235] The fusion module 1140 is used to fuse the mapping result of the reference image frame with the image frame to be enhanced to obtain an enhanced image frame of the image frame to be enhanced.
[0236] Optionally, the optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork and an output subnetwork;
[0237] The second input module 1120 includes:
[0238] A second splicing unit, used for splicing the image frame to be enhanced and the reference image frame to obtain a second splicing result;
[0239] a fourth input unit, configured to input the second concatenation result into the input subnetwork to obtain a third feature vector;
[0240] a fifth input unit, configured to input the third feature vector into the feature extraction subnetwork to obtain a fourth feature vector;
[0241] The sixth input unit is used to input the fourth feature vector into the output sub-network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced.
[0242] Optionally, the output subnetwork includes: an activation layer, an amplification layer, and an upsampling layer;
[0243] a sixth input unit, specifically configured to input the fourth eigenvector into the activation layer to obtain second optical flow information to be used;
[0244] Inputting the second optical flow information to be used into the amplification layer to obtain amplified second optical flow information to be used;
[0245] The amplified second optical flow information to be used is input into the upsampling layer to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced.
[0246] In the embodiment of the present invention, the trained optical flow information prediction network can identify the optical flow information between two image frames as realistically as possible, improve the accuracy of the output optical flow information, and further improve the quality of the enhanced image generated using the output optical flow information.
[0247] The embodiment of the present invention further provides an electronic device, such as Fig.12As shown, it includes a processor 1201 , a communication interface 1202 , a memory 1203 and a communication bus 1204 , wherein the processor 1201 , the communication interface 1202 , and the memory 1203 communicate with each other via the communication bus 1204 .
[0248] Memory 1203, used for storing computer programs;
[0249] The processor 1201 is used to implement the above-mentioned optical flow information prediction network training method and image enhancement method when executing the program stored in the memory 1203.
[0250] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0251] The communication interface is used for communication between the above electronic device and other devices.
[0252] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0253] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0254] In another embodiment provided by the present invention, a computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned optical flow information prediction network training methods and image enhancement methods are implemented.
[0255] In another embodiment provided by the present invention, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any one of the optical flow information prediction network training methods and image enhancement methods in the above embodiments.
[0256] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.
[0257] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0258] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0259] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for training an optical flow information prediction network, characterized in that: The method comprises: Acquire a first sample image frame and a second sample image frame with noise; wherein the first sample image frame and the second sample image frame are two image frames in the first sample video with an interval less than a first number; Inputting the first sample image frame and the second sample image frame into an optical flow information prediction network of an initial structure to obtain optical flow information of the second sample image frame relative to the first sample image frame; Mapping the true value image frame corresponding to the second sample image frame based on the obtained optical flow information to obtain a mapping result of the true value image frame corresponding to the second sample image frame; wherein the true value image frame corresponding to the second sample image frame has the same image content as the second sample image frame and has no noise; Based on the difference between the obtained mapping result and the true value image frame corresponding to the first sample image frame, the network parameters of the optical flow information prediction network of the initial structure are adjusted until the preset convergence conditions are reached, thereby obtaining a trained optical flow information prediction network; wherein the true value image frame corresponding to the first sample image frame is consistent with the image content of the first sample image frame and does not contain noise.
2. The method according to claim 1, characterized in that The step of acquiring a first sample image frame and a second sample image frame with noise includes: Acquire a first original image frame and a second original image frame without noise; wherein the first original image frame and the second original image frame are two image frames in the second sample video with an interval less than a first number; Performing noise adding processing on the first original image frame and the second original image frame to obtain a first sample image frame and a second sample image frame with noise; The first original image frame is a true value image frame corresponding to the first sample image frame, and the second original image frame is a true value image frame corresponding to the second sample image frame.
3. The method according to claim 2, characterized in that The step of performing noise adding processing on the first original image frame and the second original image frame to obtain a first sample image frame and a second sample image frame with noise includes: Dividing the pixel values of the first original image frame and the second original image frame by a specified value respectively to obtain a third original image frame and a fourth original image frame; Noise is added to the third original image frame and the fourth original image frame respectively to obtain a first sample image frame and a second sample image frame.
4. The method according to claim 3, characterized in that Before adding noise to the third original image frame and the fourth original image frame respectively to obtain the first sample image frame and the second sample image frame, the method further includes: Acquiring a noise parameter of a sensor that acquires the second sample video; The adding noise to the third original image frame and the fourth original image frame respectively to obtain the first sample image frame and the second sample image frame comprises: Based on the noise parameter, noise is added to the third original image frame and the fourth original image frame respectively to obtain a first sample image frame and a second sample image frame.
5. The method according to claim 4, characterized in that The adding noise to the third original image frame and the fourth original image frame respectively based on the noise parameter to obtain a first sample image frame and a second sample image frame comprises: In a case where the noise parameter includes a mean value of Poisson noise, adding Poisson noise to the third original image frame and the fourth original image frame respectively according to the noise parameter to obtain a first sample image frame and a second sample image frame; In a case where the noise parameters include a mean value and a standard deviation of Gaussian noise, Gaussian noise is added to the third original image frame and the fourth original image frame respectively according to the noise parameters to obtain a first sample image frame and a second sample image frame; When the noise parameters include the mean of Poisson noise and the mean and standard deviation of Gaussian noise, Poisson noise and Gaussian noise are added to the third original image frame and the fourth original image frame respectively according to the noise parameters to obtain a first sample image frame and a second sample image frame.
6. The method according to any one of claims 1 to 5, characterized in that: The optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork and an output subnetwork; Inputting the first sample image frame and the second sample image frame into an optical flow information prediction network of an initial structure to obtain optical flow information of the second sample image frame relative to the first sample image frame includes: splicing the first sample image frame and the second sample image frame to obtain a first splicing result; Inputting the first concatenation result into the input subnetwork to obtain a first feature vector; Inputting the first feature vector into the feature extraction subnetwork to obtain a second feature vector; The second feature vector is input into the output sub-network to obtain the optical flow information of the second sample image frame relative to the first sample image frame.
7. The method according to claim 6, characterized in that The output sub-network includes: an activation layer, an amplification layer, and an upsampling layer; Inputting the second feature vector into the output subnetwork to obtain optical flow information of the second sample image frame relative to the first sample image frame includes: Inputting the second feature vector into the activation layer to obtain first optical flow information to be used; Inputting the first optical flow information to be used into the amplification layer to obtain amplified first optical flow information to be used; The amplified first optical flow information to be used is input into the upsampling layer to obtain the optical flow information of the second sample image frame relative to the first sample image frame.
8. An image enhancement method, characterized in that: The method comprises: Acquire an image frame to be enhanced and a reference image frame; wherein the image frame to be enhanced and the reference image frame are two image frames in the same video whose interval is less than a second number; Inputting the image frame to be enhanced and the reference image frame into a pre-trained optical flow information prediction network to obtain optical flow information of the reference image frame relative to the image frame to be enhanced; wherein the optical flow information prediction network is trained based on the method of any one of claims 1 to 7 above; Mapping the reference image frame based on the obtained optical flow information to obtain a mapping result of the reference image frame; The mapping result of the reference image frame is fused with the image frame to be enhanced to obtain an enhanced image frame of the image frame to be enhanced.
9. The method according to claim 8, characterized in that The optical flow information prediction network includes: an input subnetwork, a feature extraction subnetwork and an output subnetwork; Inputting the image frame to be enhanced and the reference image frame into a pre-trained optical flow information prediction network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced, including: Splicing the image frame to be enhanced and the reference image frame to obtain a second splicing result; Inputting the second concatenation result into the input subnetwork to obtain a third feature vector; Inputting the third feature vector into the feature extraction subnetwork to obtain a fourth feature vector; The fourth feature vector is input into the output sub-network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced.
10. The method according to claim 9, characterized in that The output sub-network includes: an activation layer, an amplification layer, and an upsampling layer; Inputting the fourth feature vector into the output sub-network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced includes: Inputting the fourth eigenvector into the activation layer to obtain second optical flow information to be used; Inputting the second optical flow information to be used into the amplification layer to obtain amplified second optical flow information to be used; The amplified second optical flow information to be used is input into the upsampling layer to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced.
11. An optical flow information prediction network training device, characterized in that: The device comprises: A first acquisition module is used to acquire a first sample image frame and a second sample image frame with noise; wherein the first sample image frame and the second sample image frame are two image frames in the first sample video with an interval less than a first number; A first input module, used for inputting the first sample image frame and the second sample image frame into an optical flow information prediction network of an initial structure to obtain optical flow information of the second sample image frame relative to the first sample image frame; a first mapping module, configured to map the true value image frame corresponding to the second sample image frame based on the obtained optical flow information, to obtain a mapping result of the true value image frame corresponding to the second sample image frame; wherein the true value image frame corresponding to the second sample image frame has the same image content as the second sample image frame and has no noise; An adjustment module is used to adjust the network parameters of the optical flow information prediction network of the initial structure based on the difference between the obtained mapping result and the true value image frame corresponding to the first sample image frame, until a preset convergence condition is reached to obtain a trained optical flow information prediction network; wherein the true value image frame corresponding to the first sample image frame is consistent with the image content of the first sample image frame and does not contain noise.
12. An image enhancement device, characterized in that: The device comprises: A second acquisition module is used to acquire an image frame to be enhanced and a reference image frame; wherein the image frame to be enhanced and the reference image frame are two image frames in the same video with an interval less than a second number; A second input module is used to input the image frame to be enhanced and the reference image frame into a pre-trained optical flow information prediction network to obtain the optical flow information of the reference image frame relative to the image frame to be enhanced; wherein the optical flow information prediction network is trained based on the device of claim 11; A second mapping module, used for mapping the reference image frame based on the obtained optical flow information to obtain a mapping result of the reference image frame; The fusion module is used to fuse the mapping result of the reference image frame with the image frame to be enhanced to obtain an enhanced image frame of the image frame to be enhanced.
13. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-10 when executing a program stored in a memory.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
15. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 10 when being executed by a processor.
Citation Information
Patent Citations
Video blind denoising method and device based on deep learning
CN111539879A
Video anomaly detection method and system based on generation of collaborative discrimination network
CN113011399A
Video denoising model processing method and device, computer equipment and storage medium
CN116977200A
Video enhancement method and device, electronic equipment, storage medium and program product
CN118014862A
Image processing method, image acquisition equipment, device, medium and product
CN118711026A