Image restoration model training method and device, electronic device, readable medium
By training a neural network model based on U-Net and deep residual networks and adjusting the weight coefficients, the problem of poor image quality in video surveillance systems under low light conditions was solved, achieving improved image quality and enhanced user experience without increasing hardware.
Patent Information
- Application Number
- CN202010388892.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-09
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2040-05-09
Smart Images

Figure CN113628123B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video surveillance technology, and in particular to a training method and apparatus for an image restoration model, an electronic device, and a readable medium. Background Technology
[0002] Video surveillance systems automatically identify and store monitored images, transmitting them back to the control host via various communication networks. The control host can then perform real-time viewing, recording, playback, and retrieval of the images, enabling mobile internet-based video surveillance. However, the quality of the acquired images varies due to the influence of ambient light. In well-lit conditions, clear images with high color fidelity can be provided; however, in low-light conditions, the image quality deteriorates significantly because the camera images underexposes the light, resulting in unclear images and severely impacting the user's monitoring experience.
[0003] Currently, various manufacturers have proposed different solutions to address these issues, such as upgrading hardware configurations or adding supplementary lighting sources. However, these solutions increase hardware costs. Software-based solutions primarily focus on enhancing the brightness or restoring the color of low-light images, without comprehensively considering the problems of low brightness, low contrast, loss of detail, and color distortion in surveillance scenarios. This results in the inability to obtain optimal video surveillance images in low-light conditions, and may even render surveillance ineffective, failing to acquire useful information. Summary of the Invention
[0004] This application provides a training method and apparatus for an image restoration model, an electronic device, and a readable medium, overcoming the problem in the prior art where the acquired images are unclear due to the influence of ambient light on video surveillance imaging, preventing users from obtaining useful information.
[0005] In a first aspect, embodiments of this application provide a training method for an image restoration model. The method includes: preprocessing training images to obtain a set of low-light image samples; determining weight coefficients of the image restoration model based on the low-light image samples in the low-light image sample set and the image restoration model, wherein the image restoration model is a neural network model determined based on a U-Net network and a deep residual network; adjusting the image restoration model according to the weight coefficients, and continuing to train the adjusted image restoration model using low-light image samples until the image restoration model can restore the parameters of all low-light image samples in the low-light image sample set to a preset range.
[0006] Secondly, embodiments of this application provide a training apparatus for an image restoration model, comprising: a preprocessing module for preprocessing training images to obtain a set of low-light image samples; a weight coefficient determination module for determining the weight coefficients of the image restoration model based on the low-light image samples in the low-light image sample set and the image restoration model, wherein the image restoration model is a neural network model determined based on a U-Net network and a deep residual network; and a training adjustment module for adjusting the image restoration model according to the weight coefficients, and continuing to train the adjusted image restoration model using low-light image samples until the image restoration model can restore the parameters of all low-light image samples in the low-light image sample set to a preset range.
[0007] Thirdly, embodiments of this application provide an electronic device comprising: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0009] The method provided in this application determines and continuously adjusts the weight coefficients of an image restoration model based on low-light image samples in a low-light image sample set and the model itself. This ultimately yields an image restoration model capable of restoring all low-light image samples in the set to a preset image sample. The image restoration model is a neural network model determined based on a U-Net network and a deep residual network. This method reduces the impact of ambient light on video surveillance imaging, improves and restores acquired low-light images, obtains normal image samples, avoids image blurring, and enhances the user experience. Furthermore, it does not require additional hardware, reducing hardware costs. Attached Figure Description
[0010] The accompanying drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain the application and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed example embodiments described with reference to the accompanying drawings.
[0011] Figure 1 This is a schematic flowchart illustrating a training method for an image restoration model according to an embodiment of this application.
[0012] Figure 2 This is a schematic diagram illustrating the data reconstruction of normalized image data in this application.
[0013] Figure 3 This is a flowchart illustrating the method for determining the weight coefficients of the image restoration model in this application.
[0014] Figure 4 This is a flowchart illustrating the training method for the image restoration subnetwork constructed based on the U-Net network in this application.
[0015] Figure 5 This is an exemplary flowchart illustrating the upsampling and downsampling of an image in this application.
[0016] Figure 6 This is a flowchart illustrating the specific implementation method of the deep residual network in this application.
[0017] Figure 7 This is a flowchart of the method for processing residual samples in the residual sample set in this application.
[0018] Figure 8 This is a flowchart illustrating a training method for an image restoration model provided in another embodiment of this application.
[0019] Figure 9 This is a block diagram of the composition of a training device for an image restoration model provided in an embodiment of this application.
[0020] Figure 10 This is a block diagram of an image restoration system provided in an embodiment of this application.
[0021] Figure 11 This is a structural diagram of an exemplary hardware architecture of an electronic device for the training method and apparatus of the image restoration model according to embodiments of this application. Detailed Implementation
[0022] To enable those skilled in the art to better understand the technical solution of this application, the following describes in detail, with reference to the accompanying drawings, a training method and apparatus for an image restoration model, an electronic device, and a computer-readable medium provided in this application.
[0023] Exemplary embodiments will be described more fully below with reference to the accompanying drawings; however, these exemplary embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will enable those skilled in the art to fully understand the scope of this application.
[0024] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0025] Figure 1 This diagram illustrates a flowchart of a training method for an image restoration model according to an embodiment of this application. This image restoration model training method can be applied to an image restoration model training apparatus. Figure 1 As shown, the training method for this image restoration model includes the following steps.
[0026] Step 110: Preprocess the training images to obtain a set of low-light image samples.
[0027] Among them, the low-light image samples in the low-light image sample set are image samples obtained when the detected illuminance is less than the preset illuminance threshold.
[0028] For example, when a camera captures images at night, the ambient light is very low due to the absence of sunlight. If the detected illuminance is less than a preset illuminance threshold (e.g., 10 lux), the captured image is a low-light image. The training images captured by the camera can include low-light images and preset reference images. By preprocessing the training images, a low-light image sample set and a preset reference image sample set can be obtained. The reference image samples in the preset reference image sample set are image samples obtained when the detected illuminance is greater than or equal to the preset illuminance threshold.
[0029] The training images use RGB color mode. It should be noted that other color modes are also possible; the RGB mode used in this application is a preferred color mode, and can be set according to actual conditions. This is merely an example; other undescribed color modes are also within the scope of this application and will not be elaborated upon further.
[0030] In some specific implementations, step 110 may include: normalizing the training images to obtain normalized image data; sampling at intervals in the row and column directions of the normalized image data to obtain M-dimensional data, where M is an integer greater than or equal to 1; and performing data augmentation on each dimension of the data to obtain a low-light image sample set and a preset reference image sample set.
[0031] Specifically, the three-channel images captured by the camera (e.g., RGB images, i.e., R-channel images, G-channel images, and B-channel images) are first normalized to reduce the spatial range of the images (e.g., reduce the spatial range of the images to the range of [0, 1]) to obtain normalized image data.
[0032] Then, the normalized image data in each channel is reassembled. Figure 2 This is a schematic diagram illustrating the data reconstruction of normalized image data in this application. For example... Figure 2 As shown, the image in each channel is processed to reduce its dimensionality, that is, the image in each channel is sampled at intervals in its row and column directions to ensure that every pixel is sampled. Figure 2 In the process, the image in each channel is processed four times. Downsampling processes the image data within each channel into... The data, where H represents the original height of the image and W represents the original width of the image, is used to process the three-channel image into... By reducing the spatial resolution of the image data without losing image information, the spatial information is transferred to different dimensions for processing, achieving lossless downsampling of the image data. This effectively reduces the computational complexity of subsequent networks.
[0033] Finally, the image data is randomly divided into blocks of a certain size, and data augmentation is performed on each dimension of the data of each image block (e.g., random flipping or rotation of the data) to obtain a set of low-light image samples and a preset reference image sample set, thereby improving the richness of the data.
[0034] Step 120: Determine the weight coefficients of the image restoration model based on the low-light image samples in the low-light image sample set and the image restoration model.
[0035] It should be noted that the image restoration model described herein is a neural network model determined based on the U-Net network and the deep residual network (ResNet). In specific implementations, a first neural network model determined based on the U-Net network and a second neural network model determined based on ResNet can be cascaded to achieve the function of the image restoration model; alternatively, the image restoration model can be implemented in other different forms based on the functions of the U-Net network and ResNet. The above specific implementation methods of the image restoration model are merely illustrative examples and can be specifically set according to the specific implementation. Other undescribed specific implementation methods of image restoration models are also within the scope of protection of this application and will not be elaborated upon here.
[0036] Step 130: Adjust the image restoration model according to the weight coefficients, and continue to train the adjusted image restoration model using low-light image samples until the image restoration model can restore the parameters of all low-light image samples in the low-light image sample set to the preset range.
[0037] Specifically, based on the first weight coefficient obtained initially, the image restoration model is adjusted (e.g., relevant parameters in the image restoration model can be adjusted) to obtain an adjusted image restoration model. This allows the image restoration model to obtain restored image samples that are closer to normal images when used to restore low-light images. Then, low-light image samples are input into the adjusted image restoration model for training to obtain the second weight coefficient corresponding to the adjusted image restoration model. This second weight coefficient is then used to further adjust the adjusted image restoration model, repeating this process multiple times until the image restoration model can restore the parameters of all low-light image samples in the low-light image sample set to a preset range. After multiple training and adjustments, the final image restoration model can meet the restoration requirements of low-light image samples. That is, when low-light image samples are input into the final image restoration model, normal brightness image samples with high contrast, low noise, rich details, and high color fidelity can be obtained.
[0038] In practice, multiple validation sample sets can be used to validate the image restoration model after each adjustment, accumulating multiple validation results. Then, based on the adjustment trend of multiple validation results, it can be determined whether the current adjustment of the image restoration model according to the weight coefficients is appropriate, and the image restoration model can be fine-tuned to make it closer to the final desired image restoration model. The validation sample set includes multiple validation image samples.
[0039] In this embodiment, based on low-light image samples in a low-light image sample set and an image restoration model, the weight coefficients of the image restoration model are determined and continuously adjusted. Ultimately, an image restoration model capable of restoring all low-light image samples in the low-light image sample set to a preset image sample is obtained. This image restoration model is a neural network model determined based on a U-Net network and a deep residual network. This reduces the impact of ambient light on video surveillance imaging, improves and restores acquired low-light images to obtain normal image samples, avoids image blurring, and enhances the user experience. Furthermore, it does not require additional hardware, reducing hardware costs.
[0040] Figure 3 This diagram illustrates the method for determining the weight coefficients of the image restoration model in this application. Step 120 can be performed as follows: Figure 3 The method shown is used to achieve this, specifically including steps 121 to 123.
[0041] Step 121: Based on the U-Net network, perform lossless data reconstruction and data feature fusion on the low-light image samples in the low-light image sample set to obtain the primary restored image sample set.
[0042] In specific implementation, the following operations are performed on each low-light image sample in the low-light image sample set: First, the low-light image sample is reconstructed without loss (e.g., the image data is downsampled without loss). Then, the downsampled data and the data after deconvolution are fused by combining the downsampling results at different levels and the corresponding data features. Finally, the fused data is upsampled without loss to obtain the primary restored image sample. The primary restored image samples obtained through the above processing constitute the primary restored image sample set.
[0043] For example, Figure 4 This is a flowchart illustrating the training method for the image restoration subnetwork built based on the U-Net network in this application. Figure 4 As shown, this includes the encoding and decoding processes for image samples. Figure 4 The left side represents the encoding process, which involves preprocessing the data obtained in step 110. The data is input into an image restoration subnetwork built on a U-Net network, where a first layer of downsampling and data reconstruction is performed. For example, for... The data undergoes a 3x3 convolution and activation processing to obtain... Data, and then on The data is downsampled without loss to obtain The data is processed to obtain the data features corresponding to the first layer of downsampling. The data after the first layer of downsampling is then subjected to a second layer of downsampling and data reconstruction. The sampling process is similar to the first layer of downsampling, both being lossless downsampling. This process is repeated sequentially, for a total of four lossless downsampling operations, transforming the initial input... Data sampling to data.
[0044] Symmetrically, in Figure 4 The right side describes the decoding process of the encoded data. Specifically, it describes the decoding process of the downsampled data. The data is decoded, and corresponding to the downsampling of each layer on the left, four deconvolution upsampling operations are performed sequentially from bottom to top. Simultaneously, the high-level data features after deconvolution are concatenated with the low-level data features of the same coding stage, achieving the fusion of data features at different scales, thereby enhancing the image. For example, the first upsampling process includes: The data undergoes two 3x3 convolution calculations and activation processing to obtain... Data, and then on The data is deconvolved to obtain The data is then concatenated with the low-level features of the same stage of coding to obtain... The data is processed as described above. Four deconvolution upsampling and concatenation operations are performed sequentially to obtain the restored image data H*W*3. Finally, a lossless upsampling and data reconstruction operation is performed to obtain the initial restored image sample.
[0045] In the image restoration subnetwork built on the U-Net network, the upsampling and downsampling of image data are achieved through data reconstruction, which maps spatial data blocks to data in different dimensions.
[0046] Figure 5 This is a schematic diagram illustrating an exemplary process of upsampling and downsampling an image. For example... Figure 5 As shown, the spatial data block on the left is downsampled to obtain corresponding data in different dimensions on the right, and symmetrically, an upsampling operation is performed on the data. For example, this can be achieved using the neural network library function `Depth_To_Space`. Figure 5 Upsampling in the code, and the library function Space_To_Depth are used to implement it. Figure 5 Downsampling is used to ensure that no information is lost during the upsampling and downsampling processes, thereby improving the accuracy of data recovery.
[0047] Step 122: Input the primary restored image samples from the primary restored image sample set into the detail enhancement sub-model in the image restoration model for training, and obtain the enhanced restored image sample set.
[0048] Among them, the detail enhancement sub-model is a neural network model determined based on the deep residual network.
[0049] Step 123: Determine the weight coefficients of the image restoration model based on the enhanced restored image sample set and the preset reference image sample set.
[0050] It should be noted that the preset reference image sample set includes preset reference image samples, which can be image samples obtained when the illuminance is greater than or equal to a preset illuminance threshold. After obtaining the enhanced restored image sample, it needs to be compared with the preset reference image sample to determine whether the obtained enhanced restored image sample improves the original low-light image sample, and then determine the weight coefficients of the image restoration model to ensure that the training trend of the image restoration model is towards convergence.
[0051] In this embodiment, by using the U-Net network, lossless data reconstruction and feature fusion are performed on the low-light image samples in the low-light image sample set, enabling the low-light image samples to undergo preliminary processing and obtain a primary restored image sample set. Then, the primary restored image samples in the primary restored image sample set are input into the detail enhancement sub-model in the image restoration model for training. Lossless upsampling and downsampling operations are used to obtain an enhanced restored image sample set, ensuring that no data information is lost. Based on the enhanced restored image sample set and a preset reference image sample set, the weight coefficients of the image restoration model are determined. By adjusting the weight coefficients, the image restoration model is adjusted so that the adjusted image restoration model is closer to the requirements, that is, it can obtain restored image samples that are closer to normal images.
[0052] In one specific implementation, step 122 can also be implemented in the following manner, specifically including: performing lossless downsampling and data recombination on the primary restored image samples to obtain a primary detail enhancement training sample set; processing the primary detail enhancement training samples in the primary detail enhancement training sample set according to the deep residual network to obtain a residual sample set; and upsampling and data recombination on the residual samples in the residual sample set to obtain an enhanced restored image sample set.
[0053] For example, Figure 6 A flowchart illustrating the specific implementation method of deep residual networks is shown. Figure 6 As shown, the primary restored image samples (e.g., H*W*3 data) are input into a detail enhancement sub-model built on a ResNet network. The H*W*3 data is first downsampled twice (e.g., as shown in the diagram). Figure 5 The downsampling shown) and data reconstruction, for example, performing one step on H*W*3 data. Figure 5 After downsampling as shown, the downsampled result (i.e. Perform a 1x1 convolution operation on the data to obtain the convolution result (i.e., (data), and then perform the convolution operation on the result again as follows Figure 5 The downsampling shown in the figure obtains Data. At this point, the input primary restored image sample will undergo a 4x lossless downsampling and data reconstruction operation, reducing its resolution to [missing value]. By reducing the resolution of the primary restored image samples, the complexity of the data is reduced, making subsequent data processing easier.
[0054] Then, Data is input into N residual blocks for processing (e.g., into 14 cascaded residual blocks for processing, resulting in a deeper residual network model), where N is an integer greater than or equal to 1. The outputs of each residual block are subjected to two 1*1 convolution operations and upsampling to restore the image to its original resolution (i.e., obtaining new H*W*3 data). Finally, the primary restored image samples are combined with the data processed by the detail enhancement sub-model built on the ResNet network (i.e., the new H*W*3 data) to obtain a high-quality output image.
[0055] Through the above processing, the information of the primary restored image sample is not lost. At the same time, by processing multiple residual blocks, and with each residual block having the same data resolution, the details of the image are restored, and the processing of data details is improved, so that the details of the primary restored image sample can be enhanced, thereby improving the restoration quality of the image sample.
[0056] In some specific implementations, based on the deep residual network, the primary detail enhancement training samples in the primary detail enhancement training sample set are processed to obtain the residual sample set. This includes: performing N convolution and activation processes on the residual samples in the primary detail enhancement training sample set to obtain the primary residual sample set, where N is an integer greater than or equal to 2; and summing the primary residual samples in the primary residual sample set with the primary detail enhancement training samples in their corresponding primary detail enhancement training sample sets to obtain the residual sample set.
[0057] Through the above processing, the depth of the network can be increased without worrying about the loss of image information, and thus more information can be learned by expanding the depth of the network.
[0058] Figure 7 This diagram illustrates a method for processing residual samples in a residual sample set. Figure 7 As shown, the specific steps include the following.
[0059] Step 701: Input the data to be processed.
[0060] The data to be processed can be a residual sample from the primary detail enhancement training sample set.
[0061] Step 702: Perform a 3*3 convolution on the data to be processed to obtain the first processed data.
[0062] Step 703: Activate the first processed data to obtain activated data.
[0063] Step 704: Perform a 3*3 convolution on the activation data to obtain the second processed data.
[0064] Step 705: Add the second processed data and the data to be processed together to obtain the output data.
[0065] Through the above processing, the depth of the network can be increased without losing information in the data to be processed; by expanding the depth of the network, more data information is added, more information is learned, and the details of the image data are enhanced.
[0066] In some specific implementations, step 123 includes: comparing the enhanced restored image samples in the enhanced restored image sample set with the reference image samples in the corresponding preset reference image sample set to obtain a set of correction parameters; and determining the weight coefficients of the image restoration model based on the set of correction parameters.
[0067] For example, a set of correction parameters can be obtained through certain processing functions in a neural network. Alternatively, the training trend of an image restoration model can be obtained by comparing enhanced restored image samples with reference image samples. Based on this training trend, a set of correction parameters can be determined. By adjusting the parameters in the set of correction parameters, the weight coefficients of the image restoration model can be determined, ensuring that the adjustment trend of the weight coefficients is towards convergence.
[0068] It should be noted that the above methods for obtaining the set of correction parameters are only illustrative examples and can be set according to the specific implementation scheme. Other methods for obtaining the set of correction parameters not mentioned are also within the scope of protection of this application and will not be elaborated here.
[0069] In some specific implementations, the enhanced restored image samples in the enhanced restored image sample set and the reference image samples in the corresponding preset reference image sample set are compared one by one to obtain the set of correction parameters. This includes: determining a joint objective function based on the structural similarity function, the enhanced restored image sample set, and the preset reference image sample set; and using the joint objective function to calculate the enhanced restored image samples in the enhanced restored image sample set and the reference image samples in the corresponding preset reference image sample set to obtain the set of correction parameters.
[0070] The joint objective function is expressed by the following formula:
[0071] L=αL enh +(1-α)L ssim ,
[0072] Where L represents the value of the joint objective function, m represents the width of the image, n represents the height of the image, (i,j) represents the coordinates of a pixel in a coordinate system constructed using the width and height of the image, and I′ represents the restored enhanced image sample. GT L represents a reference image sample.enh This represents the restored enhanced image sample I′ and the reference image sample I. GT The norm of the absolute value of the difference, L ssim L represents the structural similarity function. ssim Specifically defined as Among them, x and y are two image samples with different structures, μ x μ y Let σ represent the average values of x and y, respectively. x , σ y Let σ represent the variances of x and y, respectively. xy Let x represent the covariance of x and y, α represent the weight coefficients of the image restoration model, and N represent the number of image blocks into which the image sample is divided.
[0073] By using a joint objective function to calculate the weight coefficients of the image restoration model, and adjusting these weight coefficients, it is ensured that the value of the joint objective function obtained in subsequent calculations tends to converge relative to the value of the previous joint objective function. After multiple calculations and adjustments to the weight coefficients, the final image restoration model can meet the restoration requirements of low-light image samples. That is, by inputting low-light image samples into the final image restoration model, normal image samples obtained under normal lighting conditions can be obtained that are infinitely close to those obtained under normal lighting conditions.
[0074] Figure 8 This diagram illustrates a flowchart of a training method for an image restoration model provided in another embodiment of this application. Figure 8 As shown, the specific steps may include the following.
[0075] Step 810: Preprocess the training images to obtain a set of low-light image samples.
[0076] Step 820: Determine the weight coefficients of the image restoration model based on the low-light image samples in the low-light image sample set and the image restoration model.
[0077] Step 830: Adjust the image restoration model according to the weight coefficients, and continue to train the adjusted image restoration model using low-light image samples until the image restoration model can restore the parameters of all low-light image samples in the low-light image sample set to the preset range.
[0078] It should be noted that steps 810 to 830 in this embodiment are the same as steps 110 to 130 in the previous embodiment, and will not be repeated here.
[0079] Step 840: Obtain the test image sample set.
[0080] The test sample set includes low-light test image samples. These low-light test image samples are image samples obtained when the illuminance is less than a preset illuminance threshold. For example, image samples collected at night.
[0081] Step 850: Use an image restoration model to restore the test low-light image sample to obtain the restored image sample.
[0082] Specifically, the performance parameters of the restored image samples are superior to those of the test low-light image samples. The performance parameter set includes at least one of the following: image contrast, image brightness, image resolution, and image color fidelity.
[0083] By using an image restoration model to restore test low-light image samples, the training results of the image restoration model can be judged. When the image contrast, brightness, resolution, or color reproduction of the restored image sample is enhanced, it can be determined that the current image restoration model meets the requirements. Ideally, the restored image sample can be infinitely close to the normal image sample (e.g., an image acquired during the day), ensuring that it can provide users with good image quality and improve the user experience.
[0084] Figure 9 This is a block diagram illustrating the composition of a training apparatus for an image restoration model provided in this application embodiment. Specific implementations of this apparatus can be found in the descriptions of the various method embodiments described above; repeated details will not be repeated. It is worth noting that the specific implementation of the apparatus in this embodiment is not limited to the above embodiments, and other undescribed embodiments are also within the protection scope of this apparatus.
[0085] like Figure 9 As shown, the training device for the image restoration model specifically includes: a preprocessing module 901 for preprocessing the training images to obtain a set of low-light image samples; a weight coefficient determination module 902 for determining the weight coefficients of the image restoration model based on the low-light image samples in the low-light image sample set and the image restoration model, wherein the image restoration model is a neural network model determined based on a U-Net network and a deep residual network; and a training adjustment module 903 for adjusting the image restoration model according to the weight coefficients, and continuing to train the adjusted image restoration model using low-light image samples until the image restoration model can restore the parameters of all low-light image samples in the low-light image sample set to a preset range.
[0086] In this embodiment, a weight coefficient determination module determines the weight coefficients of the image restoration model based on low-light image samples in the low-light image sample set and the image restoration model. A training and adjustment module continuously adjusts these weight coefficients to ultimately obtain an image restoration model capable of restoring all low-light image samples in the low-light image sample set to a preset image sample. This image restoration model is a neural network model determined based on a U-Net network and a deep residual network. This reduces the impact of ambient light on video surveillance imaging, improves and restores acquired low-light images to obtain normal image samples, avoids image blurring, and enhances the user experience. Furthermore, it does not require additional hardware, reducing hardware costs.
[0087] It is worth noting that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, this invention is not limited to the specific configurations and processes described in the above embodiments and shown in the figures. For ease of description and brevity, detailed descriptions of known methods are omitted here. The specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0088] Figure 10 This diagram illustrates a block diagram of an image restoration system provided in an embodiment of this application. The system specifically includes: a data acquisition device 1001, an image signal processing device 1002, a low-light image restoration device 1003, a video image encoding / decoding device 1004, and a detection device 1005.
[0089] The data acquisition device 1001 is used to acquire raw image data through a camera; the image signal processing device 1002 is used to perform image processing on the raw image data input by the data acquisition device 1001; the low-light image restoration device 1003 is used to construct a low-light image restoration model according to the training method of the image restoration model in the above method embodiment, and use the low-light image restoration model to restore low-light images; the video image encoding and decoding device 1004 is used to encode and decode images; and the detection device 1005 is used to detect the image output by the image signal processing device 1002 to determine whether the image is a low-light image.
[0090] Specifically, after the data acquisition device 1001 acquires the original image data through the camera, it inputs the original image data into the image signal processing device 1002 for a series of image processing operations, such as removing noise from the original image data or segmenting the original image data, to obtain the processed image, and outputs the processed image to the detection device 1005. The detection device 1005 detects the processed image. If it detects that the obtained image sample was acquired when the illuminance was greater than or equal to a preset illuminance threshold (e.g., a normal image), it directly inputs the image sample into the video image encoding / decoding device 1004, performs encoding / decoding operations on the image sample, and outputs the image for the user to view. If it detects that the obtained image sample was acquired when the illuminance was less than the preset illuminance threshold (e.g., a low-light image), it first outputs the image to the low-light image restoration device 1003, enabling the low-light image restoration device 1003 to restore the low-light image and obtain a restored high-light image. The parameters of the high-light image are better than those of the low-light image. Ideally, the high-light image is a normal image. Then, the high-light image is input into the video image encoding / decoding device 1004, performs encoding / decoding operations on the high-light image, and then outputs the image for the user to view.
[0091] Through the above processing, high-quality video surveillance images can be provided to users regardless of whether the lighting conditions are sufficient during the day or low light conditions at night, thereby improving the user experience.
[0092] Figure 11 This is a structural diagram of an exemplary hardware architecture of an electronic device for the training method and apparatus of the image restoration model according to embodiments of this application.
[0093] like Figure 11 As shown, the electronic device 1100 includes an input device 1101, an input interface 1102, a central processing unit 1103, a memory 1104, an output interface 1105, and an output device 1106. The input interface 1102, the central processing unit 1103, the memory 1104, and the output interface 1105 are interconnected via a bus 1107. The input device 1101 and the output device 1106 are connected to the bus 1107 via the input interface 1102 and the output interface 1105, respectively, and are thus connected to other components of the electronic device 1100.
[0094] Specifically, input device 1101 receives input information from the outside and transmits the input information to central processing unit 1103 through input interface 1102; central processing unit 1103 processes the input information based on computer-executable instructions stored in memory 1104 to generate output information, temporarily or permanently stores the output information in memory 1104, and then transmits the output information to output device 1106 through output interface 1105; output device 1106 outputs the output information to the outside of computing device 1100 for user use.
[0095] In one embodiment, Figure 11 The electronic device 1100 shown can be implemented as an electronic device that may include: a memory configured to store a program; and a processor configured to run the program stored in the memory to execute a training method for any of the image restoration models described in the above embodiments.
[0096] According to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network, and / or installed from a removable storage medium.
[0097] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0098] This document has disclosed exemplary embodiments, and although specific terminology is used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this application as set forth by the appended claims.
Claims
1. A training method for an image restoration model, comprising: Preprocess the training images to obtain a set of low-light image samples; Based on the low-light image samples in the low-light image sample set and the image restoration model, the weight coefficients of the image restoration model are determined, wherein the image restoration model is a neural network model determined based on the U-Net network and the deep residual network; Based on the weight coefficients, the image restoration model is adjusted, and the adjusted image restoration model is trained again using the low-light image samples until the image restoration model can restore the parameters of all the low-light image samples in the low-light image sample set to the preset range. The step of determining the weight coefficients of the image restoration model based on the low-light image samples in the low-light image sample set and the image restoration model includes: Based on the U-Net network, lossless data reconstruction and data feature fusion are performed on the low-light image samples in the low-light image sample set to obtain a primary restored image sample set; The primary restored image samples in the primary restored image sample set are input into the detail enhancement sub-model in the image restoration model for training to obtain an enhanced restored image sample set, wherein the detail enhancement sub-model is a neural network model determined based on the deep residual network; Based on the enhanced restored image sample set and the preset reference image sample set, the weight coefficients of the image restoration model are determined; The step of inputting the primary restored image samples from the primary restored image sample set into the detail enhancement sub-model of the image restoration model for training to obtain the enhanced restored image sample set includes: The primary restored image samples are subjected to lossless downsampling and data reconstruction to obtain a primary detail enhancement training sample set; Based on the deep residual network, the primary detail enhancement training samples in the primary detail enhancement training sample set are processed to obtain the residual sample set; The residual samples in the primary detail enhancement training sample set are subjected to N convolution and activation processes to obtain the primary residual sample set, where N is an integer greater than or equal to 2; The primary residual samples in the primary residual sample set are summed with the primary detail enhancement training samples in the corresponding primary detail enhancement training sample set to obtain the enhanced restored image sample set.
2. The method according to claim 1, characterized in that, The step of determining the weight coefficients of the image restoration model based on the enhanced restored image sample set and the preset reference image sample set includes: By comparing each enhanced restored image sample in the enhanced restored image sample set with the corresponding reference image sample in the preset reference image sample set, a set of correction parameters is obtained. The weight coefficients of the image restoration model are determined based on the set of correction parameters.
3. The method according to claim 2, characterized in that, The step of comparing each enhanced restored image sample in the enhanced restored image sample set with the corresponding reference image sample in the preset reference image sample set to obtain a set of correction parameters includes: A joint objective function is determined based on the structural similarity function, the enhanced restored image sample set, and the preset reference image sample set; Using the joint objective function, the enhanced restored image samples in the enhanced restored image sample set and the reference image samples in the corresponding preset reference image sample set are calculated to obtain the set of correction parameters.
4. The method according to any one of claims 1 to 3, characterized in that, The preprocessing of the training images to obtain a low-light image sample set includes: The training images are normalized to obtain normalized image data; The normalized image data is sampled at intervals along both the row and column directions to obtain M dimensional data, where M is an integer greater than or equal to 1. Data augmentation is performed on the data in each dimension to obtain the low-light image sample set and the preset reference image sample set.
5. The method according to any one of claims 1 to 3, characterized in that, After the step where the image restoration model can restore the parameters of all low-light image samples in the low-light image sample set to a preset range, the method further includes: Obtain a set of test image samples, which includes test low-light image samples; The image restoration model is used to restore the test low-light image sample to obtain a restored image sample, wherein the parameters in the performance parameter set of the restored image sample are better than the performance parameters of the test low-light image sample.
6. The method according to claim 5, characterized in that, The parameters in the set of performance parameters include at least one of image contrast, image brightness, and image color fidelity.
7. The method according to any one of claims 1 to 3, characterized in that, The low-light image samples in the low-light image sample set are image samples obtained when the detected illuminance is less than a preset illuminance threshold, and the reference image samples in the preset reference image sample set are image samples obtained when the detected illuminance is greater than or equal to the preset illuminance threshold.
8. The method according to any one of claims 1 to 3, characterized in that, The training images are in RGB color mode.
9. A training device for an image restoration model, comprising: The preprocessing module is used to preprocess the training images to obtain a set of low-light image samples. The weight coefficient determination module is used to determine the weight coefficients of the image restoration model based on the low-light image samples in the low-light image sample set and the image restoration model, wherein the image restoration model is a neural network model determined based on U-Net network and deep residual network; The training and adjustment module is used to adjust the image restoration model according to the weight coefficients, and continue to train the adjusted image restoration model using the low-light image samples until the image restoration model can restore the parameters of all the low-light image samples in the low-light image sample set to a preset range. The weight coefficient determination module is specifically used to perform lossless data reconstruction and data feature fusion on the low-light image samples in the low-light image sample set based on the U-Net network to obtain a primary restored image sample set; input the primary restored image samples in the primary restored image sample set into the detail enhancement sub-model in the image restoration model for training to obtain an enhanced restored image sample set, wherein the detail enhancement sub-model is a neural network model determined based on the deep residual network; and determine the weight coefficients of the image restoration model based on the enhanced restored image sample set and a preset reference image sample set. The step of inputting the primary restored image samples from the primary restored image sample set into the detail enhancement sub-model of the image restoration model for training to obtain an enhanced restored image sample set includes: performing lossless downsampling and data reconstruction on the primary restored image samples to obtain a primary detail enhancement training sample set; processing the primary detail enhancement training samples in the primary detail enhancement training sample set according to the deep residual network to obtain a residual sample set; performing N convolution and activation processes on the residual samples in the primary detail enhancement training sample set to obtain a primary residual sample set, where N is an integer greater than or equal to 2; and performing a summation operation on the primary residual samples in the primary residual sample set and their corresponding primary detail enhancement training samples in the primary detail enhancement training sample set to obtain the enhanced restored image sample set.
10. An electronic device comprising: One or more processors; A storage device having stored one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method according to any one of claims 1 to 8.
11. A computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Low-illumination reduction method based on multi-stage variational auto-encoder
CN110163815A