Model training method, image reconstruction method, related device, equipment, storage medium and computer program product

By training an image reconstruction method containing multi-network models, using image pairs of different lighting intensities for training, the problem of image reconstruction of small target objects under low light is solved, and high-quality image reconstruction effect is achieved.

CN120070856APending Publication Date: 2025-05-30CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125504.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In low-light scenarios, how to effectively achieve image reconstruction of small target objects, especially in the problems of image details, increased noise and reduced contrast.

Method used

By determining a training data set containing different illumination intensity image pairs for the target object, a multi-network model including denoising, super-resolution, reflection component, and irradiation component association is trained. The model uses parallel branch networks and decoder networks to perform image denoising and super-resolution processing, and generates high-quality reconstructed images through feature separation and stitching.

Benefits of technology

High-quality image reconstruction of small target objects under low light conditions is achieved, which improves the image detail retention and contrast, and significantly improves image availability in low light scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070856A_ABST
    Figure CN120070856A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, an image reconstruction method, a model training device, an image reconstruction device, first equipment, second equipment, a storage medium and a computer program product. The model training method comprises the steps that a first training data set is determined, and each sample in the first training data set comprises a first image and a second image for a target object; training a first model by using the first training data set, the first model being used for performing image reconstruction on the input image; the first model comprises a first network, a second network and a third network, and the first network is used for performing de-noising processing and super-resolution processing on an input image to obtain first feature data; the second network is used for obtaining second feature data and third feature data based on the first feature data; and the third network is used for splicing the second feature data and the third feature data to obtain fourth feature data, and performing convolution processing on the fourth feature data to obtain a reconstructed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and particularly to a model training method, an image reconstruction method, related devices, equipment, storage media, and computer program products. Background Art

[0002] A low-light scene refers to an image acquisition environment (i.e., an image acquisition environment) where the light is insufficient or the ambient light intensity is low. In such a scene, due to insufficient light, the image quality will be significantly limited, which will lead to problems such as loss of image details, increased noise, and reduced contrast. However, given the widespread existence of low-light scenes in various fields, such as night surveillance, unmanned driving, mobile photography, etc., researchers have been working on improving the quality and usability of low-light images.

[0003] Small target objects refer to objects with small sizes and are relatively difficult to detect compared to the background. In various fields such as aerospace, military intelligence, and intelligent surveillance, the detection and recognition of small target objects have always been extremely challenging tasks. As a basic task of small target detection, the image reconstruction of small target objects has important fundamental significance and is conducive to promoting the better progress of subsequent downstream tasks.

[0004] Given the widespread existence of low-light scenes and the practical value of small target objects in various scenes, how to achieve the image reconstruction of small target objects in low-light scenes has become an urgent problem to be solved. Summary of the Invention

[0005] To solve the related technical problems, embodiments of the present application provide a model training method, an image reconstruction method, related devices, equipment, storage media, and computer program products.

[0006] The technical solution of the embodiments of the present application is implemented as follows:

[0007] Embodiments of the present application provide a model training method, including:

[0008] Determine a first training data set, where each sample in the first training data set includes a first image and a second image of a target object, the illumination intensity corresponding to the first image is less than the illumination intensity corresponding to the second image, and the first image is an image obtained by processing the second image through a first image processing process, and the first image processing process is used to reduce the illumination intensity corresponding to the image;

[0009] Using the first training dataset, train a first model, where the first model is used to perform image reconstruction on an input image; wherein, the first model includes a first network, a second network, and a third network, the first network is used to perform denoising processing and super-resolution processing on the input image to obtain first feature data; the second network is used to obtain second feature data and third feature data based on the first feature data, the second feature data is associated with the reflection component of the input image, and the third feature data is associated with the illumination component of the input image; the third network is used to splice the second feature data and the third feature data to obtain fourth feature data, and obtain the reconstructed image by performing convolution processing on the fourth feature data.

[0010] In the above solution, the first network includes a parallel first branch network and second branch network, and also includes a decoder network; the first branch network is used to perform denoising processing on the input image to obtain fifth feature data; the second branch network is used to perform super-resolution processing on the input image to obtain sixth feature data; the decoder network is used to decode seventh feature data to obtain the first feature data; the seventh feature data is the feature data obtained by adding the fifth feature data and the sixth feature data.

[0011] In the above solution, the first network further includes a preprocessing component, and the preprocessing component is used to segment the input image into multiple patches and input the multiple patches into the first branch network and the second branch network;

[0012] Both the first branch network and the second branch network include a plurality of multi-layer perceptron mixers (MLP-Mixer, MultiLayer Perceptron-Mixer), and the number of MLP-Mixers is the same as the number of patches; each MLP-Mixer includes a sequence mixing (token-mixing) component and a channel mixing (channel-mixing) component, and the token-mixing component and the channel-mixing component are respectively used to perform feature extraction in different dimensions of the corresponding patch, and the token-mixing component includes a sparse multi-layer perceptron (SMLP, Sparse MultiLayer Perceptron);

[0013] The decoder network is a convolutional neural network (CNN) constructed based on transposed convolution technology and the parametric rectified linear unit (PReLU) activation function.

[0014] In the above solution, the second network is a dense convolutional network (DenseNet) constructed based on the bottleneck structure in the residual network (ResNet) and using strided convolution. The second network includes a parallel third branch network and a fourth branch network. The third branch network is used to encode the reflection component of the input image based on the first feature data to obtain the second feature data. The fourth branch network is used to encode the illumination component of the input image based on the first feature data to obtain the third feature data.

[0015] In the above solution, the first loss function corresponding to the first model is a function obtained by weighting the second loss function corresponding to the first network, the third loss function corresponding to the second network, the fourth loss function corresponding to the second network, and the fifth loss function corresponding to the third network. The third loss function is associated with the reflection component of the input image, and the fourth loss function is associated with the illumination component of the input image. Training the first model includes:

[0016] In each round of training, using the dynamic weight average (DWA) algorithm to adjust one or more of the weights corresponding to the second loss function, the weight corresponding to the third loss function, the weight corresponding to the fourth loss function, and the weight corresponding to the fifth loss function to obtain an updated first loss function.

[0017] Using the updated first loss function to adjust the model parameters of the first model.

[0018] The embodiment of the present application further provides an image reconstruction method, including:

[0019] Obtaining an image to be processed;

[0020] Using the first model to perform image reconstruction on the image to be processed to obtain a reconstructed image, where the first model is a model trained by using any of the above model training methods.

[0021] The embodiment of the present application further provides a model training device, including:

[0022] A first processing unit, configured to determine a first training data set, each sample in the first training data set including a first image and a second image of a target object, the illumination intensity corresponding to the first image being less than the illumination intensity corresponding to the second image, the first image being an image obtained by processing the second image through a first image processing process, the first image processing process being used to reduce the illumination intensity corresponding to the image;

[0023] A second processing unit, configured to train a first model by using the first training data set, the first model being used to perform image reconstruction on an input image; wherein, the first model includes a first network, a second network, and a third network, the first network being used to perform denoising processing and super-resolution processing on the input image to obtain first feature data; the second network being used to obtain second feature data and third feature data based on the first feature data, the second feature data being associated with the reflection component of the input image, and the third feature data being associated with the illumination component of the input image; the third network being used to splice the second feature data and the third feature data to obtain fourth feature data, and to obtain a reconstructed image by performing convolution processing on the fourth feature data.

[0024] An embodiment of the present application further provides an image reconstruction device, including:

[0025] An acquisition unit, configured to acquire an image to be processed;

[0026] A reconstruction unit, configured to perform image reconstruction on the image to be processed by using a first model to obtain a reconstructed image, the first model being a model trained by using any of the above model training methods.

[0027] An embodiment of the present application further provides a first device, including: a first processor and a first memory for storing a computer program that can run on the processor,

[0028] wherein, when the first processor is used to run the computer program, it executes the steps of any of the above model training methods.

[0029] An embodiment of the present application further provides a second device, including: a second processor and a second memory for storing a computer program that can run on the processor,

[0030] wherein, when the second processor is used to run the computer program, it executes the steps of any of the above image reconstruction methods.

[0031] An embodiment of the present application further provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any of the above model training methods, or implements the steps of any of the above image reconstruction methods.

[0032] An embodiment of the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of any of the above model training methods or implements the steps of any of the above image reconstruction methods.

[0033] The model training method, image reconstruction method, related device, equipment, storage medium and computer program product provided by the embodiments of the present application. The model training method includes: determining a first training data set, each sample in the first training data set includes a first image and a second image for a target object, the illumination intensity corresponding to the first image is less than the illumination intensity corresponding to the second image, the first image is an image obtained by processing the second image through a first image processing process, and the first image processing process is used to reduce the illumination intensity of the image; using the first training data set to train a first model, the first model is used to perform image reconstruction on the input image; wherein, the first model includes a first network, a second network and a third network, the first network is used to perform denoising processing and super-resolution processing on the input image to obtain first feature data; the second network is used to obtain second feature data and third feature data based on the first feature data, the second feature data is associated with the reflection component of the input image, and the third feature data is associated with the illumination component of the input image; the third network is used to splice the second feature data and the third feature data to obtain fourth feature data, and through convolutional processing on the fourth feature data, obtain the reconstructed image. The solution provided by the embodiments of the present application uses a specific data set (i.e., the first training data set) to train a first model for performing image reconstruction on the input image. Each sample in this data set includes an image pair (i.e., two images) for a specific target object and corresponding to different illumination intensities. The first model includes a network for performing denoising processing and super-resolution processing (i.e., the first network), a network for performing feature separation associated with the reflection component and illumination component of the image (i.e., the second network), and a network for obtaining the reconstructed image through further convolution (i.e., the third network); in this way, it is possible to effectively integrate various processing stages such as denoising, super-resolution, feature separation, and reconstruction for the image and / or feature data, so as to obtain high-quality image reconstruction results, and further improve the quality and usability of the reconstructed image in some specific scenarios where the image reconstruction effect of the related technology is limited. For example, it can effectively implement the image reconstruction of small target objects in low-light scenarios. Description of the Drawings

[0034] Figure 1 It is a schematic flowchart of the model training method of the embodiment of the present application;

[0035] Figure 2 Schematic flow diagram of the image reconstruction method according to the embodiment of the present application;

[0036] Figure 3 Schematic diagram of the overall structure of the neural network model of the application example of the present application;

[0037] Figure 4 Schematic diagram of the denoising and super-resolution joint network structure of the application example of the present application;

[0038] Figure 5 Schematic diagram of the Token Mixing structure of the application example of the present application;

[0039] Figure 6 Schematic diagram of the Channel Mixing structure of the application example of the present application;

[0040] Figure 7 Schematic diagram of the SMLP structure of the application example of the present application;

[0041] Figure 8 Schematic diagram of the feature separation network structure of the application example of the present application;

[0042] Figure 9 Schematic diagram of the BottleNeck structure in ResNet of the application example of the present application;

[0043] Figure 10 Schematic diagram of the feature refinement network structure of the application example of the present application;

[0044] Figure 11 Schematic diagram of the model training device structure according to the embodiment of the present application;

[0045] Figure 12 Schematic diagram of the image reconstruction device structure according to the embodiment of the present application;

[0046] Figure 13 Schematic diagram of the first device structure according to the embodiment of the present application;

[0047] Figure 14 Schematic diagram of the second device structure according to the embodiment of the present application;

[0048] Figure 15 Schematic diagram of the model training and image reconstruction system structure according to the embodiment of the present application. Detailed implementation manners

[0049] The present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0050] Based on the related art, the following two directions of solutions can be considered to implement the image reconstruction of small target objects in low-light scenes:

[0051] Solution 1, a solution based on bionics and mathematical models;

[0052] Solution 2, a solution based on neural networks.

[0053] Among them, for the specific implementation of Solution 1, first, considering that the human visual system has a constant perception ability for the colors of objects under different lighting conditions. Second, considering that the perception of object colors by the human visual system not only depends on the spectral information reflected by the object, but also is related to the observed lighting conditions. Therefore, it is possible to simulate the human brain, decompose the visual field into a series of filters of different scales, process the image at different scales, and extract the reflection component and the illumination component. Based on the above theory, the original image I(x, y) can be decomposed into the product of the reflection image R(x, y) and the illumination image L(x, y): I = R * L. After that, the separated R and L can be respectively subjected to image enhancement and illumination adjustment to enhance the local contrast and color gradient of the image. This solution can be specifically implemented by algorithms such as the multi-scale Retina response algorithm (which can be expressed in English as the Retinex Algorithm, or also called the Retinex algorithm), the single-scale Retinex algorithm, etc.; however, these algorithms generally have defects such as weak sharpening ability and poor color restoration effect.

[0054] For the specific implementation of Solution 2, usually, a large number of low-light and high-light image pairs of data can be input to the model, so that the model can learn the feature representations of the illumination and reflection images from a large amount of data, and separate and perform subsequent processing on them. The commonly used neural network model methods can include the following two categories:

[0055] 1) Based on the encoder-decoder structure, the original image pair of data is encoded and then decoded and separated, and then combined with downstream tasks to perform the final image reconstruction task;

[0056] 2) Based on the generative adversarial network method, that is, using a generator to generate a reconstructed image, and then using a discriminator to evaluate the difference between the generated image and the real image, and reducing the difference between the generated image and the real image through iteration; in this way, through adversarial training, the network can generate more real and natural reconstructed images.

[0057] However, the above commonly used neural network algorithms also face problems such as a single processing stage and relatively limited image reconstruction effects for small target objects in low-light scenarios.

[0058] Combined with the above description, it can be seen that in the related technologies, the following problems may exist in the process of realizing the image reconstruction of small target objects in low-light scenarios:

[0059] Problem 1. Although it is possible to use a neural network model based on highlight and low-light image pairs of data to achieve image enhancement under low-light conditions, when the low-light image restoration method of related technologies processes images containing small objects, since small objects are difficult to identify compared to the background, it usually leads to loss of details or blurring. That is, the colors of the restored images of small objects usually appear inaccurate or unnatural, affecting the authenticity and quality of the images and resulting in insufficient reconstruction effects for small-object images.

[0060] Problem 2. Although it is possible to implement a low-light image network model that extracts and fuses local and global features, small objects often have high contrast and large brightness differences. Without corresponding feature processing, it is usually difficult to solve the processing differences, and the reconstruction effect of the model will also be limited to a certain extent. In addition, when processing small-object images, it usually leads to problems such as over-enhancement or uneven brightness, making the image look unbalanced or overly prominent. Moreover, if the small objects in the image are in an unevenly lit background, traditional methods may not be able to accurately separate the reflection component and illumination component of the small objects, resulting in the enhancement of small objects being affected by the background and the reconstruction result being unsatisfactory. Generally speaking, there are still certain deficiencies in the illumination processing of reconstructed images under low-light conditions.

[0061] Problem 3. Classical image reconstruction methods based on mathematical optimization usually involve the selection of parameters, such as filter size, relevant thresholds, etc. The selection of these parameters is usually based on experience or a trial-and-error approach, which needs to be adjusted for different images and application scenarios, is not easy to generalize to different datasets and tasks, and requires multiple image filtering and operations, which will lead to a high computational complexity. In addition, in some specific scenarios, such as in the scenario of processing a large number of images, the time and memory overheads of these algorithms will be large. Moreover, these algorithms often produce unsatisfactory results when processing different types of images, images with different light intensities and light distributions, and the generalization scenario is weak. Generally speaking, classical model-based methods have problems such as high algorithm complexity, a large parameter space, and limited generalization scenarios.

[0062] Problem 4. Existing related algorithms usually only focus on image reconstruction or single-stage scene processing, and do not have a complete process paradigm for multi-stage image reconstruction. They may not involve image denoising, super-resolution and other processing processes at all. This makes it still a challenge to design a complete process paradigm for multi-stage image reconstruction, which requires comprehensively considering the mutual influence and correlation between different tasks and selecting appropriate algorithms and technologies to implement the processing of each stage to achieve a higher level of image reconstruction effect.

[0063] In summary, regarding how to achieve image reconstruction of small target objects under low-light scenarios, related technologies have not yet had an effective solution.

[0064] Based on this, in various embodiments of the present application, a first model for image reconstruction of an input image is trained using a specific dataset. Each sample in the dataset includes an image pair (i.e., two images) for a specific target object and corresponding to different light intensities. The first model includes a network for denoising and super-resolution processing, a network for feature separation associated with the reflection component and illumination component of the image, and a network for obtaining a reconstructed image through further convolution. In this way, various processing stages such as denoising, super-resolution, feature separation, and reconstruction of the image and / or feature data can be effectively integrated, so as to obtain high-quality image reconstruction results, and further improve the quality and usability of the reconstructed image in some specific scenarios where the image reconstruction effect in the related art is limited. For example, the image reconstruction of small target objects in low-light scenarios can be effectively realized.

[0065] It should be noted that, in various embodiments of the present application, "one or more / a plurality of" means at least one / at least one item, and "a plurality of" means at least two / at least two items.

[0066] An embodiment of the present application provides a model training method, as Figure 1 shown, the method includes:

[0067] Step 101: Determine a first training dataset. Each sample in the first training dataset includes a first image and a second image for a target object. The light intensity corresponding to the first image is less than the light intensity corresponding to the second image. The first image is an image obtained by processing the second image through a first image processing process, and the first image processing process is used to reduce the light intensity corresponding to the image.

[0068] Step 102: Use the first training dataset to train a first model for image reconstruction of an input image. The first model includes a first network, a second network, and a third network. The first network is used to perform denoising and super-resolution processing on the input image to obtain first feature data. The second network is used to obtain second feature data and third feature data based on the first feature data. The second feature data is associated with the reflection component of the input image, and the third feature data is associated with the illumination component of the input image. The third network is used to splice the second feature data and the third feature data to obtain fourth feature data, and obtain a reconstructed image by performing convolution processing on the fourth feature data.

[0069] In practical applications, in order to enable the first model to effectively implement image reconstruction of small target objects in low-light scenarios, for each sample in the first training dataset, the target object may include small target objects, such as mice, moths, flowers, etc. of a specific size. The specific types and size ranges of the small target objects can be set according to needs (such as image reconstruction requirements, etc.), and the embodiments of the present application do not limit this. Additionally, it can be understood that the illumination intensity corresponding to the first image is less than the illumination intensity corresponding to the second image, which means that the first image can be a low-light image and the second image can be a high-light image. Correspondingly, the first image processing process can be understood as a low-light processing process. Here, the images (i.e., the first image and the second image) can also be referred to as pictures, and the determination rule for the illumination level (i.e., the magnitude of the illumination intensity) of the images can also be set according to needs (such as image reconstruction requirements, etc.), and the embodiments of the present application do not limit this either.

[0070] In practical applications, the first network can also be referred to as a denoising and super-resolution joint network, etc., the second network can also be referred to as a feature separation network, etc., and the third network can also be referred to as a feature refinement network, etc. The embodiments of the present application do not limit the specific names of the first network, the second network, and the third network, as long as their functions are achieved. Additionally, it can be understood that the input image can be decomposed into the product of a reflection image R(x, y) and an illumination image L(x, y): I = R * L. The reflection image R can represent the detailed information carried by the input image, and the illumination image L can represent the illumination environment in which the target object in the input image is located. Among them, the reflection image R can also be referred to as a reflection component, a reflection part, or a reflection component, etc., and the illumination image L can also be referred to as an illumination / lighting component, an illumination / lighting part, or an illumination / lighting component, etc. In other words, the second feature data is associated with the reflection component of the input image, which means that the second feature data is associated with the reflection image R of the input image; the third feature data is associated with the illumination component of the input image, which means that the third feature data is associated with the illumination image L of the input image.

[0071] In practical applications, in order to improve data processing efficiency and simultaneously implement denoising processing and super-resolution processing, the first network can be implemented using a parallel dual-branch network architecture.

[0072] Based on this, in one embodiment, the first network may include a parallel first branch network and second branch network, and may further include a decoder network; the first branch network is used to denoise the input image to obtain fifth feature data; the second branch network is used to perform super-resolution processing on the input image to obtain sixth feature data; the decoder network is used to decode seventh feature data to obtain the first feature data; the seventh feature data is the feature data obtained by adding the fifth feature data and the sixth feature data.

[0073] Among them, in practical applications, the first branch network may also be referred to as a denoising network, etc., and the second branch network may also be referred to as a super-resolution network, etc. The embodiments of the present application do not limit the specific names of the first branch network and the second branch network, as long as their functions are realized.

[0074] In practical applications, both the first branch network and the second branch network can perform feature extraction on the input image in multiple dimensions, so as to fully extract the features of the input image and further improve the quality and usability of the reconstructed image.

[0075] Based on this, in one embodiment, the first network may further include a preprocessing component, which is used to segment the input image to obtain multiple patches, and input the multiple patches into the first branch network and the second branch network;

[0076] Both the first branch network and the second branch network may include multiple MLP-Mixers, and the number of MLP-Mixers is the same as the number of patches; each MLP-Mixer may include a token-mixing component and a channel-mixing component, and the token-mixing component and the channel-mixing component are respectively used to perform feature extraction in different dimensions of the corresponding patch, and the token-mixing component includes an SMLP;

[0077] The decoder network is a CNN constructed based on transposed convolution technology and the PReLU activation function.

[0078] In actual application, the token-mixing component and the channel-mixing component are respectively used for feature extraction in different dimensions of the corresponding patches, which means that through the token-mixing component and the channel-mixing component, horizontal and vertical feature extraction for the multiple patches can be achieved. In addition, since the token-mixing component includes SMLP, and the decoder network is a CNN constructed based on the transposed convolution technique and the PReLU activation function, the first network can fully extract the features of the input image on the premise of effectively reducing the number of network parameters.

[0079] In actual application, the second network can be implemented based on ResNet and DenseNet. And, in order to further improve the data processing efficiency and effectively separate the reflection feature (i.e., the second feature data) from the illumination feature (i.e., the third feature data), the second network can also include a parallel dual-branch network architecture.

[0080] Based on this, in one embodiment, the second network can be a DenseNet constructed based on the Bottleneck structure in ResNet and using the strided convolution method. The second network includes a parallel third branch network and a fourth branch network; the third branch network is used to encode the reflection component of the input image based on the first feature data to obtain the second feature data; the fourth branch network is used to encode the illumination component of the input image based on the first feature data to obtain the third feature data.

[0081] Among them, in actual application, the third branch network can also be called a reflection refinement network, etc., and the fourth branch network can also be called an illumination adjustment network, etc. The embodiments of the present application do not limit the specific names of the third branch network and the fourth branch network, as long as their functions are realized. In addition, the specific structures of the second network, the third branch network, and the fourth branch network can be set according to needs (such as image reconstruction requirements, etc.), and the embodiments of the present application also do not limit this.

[0082] In practical applications, it can be understood that different networks in the first model (i.e., the first network, the second network, and the third network) can correspond to different loss functions. By weighting these loss functions, the overall loss function of the first model (which can be denoted as the first loss function in subsequent descriptions) can be obtained. By minimizing the first loss function, the first model can be solved and evaluated, that is, the training and optimization of the first model are completed. Among them, when training the first model, the DWA algorithm can be used to adjust the weights of the loss functions of each network during each round of training, so as to update the first loss function, and thus the training effect of the first model can be further improved, that is, the quality and usability of the reconstructed image can be further improved.

[0083] Based on this, in one embodiment, the first loss function corresponding to the first model can be a function obtained by weighting the second loss function corresponding to the first network, the third loss function corresponding to the second network, the fourth loss function corresponding to the second network, and the fifth loss function corresponding to the third network. The third loss function is associated with the reflection component of the input image, and the fourth loss function is associated with the illumination component of the input image;

[0084] Correspondingly, training the first model may include:

[0085] During each round of training, use the DWA algorithm to adjust one or more of the weights corresponding to the second loss function, the weight corresponding to the third loss function, the weight corresponding to the fourth loss function, and the weight corresponding to the fifth loss function to obtain an updated first loss function;

[0086] Use the updated first loss function to adjust the model parameters of the first model.

[0087] Among them, in practical applications, the specific formulas corresponding to the first loss function, the second loss function, the third loss function, the fourth loss function, and the fifth loss function can be set according to needs (such as image reconstruction requirements, etc.), and the embodiments of the present application do not limit this.

[0088] In practical applications, from the perspective of hardware implementation, the model training method provided by the embodiments of the present application can be implemented by a first device, and the first device can include a server, etc. And after the first model is trained, the first model can be deployed to a second device, and the second device uses the first model to implement image reconstruction, that is, the second device schedules and optimizes the first model according to needs.

[0089] Among them, the second device may be the same as or different from the first device. Exemplarily, the second device may include a server, an imaging device (such as a surveillance camera, etc.). Here, with the wide popularization of imaging devices, imaging devices are applied to various scenarios in real life. For example, in the fields of smart urban management, transparent kitchens and bright kitchens, and general security, visual analysis and early warning are carried out on target objects in the scenarios, including low-light scenarios such as at night and in basements. Under low-light conditions, the imaging devices in the related art will produce situations such as high noise and low resolution for the imaging of objects. In particular, the imaging of small target objects such as mice and moths will be greatly distorted, restricting their application in low-light situations. By deploying the first model to the imaging device, it is possible to effectively implement image reconstruction of small target objects in low-light scenarios, thereby significantly improving the performance of the imaging device for video surveillance under low-light conditions.

[0090] Correspondingly, an embodiment of the present application further provides an image reconstruction method, which is applied to a second device (such as a server, an imaging device, etc.), as Figure 2 shown. The method includes:

[0091] Step 201: Obtain an image to be processed;

[0092] Step 202: Use the first model to perform image reconstruction on the image to be processed to obtain a reconstructed image. In other words, input the image to be processed into the first model to obtain the reconstructed image output by the first model. The first model is a model trained by using the model training method provided by one or more of the above technical solutions.

[0093] The model training method provided by the embodiments of the present application determines a first training data set. Each sample in the first training data set includes a first image and a second image of a target object. The illumination intensity corresponding to the first image is less than that corresponding to the second image. The first image is an image obtained by processing the second image through a first image processing process, and the first image processing process is used to reduce the illumination intensity of the image. Using the first training data set, a first model is trained. The first model is used to perform image reconstruction on the input image. Among them, the first model includes a first network, a second network, and a third network. The first network is used to perform denoising processing and super-resolution processing on the input image to obtain first feature data. The second network is used to obtain second feature data and third feature data based on the first feature data. The second feature data is associated with the reflection component of the input image, and the third feature data is associated with the illumination component of the input image. The third network is used to splice the second feature data and the third feature data to obtain fourth feature data, and through convolutional processing on the fourth feature data, a reconstructed image is obtained. The solution provided by the embodiments of the present application uses a specific data set (i.e., the first training data set) to train a first model for performing image reconstruction on the input image. Each sample in this data set includes an image pair (i.e., two images) for a specific target object and corresponding to different illumination intensities. The first model includes a network for performing denoising processing and super-resolution processing (i.e., the first network), a network for performing feature separation associated with the reflection component and illumination component of the image (i.e., the second network), and a network for obtaining a reconstructed image through further convolution (i.e., the third network). In this way, it is possible to effectively integrate various processing stages such as denoising, super-resolution, feature separation, and reconstruction of the image and / or feature data, so as to obtain high-quality image reconstruction results, and further improve the quality and usability of the reconstructed image in some specific scenarios where the image reconstruction effect of the related technology is limited. For example, it can effectively realize the image reconstruction of small target objects in low-light scenarios.

[0094] The following further describes the present application in detail with application examples.

[0095] This application example fully considers the application scenario and data characteristics of small target objects under low-light conditions. Aiming at the innovation of existing related algorithms focusing on one or more specific stages of image reconstruction and the problems such as the need for complex parameter combination searches for optimal parameter matching, a new neural network model that organically combines image super-resolution, image denoising, and low-light image reconstruction (i.e., the above-mentioned first model, which can also be understood as a new algorithm) is proposed, and the adaptive parameter method is used to solve the drawback of the complex search for the optimal parameter combination. The overall structure of this neural network model is as Figure 3As shown, it includes a denoising and super-resolution joint network (i.e., the first network mentioned above), a feature separation network (i.e., the second network mentioned above), and a feature refinement network (i.e., the third network mentioned above); among them, the denoising and super-resolution joint network includes a denoising network (i.e., the first branch network mentioned above) and a super-resolution network (i.e., the second branch network mentioned above); the feature separation network includes a reflection refinement network (i.e., the third branch network mentioned above) and an illumination adjustment network (i.e., the fourth branch network mentioned above). In other words, in the subsequent description of this application example, the first network mentioned above will be referred to as the denoising and super-resolution joint network, the second network mentioned above will be referred to as the feature separation network, the third network mentioned above will be referred to as the feature refinement network, the first branch network mentioned above will be referred to as the denoising network, the second branch network mentioned above will be referred to as the super-resolution network, the third branch network mentioned above will be referred to as the reflection refinement network, and the fourth branch network mentioned above will be referred to as the illumination adjustment network.

[0096] Based on the above structure, first, image pair data of low-light and high-light small target objects can be constructed (i.e., a low-light image and a high-light image for the same small target object, that is, the first image and the second image mentioned above), and sent into the above neural network model. Through the denoising and super-resolution process in the first stage, the original small target object is enlarged and super-resolution reconstructed (which can also be understood as super-resolution reconstruction) to obtain a preliminary super-resolution image; at the same time, the network can remove noise from the feature data and perform condensation coding of the feature data to realize the function of the encoder. After that, the encoded mixed data can be fully decoded, and then through the feature separation process in the second stage, the illumination part and the reflection part are encoded in the form of a dual-branch network architecture to learn the feature representation method for separating the original data. Finally, the two separated parts of data can be spliced and input into an encoder-decoder network architecture (i.e., the feature refinement network) for decoding and restoration to complete the feature refinement process in the third stage. Here, in order to streamline the number of network parameters, the original input low-light and high-light data pairs can share the same network parameters in the super-resolution network, denoising network, and feature separation network parts.

[0097] In this application example, the structure of the denoising and super-resolution joint network is as Figure 4 shown. The following will combine Figure 4 to describe the denoising and super-resolution joint network in detail.

[0098] In this application example, since small target objects are usually small in size relative to the background and thus contain less pixel information, magnifying the image and enhancing the pixels of small target objects will contribute to subsequent image reconstruction. Here, the commonly used denoising and super-resolution networks in related technologies are usually implemented based on the classic CNN network architecture. However, the classic CNN network may face disadvantages such as insufficient spatial information mining and limited global receptive fields. Considering the characteristics of the original image being taken in a low-light scene, especially the large impact of background information on the imaging of small target objects, for the reconstruction task of small target objects, the global information of the image is very important, which is conducive to separating small target objects from background information and reconstructing the detailed information of small target objects. Therefore, this application example uses SMLP as part of the decoder network.

[0099] Specifically, as Figure 4 shown, first, the original image can be upsampled (which can be expressed in English as Upsampling), which will facilitate the subsequent patch partition of the image (which can be expressed in English as Patch Partition, equivalent to the function of the above-mentioned preprocessing component). After that, drawing on the ideas in the field of natural language processing, the original image can be segmented into multiple patches and linearly embedded (which can be expressed in English as Linear Embedding), and SMLP can be used to extract information from each patch. During the information extraction process, the idea of MLP-Mixer can be adopted, and the patch can be divided into two parts: sequence mixing (Token Mixing, that is, the above-mentioned token-mixing component) and channel mixing (Channel Mixing, that is, the above-mentioned channel-mixing component); in other words, each part can use Token Mixing and Channel Mixing to extract horizontal and vertical information from the segmented patches to obtain the global information of the feature data and the relevant information between channels.

[0100] Among them, as Figure 5 shown, Token Mixing can mainly be composed of SMLP in cooperation with DWConv and the Bottleneck structure, which is used to extract information at the same position of different patches, so as to learn the global information of the picture like Tokens in the NLP field. And as Figure 6 shown, Channel Mixing is to extract information between different channels, which is used to extract the correlation between the picture itself and different channels. As Figure 7As shown, it can be understood that the key component SMLP in this application example consists of a double-bracket structure. By utilizing the linear complexity of the MLP algorithm and the method of image size reconstruction, the number of network parameters can be significantly reduced. Compared with the Vision in Transformer structure, this architecture can, on the one hand, reduce the number of network parameters, and on the other hand, retain its advantage of global information extraction.

[0101] In addition, as Figure 4 shown, this application example proposes the idea of using a high-light image (i.e., the above-mentioned second image) as a guiding map, and putting the data pair containing the low-light image (i.e., the above-mentioned first image) and the high-light image into a joint network for learning. Since this application example uses a joint network for image processing, in order to more fully extract the information in the original data, a double-branch SMLP network architecture can be adopted to redundantly extract the information in the original data in the form of two sub-encoders respectively, and then add the feature data encoded by the two sub-encoders as the final encoded data and put it into the decoder.

[0102] To simplify the network structure, the decoder used in this application example can be a pure CNN decoding network stacked with 7-layer transposed convolution and PReLU (i.e., the above-mentioned decoder network). Since the encoder used in this application example can extract the information of the encoded data relatively fully, in order to better decode the information, this application example can adopt a variant of the ReLU activation function - the activation function PReLU that can self-learn hyperparameters to construct the decoder network to improve its flexibility and performance. Based on the above explanation,

[0103] this application example can obtain the loss function of the joint network (i.e., the above-mentioned second loss function) as:

[0104]

[0105] where SUDN l and SUDN h respectively represent the low-light feature map and the high-light feature map output by the super-resolution network.

[0106] In this application example, the structure of the feature separation network is as Figure 8 shown. The following will combine Figure 8 to describe the feature separation network in detail.

[0107] In this application example, the original image can be considered as a combination of an illumination part L and a reflection part R. Among them, the reflection part can represent the detailed information carried by the image, while the illumination part can represent the lighting environment in which the objects in the image are located. Therefore, by enhancing the reflection part of the image and adjusting the illumination part of the image, the real image data in a high-light scene can be reconstructed. Here, S-Net can be used to represent the feature separation network, and the feature separation network in this application example can be expressed as:

[0108]

[0109] As Figure 8 shown, the feature separation network is implemented using a dual-branch network architecture. By combining the Bottleneck structure in ResNet and the use of the dual-branch structure, it can effectively separate the illumination part and the reflection part information, that is, through different network branches, the feature representations of the illumination part signal and the reflection part signal can be obtained respectively from the original mixed signal; and, based on the idea of DenseNet, in order to better purify these two parts of the signal, the method of adding strided feature maps can be used to make full use of the illumination and reflection information retained in the mixed signal, so as to make a more sufficient light adjustment to the illumination part and a more sufficient restoration of the detailed information of the reflection part. Here, the Bottleneck structure in ResNet can be as Figure 9 shown, including two 1*1 convolutions and one 3*3 convolution.

[0110] Among them, since the dual-branch network needs to enhance the image of the reflection part and adjust the illumination of the illumination part, the functions of the two network branches are different. Therefore, in order to better enable the branch network to learn more information, the corresponding high-light image information is required as a guide. Here, L l 、L h 、R l 、R h are used to represent the illumination part of the low light, the illumination part of the high light, the reflection part of the low light, and the reflection part of the high light respectively. Using the 1-norm as the loss function, the illumination loss function (i.e., the above fourth loss function) and the reflection loss function (i.e., the above third loss function) are as follows:

[0111]

[0112] In this application example, the structure of the feature refinement network is as Figure 10 shown. The following will describe the feature refinement network in detail in combination with Figure 10 .

[0113] Specifically, based on the bionic hypothesis that an image is divided into an illumination part and a reflection part, this application example can splice the illumination adjustment feature data (i.e., the above-mentioned third feature data) and the reflection reconstruction feature data (i.e., the above-mentioned second feature data) output by the dual branches of the feature separation network, and then send them into the feature refinement network for subsequent processing. As Figure 10 shown, the feature refinement network can first encode the input spliced data (i.e., the above-mentioned fourth feature data) from the feature format of 512*512*6 into feature data of 32*32*128 using four layers of lightweight convolution, and then decode it into the finally reconstructed image data of 512*512*3 using four layers of convolution. In order to enable the feature refinement network to fully learn the feature information of the image, this application example can finally use the real data of 512*512*3 as a guide to construct a loss function (i.e., the above-mentioned fifth loss function) as follows:

[0114]

[0115] where RE represents the reconstructed image and GT represents the real data.

[0116] The adaptive loss function weight balancing strategy of this application example is described below.

[0117] Since this application example combines the image processing processes of multiple stages (i.e., the above-mentioned first stage, second stage, and third stage), the constructed low-light small target image reconstruction system (i.e., the above-mentioned first model) has more loss functions. How to balance the weights of these loss functions reflects the importance of each stage in the entire system. In order to avoid the disadvantages of complex parameter space search, this application example proposes an adaptive loss function weight balancing strategy. The final loss function of this application example (i.e., the above-mentioned first loss function) can be expressed as:

[0118]

[0119] where α, β, μ, and ω respectively represent the weights of the loss functions of each corresponding stage. Given that the loss functions adopted in this application example are all 1-norms, the properties and smoothness of the loss functions are relatively consistent. Therefore, this application example can adopt the DWA algorithm as the adjustment method for the weights of the adaptive loss function. Specifically, i can be used to represent each processing stage, t can be used to represent the training time, represents the weights of the loss functions of each stage, and represents the losses of each stage. Then, as shown in Formulas 7-9, the weight can be dynamically updated in different rounds using the adjustment coefficient λ i :

[0120]

[0121] In this application example, the ratio of the loss of the current task to the loss of the previous time can be used as the adjustment coefficient λ. If the error gap between the two times is large, the adjustment coefficient will also be larger. This indicates that the loss of the current task decreases slowly, which also means that the training of this task is more difficult and the model should pay more attention. The use of the softmax function is to amplify and normalize each loss function, which is beneficial to the adaptive adjustment of the model. Through the adaptive weight balance adjustment strategy, the search for complex weight parameters can be avoided, and the automation level of the system can be greatly improved.

[0122] To sum up, this application example proposes a new multi-stage fusion method and training strategy for reconstructing low-light small target objects in response to the challenges faced by the related technology (i.e., problems 1 to 4). First, by constructing a method for image data pairs of low-light and high-light small target objects, the model is trained to learn the feature representation method of the illumination and reflection of small target images, so as to effectively separate the illumination and reflection parts of the image. Secondly, a multi-stage dual-branch network architecture is designed to organically integrate the image reconstruction process, the image denoising process, and the image super-resolution process, so as to obtain high-quality reconstruction results. Finally, the loss functions of each stage are adaptively balanced in order to balance the positions occupied by the processing processes of each stage in the network, so as to obtain the best reconstruction results, and a high model generalization ability can be achieved without selecting complex parameter combinations. By organically integrating the multi-stage image processing process, the introduction of the dual-branch network architecture, and the adaptive balancing of the loss functions of each stage, this application example can effectively realize the image reconstruction of small target objects under low-light conditions and is superior to the existing algorithms of the related technology; at the same time, by adopting a new network architecture combining SMLP and bionic ideas and an adaptive network training strategy, this application example can greatly improve the reconstruction effect of the network on the original signal and reduce the overall training time, which helps to fully strip the signals in the mixed signals and efficiently and robustly reconstruct the original noisy small object image data.

[0123] Specifically, the solution provided by this application example has the following advantages:

[0124] 1) This application example proposes a new multi-stage network architecture combining CNN and SMLP for reconstructing low-light small target objects, which combines super-resolution, denoising, illumination adjustment, reflection refinement, and image feature reconstruction; solving the problem of small object image reconstruction under low-light conditions in a multi-stage manner is beneficial to the performance improvement of imaging devices under low-light conditions;

[0125] 2) This application example proposes a new bionics-based low-light feature representation network. Based on the bionics information of the human eye visual system, this application example divides low-light imaging pictures into a reflection part and an illumination part. By combining the network architectures of ResNet and DenseNet, a dual-branch network is used to adjust the illumination and refine the features (i.e., reflectance refinement) of the feature data in parallel. Finally, high-quality data of the original imaging under high-light conditions can be obtained;

[0126] 3) This application example proposes a new dual-branch SMLP neural network for joint denoising and super-resolution. First, a strategy of patch segmentation of the original image is adopted. The segmented patches are used for horizontal and vertical information extraction by Token Mixing and Channel Mixing constructed by SMLP for subsequent encoding. Then, the encoded feature data are added and put into a decoder constructed by transposed convolution and PReLU activation function for full decoding, and sent to the subsequent illumination and reflectance refinement parts for feature learning. Through the processing of the super-resolution and denoising joint network, small target objects can be fully separated and reconstructed from the background information for the reconstruction of subsequent high-light images;

[0127] 4) This application example proposes a new strategy for multi-stage low-light small target object image reconstruction. First, the method of using pairs of low-light and high-light image data is adopted to share the training process of the early super-resolution, denoising, and signal separation stages, effectively reducing the number of network parameters and improving the reconstruction effect of the network. Secondly, real images are used to make up for the lack of details in low-light images, which more efficiently helps the multi-stage system to reconstruct small target object images. Finally, an adaptive loss function weight balancing strategy is adopted to avoid the complicated search in the parameter space and improve the training effect of the network.

[0128] To implement the model training method of the embodiments of this application, the embodiments of this application also provide a model training device, which is set on the first device, as Figure 11 shown. The device includes:

[0129] A first processing unit 1101, configured to determine a first training data set. Each sample in the first training data set includes a first image and a second image of a target object. The illumination intensity corresponding to the first image is less than the illumination intensity corresponding to the second image. The first image is an image obtained by processing the second image through a first image processing process, and the first image processing process is used to reduce the illumination intensity of the image;

[0130] A second processing unit 1102 is configured to train a first model by using the first training data set. The first model is used to perform image reconstruction on an input image. The first model includes a first network, a second network, and a third network. The first network is configured to perform denoising processing and super-resolution processing on the input image to obtain first feature data. The second network is configured to obtain second feature data and third feature data based on the first feature data. The second feature data is associated with a reflection component of the input image, and the third feature data is associated with an illumination component of the input image. The third network is configured to splice the second feature data and the third feature data to obtain fourth feature data, and obtain a reconstructed image by performing convolution processing on the fourth feature data.

[0131] Wherein, in one embodiment, a first loss function corresponding to the first model is a function obtained by weighting a second loss function corresponding to the first network, a third loss function corresponding to the second network, a fourth loss function corresponding to the second network, and a fifth loss function corresponding to the third network. The third loss function is associated with a reflection component of the input image, and the fourth loss function is associated with an illumination component of the input image.

[0132] Accordingly, the second processing unit 1102 is further configured to:

[0133] In the training process of each round, use the DWA algorithm to adjust one or more of the weights corresponding to the second loss function, the weights corresponding to the third loss function, the weights corresponding to the fourth loss function, and the weights corresponding to the fifth loss function to obtain an updated first loss function.

[0134] Use the updated first loss function to adjust the model parameters of the first model.

[0135] In practical applications, the first processing unit 1101 and the second processing unit 1102 may be implemented by a processor in a model training device.

[0136] It should be noted that when the model training device provided in the above embodiment performs model training, only the above division of each program module is used for illustration. In practical applications, the above processing may be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the model training device provided in the above embodiment and the model training method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0137] To implement the image reconstruction method of the embodiments of the present application, an image reconstruction device is further provided in the embodiments of the present application. The device is disposed on a second device, such as Figure 12 shown. The device includes:

[0138] An acquisition unit 1201, configured to acquire an image to be processed;

[0139] A reconstruction unit 1202, configured to perform image reconstruction on the image to be processed by using a first model to obtain a reconstructed image, where the first model is a model trained by using the model training method provided in the above one or more technical solutions.

[0140] In practical applications, the acquisition unit 1201 and the reconstruction unit 1202 may be implemented by a processor in the image reconstruction device.

[0141] It should be noted that: when the image reconstruction device provided in the above embodiment performs image reconstruction, only the division of the above program modules is used for illustration. In practical applications, the above processing may be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the image reconstruction device provided in the above embodiment and the image reconstruction method embodiment belong to the same concept. For the specific implementation process, refer to the method embodiment, which will not be elaborated here.

[0142] Based on the hardware implementation of the above program modules, and to implement the model training method of the embodiments of the present application, a first device is further provided in the embodiments of the present application, such as Figure 13 shown. The first device 1300 includes:

[0143] A first communication interface 1301, capable of performing information interaction with other electronic devices (such as a second device, etc.);

[0144] A first processor 1302, connected to the first communication interface 1301 to implement information interaction with other electronic devices, and configured to execute the model training method provided in the above one or more technical solutions when running a computer program;

[0145] A first memory 1303, where the computer program is stored on the first memory 1303.

[0146] Specifically, the first processor 1302 is configured to:

[0147] Determine a first training dataset, where each sample in the first training dataset includes a first image and a second image of a target object, the illumination intensity corresponding to the first image is less than the illumination intensity corresponding to the second image, the first image is an image obtained by processing the second image through a first image processing process, and the first image processing process is used to reduce the illumination intensity corresponding to the image;

[0148] Use the first training dataset to train a first model, where the first model is used to perform image reconstruction on an input image; wherein, the first model includes a first network, a second network, and a third network, the first network is used to perform denoising processing and super-resolution processing on the input image to obtain first feature data; the second network is used to obtain second feature data and third feature data based on the first feature data, the second feature data is associated with the reflection component of the input image, and the third feature data is associated with the illumination component of the input image; the third network is used to splice the second feature data and the third feature data to obtain fourth feature data, and obtain a reconstructed image by performing convolution processing on the fourth feature data.

[0149] Wherein, in one embodiment, the first loss function corresponding to the first model is a function obtained by weighting the second loss function corresponding to the first network, the third loss function corresponding to the second network, the fourth loss function corresponding to the second network, and the fifth loss function corresponding to the third network, the third loss function is associated with the reflection component of the input image, and the fourth loss function is associated with the illumination component of the input image;

[0150] Correspondingly, the first processor 1302 is further configured to:

[0151] During the training process of each round, use the DWA algorithm to adjust one or more of the weights corresponding to the second loss function, the weights corresponding to the third loss function, the weights corresponding to the fourth loss function, and the weights corresponding to the fifth loss function to obtain an updated first loss function;

[0152] Use the updated first loss function to adjust the model parameters of the first model.

[0153] It should be noted that: the specific processing process of the first processor 1302 can be understood with reference to the above method and will not be elaborated here.

[0154] Of course, in practical applications, the various components in the first device 1300 are coupled together through the bus system 1304. It can be understood that the bus system 1304 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 1304 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 13 all kinds of buses are labeled as the bus system 1304.

[0155] The first memory 1303 in the embodiment of the present application is used to store various types of data to support the operation of the first device 1300. Examples of these data include: any computer program for operating on the first device 1300.

[0156] The method disclosed in the embodiment of the present application described above can be applied to or implemented by the first processor 1302. The first processor 1302 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the first processor 1302 or instructions in software form. The first processor 1302 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The first processor 1302 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiment of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the first memory 1303. The first processor 1302 reads the information in the first memory 1303 and combines its hardware to complete the steps of the foregoing method.

[0157] In an exemplary embodiment, the first device 1300 may be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components, and is used to execute the foregoing method.

[0158] Based on the hardware implementation of the foregoing program modules, and in order to implement the image reconstruction method of the embodiments of the present application, the embodiments of the present application further provide a second device, as Figure 14 shown. The second device 1400 includes:

[0159] A second communication interface 1401, capable of interacting with other electronic devices (such as the first device, etc.);

[0160] A second processor 1402, connected to the second communication interface 1401 to implement information interaction with other electronic devices, and is used to execute the image reconstruction method provided by the above one or more technical solutions when running a computer program;

[0161] A second memory 1403, on which the computer program is stored.

[0162] Specifically, the second processor 1402 is used for:

[0163] Obtain an image to be processed;

[0164] Use a first model to perform image reconstruction on the image to be processed to obtain a reconstructed image, where the first model is a model trained by using the model training method provided by the above one or more technical solutions.

[0165] It should be noted that the specific processing process of the second processor 1402 may be understood with reference to the foregoing method, and will not be elaborated here.

[0166] Of course, in practical applications, the various components in the second device 1400 are coupled together through the bus system 1404. It can be understood that the bus system 1404 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 14 all kinds of buses are labeled as the bus system 1404.

[0167] The second memory 1403 in the embodiment of the present application is used to store various types of data to support the operation of the second device 1400. Examples of such data include: any computer program for operating on the second device 1400.

[0168] The method disclosed in the above embodiment of the present application can be applied to or implemented by the second processor 1402. The second processor 1402 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the second processor 1402 or the instructions in software form. The second processor 1402 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The second processor 1402 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiment of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the second memory 1403. The second processor 1402 reads the information in the second memory 1403 and combines its hardware to complete the steps of the foregoing method.

[0169] In an exemplary embodiment, the second device 1400 may be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components for executing the foregoing method.

[0170] It can be understood that the memories (the first memory 1303 and the second memory 1403) in the embodiments of the present application can be volatile memories or non-volatile memories, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.

[0171] To implement the method provided by the embodiments of the present application, the embodiments of the present application further provide a model training and image reconstruction system, as Figure 15 shown. The system includes: a first device 1501 and a second device 1502.

[0172] Here, it should be noted that: the specific processing procedures of the first device 1501 and the second device 1502 have been described in detail above and will not be elaborated here.

[0173] In an exemplary embodiment, the embodiments of the present application further provide a storage medium, namely a computer storage medium, specifically a computer-readable storage medium. For example, it includes a first memory 1303 storing a computer program, and the above computer program can be executed by a first processor 1302 of the first device 1300 to complete the steps of any of the foregoing model training methods. Another example is a second memory 1403 storing a computer program, and the above computer program can be executed by a second processor 1402 of the second device 1400 to complete the steps of any of the foregoing image reconstruction methods. The computer-readable storage medium can be a FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.

[0174] In an exemplary embodiment, the embodiments of the present application further provide a computer program product, including a computer program. The computer program can be executed by a first processor 1302 of the first device 1300 to complete the steps of any of the foregoing model training methods; or the computer program can be executed by a second processor 1402 of the second device 1400 to complete the steps of any of the foregoing image reconstruction methods.

[0175] It should be noted that: "first", "second", etc. are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0176] In addition, among the technical solutions described in the embodiments of the present application, they can be combined arbitrarily without conflict.

[0177] The above is only a preferred embodiment of the present application and is not used to limit the protection scope of the present application.

Claims

1. A model training method, characterized in that: include: Determine a first training data set, each sample in the first training data set includes a first image and a second image of a target object, the light intensity corresponding to the first image is less than the light intensity corresponding to the second image, the first image is an image obtained by processing the second image through a first image processing process, and the first image processing process is used to reduce the light intensity corresponding to the image; The first training data set is used to train a first model, and the first model is used to reconstruct an input image; wherein the first model includes a first network, a second network, and a third network, and the first network is used to perform denoising and super-resolution processing on the input image to obtain first feature data; the second network is used to obtain second feature data and third feature data based on the first feature data, and the second feature data is associated with the reflection component of the input image, and the third feature data is associated with the illumination component of the input image; the third network is used to splice the second feature data and the third feature data to obtain fourth feature data, and obtain a reconstructed image by performing convolution processing on the fourth feature data.

2. The method according to claim 1, characterized in that: The first network includes a first branch network and a second branch network in parallel, and also includes a decoder network; the first branch network is used to perform denoising on the input image to obtain fifth feature data; the second branch network is used to perform super-resolution processing on the input image to obtain sixth feature data; the decoder network is used to decode the seventh feature data to obtain the first feature data; the seventh feature data is feature data obtained by adding the fifth feature data and the sixth feature data.

3. The method according to claim 2, characterized in that The first network further includes a preprocessing component, which is used to segment the input image to obtain a plurality of patches, and input the plurality of patches into the first branch network and the second branch network; The first branch network and the second branch network both include a plurality of multi-layer perceptron mixers MLP-Mixer, the number of the MLP-Mixers being the same as the number of the patches; each MLP-Mixer includes a sequence mixing token-mixing component and a channel mixing channel-mixing component, the token-mixing component and the channel-mixing component are respectively used to extract features of different dimensions of the corresponding patches, and the token-mixing component includes a sparse multi-layer perceptron SMLP; The decoder network is a convolutional neural network (CNN) constructed based on the transposed convolution technique and the parameterized rectified linear unit (PReLU) activation function.

4. The method according to claim 1, characterized in that The second network is a dense convolutional network DenseNet constructed based on the bottleneck BottleNeck structure in the residual network ResNet and using a strided convolution method, and the second network includes a parallel third branch network and a fourth branch network; The third branch network is used to encode the reflection component of the input image based on the first feature data to obtain the second feature data; the fourth branch network is used to encode the illumination component of the input image based on the first feature data to obtain the third feature data.

5. The method according to claim 1, characterized in that The first loss function corresponding to the first model is a function obtained by weighting a second loss function corresponding to the first network, a third loss function corresponding to the second network, a fourth loss function corresponding to the second network, and a fifth loss function corresponding to the third network, the third loss function is associated with a reflection component of the input image, and the fourth loss function is associated with an illumination component of the input image; The training of the first model comprises: In each round of training, a dynamic weighted average DWA algorithm is used to adjust one or more of the weights corresponding to the second loss function, the weights corresponding to the third loss function, the weights corresponding to the fourth loss function, and the weights corresponding to the fifth loss function to obtain an updated first loss function; The updated first loss function is used to adjust the model parameters of the first model.

6. An image reconstruction method, characterized in that: include: Get the image to be processed; The image to be processed is reconstructed using a first model to obtain a reconstructed image, wherein the first model is a model trained using the model training method described in any one of claims 1 to 5.

7. A model training device, characterized in that: include: A first processing unit is used to determine a first training data set, each sample in the first training data set includes a first image and a second image of a target object, the light intensity corresponding to the first image is less than the light intensity corresponding to the second image, the first image is an image obtained by processing the second image through a first image processing process, and the first image processing process is used to reduce the light intensity corresponding to the image; The second processing unit is used to train a first model using the first training data set, and the first model is used to reconstruct an input image; wherein the first model includes a first network, a second network and a third network, and the first network is used to perform denoising and super-resolution processing on the input image to obtain first feature data; the second network is used to obtain second feature data and third feature data based on the first feature data, the second feature data is associated with a reflection component of the input image, and the third feature data is associated with an illumination component of the input image; the third network is used to splice the second feature data and the third feature data to obtain fourth feature data, and obtain a reconstructed image by performing convolution processing on the fourth feature data.

8. An image reconstruction device, characterized in that: include: An acquisition unit, used for acquiring an image to be processed; A reconstruction unit is used to reconstruct the image to be processed using a first model to obtain a reconstructed image, wherein the first model is a model trained using the model training method described in any one of claims 1 to 5.

9. A first device, characterized in that: include: a first processor and a first memory for storing a computer program executable on the processor, Wherein, when the first processor is used to run the computer program, the steps of the method described in any one of claims 1 to 5 are executed.

10. A second device, characterized in that: include: a second processor and a second memory for storing a computer program executable on the processor, Wherein, when the second processor is used to run the computer program, it executes the steps of the method according to claim 6.

11. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 5, or implements the steps of the method according to claim 6.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 5, or implements the steps of the method according to claim 6.