Model training method and device, image reconstruction method and device, electronic equipment and medium

By combining segmentation network and diffusion network, using segmentation mask probability training images as prior knowledge, the problem of poor image reconstruction quality in the prior art is solved, and the recognition and reconstruction accuracy of edge blurred objects such as flames and smoke is improved.

CN120411686AActive Publication Date: 2025-08-01PENG CHENG LAB

Patent Information

Application Number
CN202510814055.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-08-01
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing image object removal technology deals with objects with blurred boundaries and translucent properties, and the image reconstruction quality is poor, especially the edge blur of objects such as flames and smoke, resulting in poor texture interference and reconstruction quality.

Method used

The segmentation network module is used for image segmentation processing, and the segmentation mask probability training image is generated as a prior knowledge. It is combined with the diffusion network module to perform noise addition and denoising processing. It is iteratively updated by calculating the total loss, which improves the occlusion area recognition ability and background structure mining, and improves the image reconstruction quality.

Benefits of technology

The processing accuracy and effect of the image reconstruction model for edge blurred objects is significantly improved, and the quality of image reconstruction is enhanced, especially in occlusion areas such as flame and smoke and retention capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411686A_ABST
    Figure CN120411686A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, an image reconstruction method and device, electronic equipment and a medium, and relates to the technical field of neural networks. The method comprises the following steps: inputting an initial training image into an initial image reconstruction model; performing image segmentation processing through a segmentation network module which completes pre-training to obtain a segmentation mask probability training image, and inputting the segmentation mask probability training image and the initial training image into a noise adding unit of a diffusion network module to obtain a noise-added training image; inputting the noise-added training image into a denoising unit of a diffusion network module to obtain a predicted training noise and a reconstructed training image; and based on the reconstructed training image, the initial training image, the predicted training noise, the segmented mask probability training image and the real label image, calculating to obtain total loss, and based on the total loss, carrying out iterative updating on the initial image reconstruction model to obtain a trained target image reconstruction model. The method can improve the reconstruction quality of an image reconstruction model for an image with an occlusion area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of neural networks, and particularly relates to a model training method, an image reconstruction method, an apparatus, an electronic device, and a medium. Background Art

[0002] Image object removal is a technology aimed at removing specific objects from an image and reconstructing the integrity of its occluded background, and is widely applied in fields such as monitoring, security, and medical image analysis. This technology provides important support for the application of computer vision systems in multiple fields by improving the integrity and usability of images.

[0003] In related technologies, with the rapid development of deep learning, generative adversarial networks (GANs) and their derivative models have been widely applied to image object removal tasks. When these networks and their derivatives process target objects with clear boundaries and distinct contours, significant research results have been achieved. However, for images of objects with blurred boundaries and semi-transparent characteristics such as flames and smoke, the quality of image reconstruction by current networks and their derivative models is relatively poor. Summary of the Invention

[0004] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides a model training method, an image reconstruction method, an apparatus, an electronic device, and a medium, which can improve the reconstruction quality of an image reconstruction model for an image with an occluded area.

[0005] To achieve the above object, a first aspect embodiment of the present application provides an image reconstruction model training method, including: Obtaining an initial training image with an occluded area and a corresponding ground truth label image of the initial training image; Inputting the initial training image into an initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a pre-trained segmentation network module; Performing image segmentation processing on the initial training image through the segmentation network module to obtain a segmentation mask probability training image; wherein, a first pixel of the segmentation mask probability training image corresponds to a second pixel of the initial training image one by one, and a value of the first pixel represents a probability that the corresponding second pixel is an occluded area; Inputting the segmentation mask probability training image and the initial training image into a noise addition unit of the diffusion network module to obtain a noise-added training image; Inputting the noise-added training image into a denoising unit of the diffusion network module to obtain a predicted training noise and a reconstructed training image; Calculate a total loss based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the true label image; Iteratively update the initial image reconstruction model based on the total loss to obtain a trained target image reconstruction model.

[0006] According to some embodiments of the first aspect of the present application, the step of inputting the segmentation mask probability training image and the initial training image into the noise addition unit of the diffusion network module to obtain a noise-added training image includes: Perform a per-pixel weighting operation on the initial training image and the segmentation mask probability training image to obtain a weighted training image; Perform noise addition processing on the weighted training image through the noise addition unit to obtain the noise-added training image.

[0007] According to some embodiments of the first aspect of the present application, the step of inputting the noise-added training image into the denoising unit of the diffusion network module to obtain the predicted training noise and the reconstructed training image includes: Perform noise prediction processing on the noise-added training image based on the denoising unit to obtain the predicted training noise; Perform image reconstruction processing based on the predicted training noise and the noise-added training image to obtain the reconstructed training image.

[0008] According to some embodiments of the first aspect of the present application, the step of calculating the total loss based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the true label image includes: Calculate an image reconstruction loss based on the reconstructed training image and the initial training image; Calculate a denoising loss based on the predicted training noise; Calculate a mask consistency loss based on the predicted training noise and the segmentation mask probability training image; Calculate a segmentation loss based on the segmentation mask probability training image and the true label image; Calculate a sparsity loss based on the segmentation mask probability training image; Obtain the total loss based on the image reconstruction loss, the denoising loss, the mask consistency loss, the segmentation loss, and the sparsity loss.

[0009] According to some embodiments of the first aspect of the present application, the step of calculating the image reconstruction loss based on the reconstructed training image and the initial training image includes: The image reconstruction loss is calculated based on the reconstructed training image, the initial training image and a reconstruction loss function; the reconstruction loss function is: ; in, L recon Characterize the image reconstruction loss, I output characterizing the reconstructed training image, I org The initial training image is characterized.

[0010] According to some embodiments of the first aspect of the present application, the calculating the denoising loss based on the predicted training noise includes: The denoising loss is calculated based on the predicted training noise and the denoising loss function; the denoising loss function is: ; in, L denoise represents the denoising loss; E represents the mathematical expectation; is Gaussian noise; represents the predicted training noise, y0 represents the noisy training image; y t represents the noisy training image obtained by the diffusion network module at time step t; Mprob represents the segmentation mask probability training image.

[0011] According to some embodiments of the first aspect of the present application, the calculating the mask consistency loss based on the predicted training noise and the segmentation mask probability training image includes: The mask consistency loss is calculated based on the predicted training noise, the segmentation mask probability training image and the mask consistency loss function; the mask consistency loss function is: ; in, L mask is the mask consistency loss, E represents the mathematical expectation; is Gaussian noise, Characterize the prediction training noise, y t represents the noisy training image obtained by the diffusion network module at time step t; Mprob represents the segmentation mask probability training image, and y0 represents the noisy training image when the time step is 0.

[0012] According to some embodiments of the first aspect of the present application, the calculating the segmentation loss based on the segmentation mask probability training image and the true label image includes: Based on the segmentation mask probability training image, the ground truth label image, and the segmentation loss function, calculate the segmentation loss; the segmentation loss function is: ; where L seg represents the segmentation loss, represents the probability of pixel i in the segmentation mask probability training image; represents the probability of pixel i in the ground truth label image; N is the total number of pixels in the segmentation mask probability training image.

[0013] To achieve the above object, the second aspect embodiment of the present application provides an image reconstruction method, including: Obtain a to-be-reconstructed image with an occlusion area; Input the to-be-reconstructed image into a target image reconstruction model to obtain a target reconstructed image; wherein, the target image reconstruction model is trained according to the image reconstruction model training method of any one of the first aspect embodiments.

[0014] To achieve the above object, the third aspect embodiment of the present application provides an image reconstruction model training device, including: An acquisition module, configured to acquire an initial training image with an occlusion area and the ground truth label image corresponding to the initial training image; A first input module, configured to input the initial training image into an initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a pre-trained segmentation network module; A segmentation module, configured to perform image segmentation processing on the initial training image through the segmentation network module to obtain a segmentation mask probability training image; wherein, the first pixel of the segmentation mask probability training image corresponds one-to-one with the second pixel of the initial training image, and the value of the first pixel represents the probability that the corresponding second pixel is an occlusion area; A second input module, configured to input the segmentation mask probability training image and the initial training image into the noise addition unit of the diffusion network module to obtain a noise-added training image; A third input module, configured to input the noise-added training image into the denoising unit of the diffusion network module to obtain a predicted training noise and a reconstructed training image; A calculation module, configured to calculate a total loss based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the ground truth label image; An iterative update module, configured to iteratively update the initial image reconstruction model based on the total loss to obtain a trained target image reconstruction model.

[0015] To achieve the above object, an embodiment of the fourth aspect of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the image reconstruction model training method according to any one of the embodiments of the first aspect, or the image reconstruction method according to the embodiment of the second aspect.

[0016] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the image reconstruction model training method according to any one of the embodiments of the first aspect, or the image reconstruction method according to the embodiment of the second aspect.

[0017] According to the model training method, image reconstruction method, device, electronic device and medium of the embodiments of the present application, the initial training image is first input into the initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a pre-trained segmentation network module; image segmentation processing is performed by the pre-trained segmentation network module to obtain a segmentation mask probability training image, and the segmentation mask probability training image and the initial training image are input into the denoising unit of the diffusion network module to obtain a noisy training image; the noisy training image is input into the denoising unit of the diffusion network module to obtain predicted training noise and a reconstructed training image; based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the true label image, the total loss is calculated, and the initial image reconstruction model is iteratively updated based on the total loss to obtain a trained target image reconstruction model. In this process, since the segmentation network module is pre-trained, the accuracy of the segmentation network module is relatively high. In the segmentation mask probability training image output by the segmentation network module, the first pixel of the segmentation mask probability training image corresponds one-to-one to the second pixel of the initial training image, and the value of the first pixel represents the probability that the corresponding second pixel is an occluded area. Therefore, the information represented by the segmentation mask probability training image can be used as prior knowledge to help the diffusion network module learn the intrinsic mapping mechanism between the transparency of the occluded area and the information represented by the segmentation mask probability training image based on prior knowledge to accurately distinguish the edge and background pixels of the occluded area, thereby improving the ability to recognize occluded areas and overcome texture interference. Under the guidance of prior knowledge, the diffusion network module can deeply explore the background structure information that may exist in the occluded area, and while removing blurring objects such as flames and smoke, reasonably integrate this background information into the image reconstruction process. In this way, the diffusion network module can better preserve the background characteristics and logical structure of the original scene of the initial training image, help the reconstruction model solve the edge blur dilemma, improve the accuracy and effect of edge processing of the image reconstruction model, and thus significantly enhance the quality of image reconstruction. In this process, the information represented by the segmentation mask probability training image can divide the initial training image into occluded areas and non-occluded areas. In the denoising unit, the occluded areas are guided by noise, prompting the image reconstruction model to generate natural and continuous results in the subsequent denoising process.

[0018] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present application is further described below with reference to the accompanying drawings and embodiments, wherein: Figure 1 This is a flowchart of the steps of the image reconstruction model training method according to an embodiment of the present application; Figure 2Schematic diagram of the framework of the image reconstruction model according to an embodiment of the present application; Figure 3 Schematic diagram of a specific process of step S140 according to an embodiment of the present application; Figure 4 Schematic diagram of a specific process of step S150 according to an embodiment of the present application; Figure 5 Schematic diagram of a specific process of step S160 according to an embodiment of the present application; Figure 6 Schematic diagram of the step flow of the image reconstruction method according to an embodiment of the present application; Figure 7 Schematic diagram of the structure of the image reconstruction apparatus according to an embodiment of the present application; Figure 8 Schematic diagram of the structure of the image reconstruction model training apparatus according to an embodiment of the present application; Figure 9 Schematic diagram of the structure of the electronic device according to an embodiment of the present application. Detailed description of the specific implementation

[0020] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.

[0021] In the description of the present application, it should be understood that the orientation or positional relationship indicated by terms such as up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present application.

[0022] In the description of the present application, the meaning of several is more than one, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0023] In the description of the present application, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the present application in combination with the specific content of the technical solution.

[0024] In the description of this application, reference to the terms "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples.

[0025] Some methods in the related art have many limitations when processing translucent objects such as flames. As far as edge processing is concerned, when removing translucent objects such as flames, the edge blurring phenomenon is extremely prominent. This is mainly due to the fact that the translucent characteristics cause complex pixel changes in the transition area between the edge of the object and the background, and the texture information of the occluded area is intertwined with the background, resulting in serious texture interference, which greatly affects the accuracy and effect of the removal. In terms of occlusion processing and completion, the currently commonly used strategy is to directly remove the occluded area of translucent objects such as flames, and then fill in the gaps. However, this method has obvious drawbacks. It does not fully consider the background structure information contained in the translucent area. This information is lost during the removal and completion process, resulting in the destruction of the logical correlation between the reconstructed image and the original scene, ultimately resulting in poor image reconstruction quality.

[0026] Based on this, embodiments of the present application provide a model training method, image reconstruction method, apparatus, electronic device, and medium that can improve the reconstruction quality of an image reconstruction model for images with occluded areas. The information representing the segmentation mask probability training image output by the segmentation network module in the model can serve as prior knowledge. Under the guidance of this prior knowledge, the diffusion network module can deeply explore the background structure information that may exist in the occluded area. In this way, the diffusion network module can better preserve the background features and logical structure of the original scene of the initial training image, thereby significantly enhancing the quality of image reconstruction.

[0027] The first aspect of the present application provides an image reconstruction model training method. The image reconstruction model training method of the embodiment of the present application can be applied to a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a router, a programmable switch, a network card, etc.; the server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers. It can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the image reconstruction model training method, etc., but is not limited to the above forms.

[0028] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific operations or implement specific abstract data types. This application can also be practiced in a distributed computing environment where operations are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0029] Referring to Figure 1 , Figure 1 is a schematic flowchart of the steps of the image reconstruction model training method according to an embodiment of this application. The image reconstruction model training method may include but is not limited to steps S110 to S170.

[0030] Step S110, obtain an initial training image with an occlusion area and a corresponding ground truth image of the initial training image; It should be noted that an initial training image with an occlusion area can be obtained by photographing an object with flames or smoke through a camera, and the occlusion area is the flame part or the smoke part. In addition, the occlusion area can also be a cloud part, a water vapor part, a translucent glass part, a mosaic part, etc. After obtaining the initial training image, the initial training image is manually annotated to obtain the corresponding ground truth image of the initial training image. For example, the ground truth image is obtained by annotating based on the initial training image. The pixels in the occlusion area of the initial training image are annotated as 1, and the pixels in the non-occlusion area are annotated as 0. In this way, the corresponding ground truth image of the initial training image is obtained. In another embodiment, the initial training image and the corresponding ground truth image are obtained from a publicly available material library. There are various ways to obtain them from the material library. For example, the initial training image and the corresponding ground truth image can be obtained by accessing the material library through a custom application.

[0031] Step S120, input the initial training image into an initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a pre-trained segmentation network module; Step S130: Perform image segmentation processing on the initial training image through a segmentation network module to obtain a segmentation mask probability training image. The first pixel of the segmentation mask probability training image corresponds one-to-one with the second pixel of the initial training image, and the value of the first pixel represents the probability that the corresponding second pixel is an occluded area. It should be noted that referring to Figure 2 , Figure 2 is a schematic framework diagram of the image reconstruction model of the embodiment of the present application. The segmentation network module in the image reconstruction model adopts an encoder-decoder architecture with an attention mechanism. Specifically, in step S130, the initial training image is first converted into an input sequence. The encoder is responsible for converting the input sequence into a fixed-length context vector, and the decoder generates an output sequence based on this context vector. The processing process of the attention mechanism is expressed as: ; where Q, K, and V respectively represent the query vector, key vector, and value vector of the initial training image, and d k represents the model dimension. Through the attention mechanism, the segmentation network module can capture long-range dependencies in the initial training image, thereby better identifying occluded areas. The self-attention mechanism is incorporated into the encoder part as a feature enhancement module. This module captures the global context information of the occluded area through global feature modeling and enhances the feature saliency of the flame area through the channel attention mechanism. The processing of the segmentation network module can be expressed as: ; where Mprob represents the segmentation mask probability training image, f seg represents the segmentation network module, I orgCharacterize the initial training image. The segmentation mask probability training image represents the spatial distribution of the occluded area (i.e., flame or smoke). The value of each first pixel in the segmentation mask probability training image ranges from [0, 1]. Among them, a value close to 1 indicates a higher probability that the corresponding second pixel in the initial training image is in the occluded area, while a value close to 0 indicates a lower probability that the corresponding second pixel in the initial training image is in the occluded area. The value of the first pixel is greater than or equal to 0 and less than or equal to 1. For example, if the value of a first pixel is 0.9, the probability that the corresponding second pixel in the initial training image is in the occluded area is 90%; if the value of a first pixel is 0.01, the probability that the corresponding second pixel in the initial training image is in the occluded area is 1%. The segmentation mask probability training image not only provides the explicit position of the flame area but also provides spatial guidance information for the next-stage denoising learning. Usually, the occluded area only occupies a part of the image. Therefore, a sparsity constraint is imposed on the distribution of the segmentation mask probability training image to prevent the segmentation network module from generating a large range of unnecessary masks.

[0032] In one embodiment, the segmentation network module adopts the Attention U-Net network, which is a deep learning model that introduces an attention mechanism based on U-Net. U-Net is a convolutional neural network architecture commonly used in medical image segmentation, consisting of an encoder and a decoder. The feature maps of the encoder are passed to the decoder through skip connections to retain more spatial information.

[0033] Step S140: Input the segmentation mask probability training image and the initial training image into the noise addition unit of the diffusion network module to obtain a noisy training image; Step S150: Input the noisy training image into the denoising unit of the diffusion network module to obtain a predicted training noise and a reconstructed training image; In some embodiments, the diffusion network module includes a noise addition unit and a denoising unit. The diffusion network module adopts the DDPM (Denoising Diffusion Probabilistic Models) diffusion model. The diffusion network DDPM is a generative model based on variational inference and Markov chains, mainly used for image generation tasks. It generates high-quality images by learning to gradually add noise to the data (forward diffusion process) and then gradually remove the noise to restore the original data (reverse diffusion process). The noise addition unit is used to perform the forward diffusion process, and the denoising unit is used to perform the reverse diffusion process. Forward diffusion process: Starting from the original data, through a series of noise addition steps, the data is gradually transformed into Gaussian noise. This process can be regarded as a destructive process for the data, adding noise at each step and making the data closer and closer to random noise. Reverse diffusion process: Starting from pure noise, through a series of denoising steps, the original data is gradually restored. This process learns a neural network to predict the denoising process at each step, thus effectively restoring a clear image from the noise.

[0034] Step S160, calculate the total loss based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the true label image. Step S170, iteratively update the initial image reconstruction model based on the total loss to obtain the trained target image reconstruction model.

[0035] It should be noted that steps S110 to S170 are repeatedly executed. Each time the total loss is calculated, the parameter set of the image reconstruction model is updated once. For example, after obtaining the total loss, the gradient descent algorithm is used to update the parameter set. In this way, the iterative update of the image reconstruction model is realized until the number of iterations reaches the preset iteration threshold.

[0036] The image reconstruction model training method of the embodiment of the present application completes the training of the initial image reconstruction model through the above-mentioned steps S110 to S170, thereby obtaining a trained target image reconstruction model. First, the initial training image is input into the initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a pre-trained segmentation network module; image segmentation processing is performed by the pre-trained segmentation network module to obtain a segmentation mask probability training image, the segmentation mask probability training image and the initial training image are input into the denoising unit of the diffusion network module to obtain a noisy training image; the noisy training image is input into the denoising unit of the diffusion network module to obtain predicted training noise and a reconstructed training image; based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the true label image, the total loss is calculated, and the initial image reconstruction model is iteratively updated based on the total loss to obtain a trained target image reconstruction model. In this process, since the segmentation network module is pre-trained, the accuracy of the segmentation network module is relatively high. In the segmentation mask probability training image output by the segmentation network module, the first pixel of the segmentation mask probability training image corresponds one-to-one to the second pixel of the initial training image, and the value of the first pixel represents the probability that the corresponding second pixel is an occluded area. Therefore, the information represented by the segmentation mask probability training image can be used as prior knowledge to help the diffusion network module learn the intrinsic mapping mechanism between the transparency of the occluded area and the information represented by the segmentation mask probability training image based on prior knowledge to accurately distinguish the edge and background pixels of the occluded area, thereby improving the ability to recognize occluded areas and overcome texture interference. Under the guidance of prior knowledge, the diffusion network module can deeply explore the background structure information that may exist in the occluded area, and while removing blurring objects such as flames and smoke, reasonably integrate this background information into the image reconstruction process. In this way, the diffusion network module can better preserve the background characteristics and logical structure of the original scene of the initial training image, help the reconstruction model solve the edge blur dilemma, improve the accuracy and effect of edge processing of the image reconstruction model, and thus significantly enhance the quality of image reconstruction. In this process, the information represented by the segmentation mask probability training image can divide the initial training image into occluded areas and non-occluded areas. In the denoising unit, the occluded areas are guided by noise, prompting the image reconstruction model to generate natural and continuous results in the subsequent denoising process.

[0037] In some embodiments, reference Figure 3 , Figure 3 This is a specific flow chart of step S140 in an embodiment of the present application. Step S140 may include but is not limited to step S310 and step S320.

[0038] Step S310, performing a pixel-by-pixel weighting operation on the initial training image and the segmentation mask probability training image to obtain a weighted training image; In one embodiment, the pixels in the initial training image correspond one-to-one with the pixels in the segmentation mask probability training image. The process of per-pixel weighting operation is as follows: multiply the value of the first pixel in the initial training image by the value of the corresponding second pixel in the segmentation mask probability training image. The per-pixel weighting operation can weaken the contribution of the occluded region and add noise to the background region. The initial training image is segmented into known regions (non-occluded regions) and unknown regions (occluded regions) under the action of the segmentation mask probability training image. In the subsequent noise addition process and denoising process, in the weighted training image, the value of the first pixel in the segmentation mask probability training image serves as the weight for the noise addition process, realizing adding more noise to the unknown region to learn how to complete it. The known region is only slightly adjusted so that the semantic information in the background part of the initial training image will not be lost due to excessive noise. The region to be removed is guided by high noise, prompting the model to generate natural and continuous results in the subsequent denoising process.

[0039] Step S320, perform noise addition processing on the weighted training image through a noise addition unit to obtain a noisy training image.

[0040] In one embodiment, the noise addition processing process is expressed as: ; where is the noise attenuation coefficient; is the Gaussian noise used to remove the occluded region, is the low-weight noise in the non-occluded region, Mprob represents the segmentation mask probability training image, I org characterizes the initial training image, characterizes the noisy training image. In this way, more noise is added to the occluded region to quickly degrade it into pure noise, more image information is retained in the background region, and unnecessary noise is reduced. Through the noise addition processing process, since the first pixel in the segmentation mask probability training image corresponds one-to-one with the second pixel in the initial training image, and the value of the first pixel represents the probability that the corresponding second pixel is an occluded region, therefore, combining the expression of the noise addition processing process, the value of the first pixel is equivalent to the weight of the noise. The noise addition process realizes adding Gaussian noise to the occluded region, and the greater the probability of the occluded region, the greater the added noise, while the non-occluded region is added with low-weight noise. The noisy training image contains the mixed information of the occluded region and the background region (i.e., the non-occluded region), enabling the diffusion network module to better learn the features of eliminating the occluded region.

[0041] ​​​​The present application first performs a pixel-by-pixel weighted operation on the initial training image and the segmentation mask probability training image to obtain a weighted training image through steps S310 to S320, and then performs noise processing on the weighted training image. In this way, the weighted training image can cover the probability information of the occluded area in the segmentation mask probability training image, so that the information represented by the segmentation mask probability training image output by the segmentation network module can be used as prior knowledge to help the diffusion network module learn the intrinsic mapping mechanism between the transparency of the occluded area and the information represented by the segmentation mask probability training image based on the prior knowledge to accurately distinguish the edge of the occluded area and the background pixels, so as to improve the occluded area recognition ability and overcome texture interference ability.

[0042] In some embodiments, reference Figure 4 , Figure 4 1 is a specific flow chart of step S150 of the embodiment of the present application. Step S150 may include but is not limited to steps S410 to S420.

[0043] Step S410, performing noise prediction processing on the noisy training image based on the denoising unit to obtain predicted training noise; Step S420 : performing image reconstruction processing based on the predicted training noise and the noisy training image to obtain a reconstructed training image.

[0044] In one embodiment, the image reconstruction process is expressed as follows: ; ; It is worth noting that Mprob represents the segmentation mask probability training image, I org Representing the initial training image, in the diffusion network module, the image reconstruction process is to perform t-step denoising processing, where The diffusion network module generates content based on the non-occluded area in step t-1; t-1 Represents the denoised image generated at step t-1, y t represents the denoised image generated at step t, and y t The image reconstruction process adds more noise to unknown (occluded) areas to learn how to complete them, while making only slight adjustments to known (non-occluded) areas to ensure that semantic information in the background is not lost due to excessive noise.

[0045] , The diffusion network module generates content based on the occluded area at step t-1. Characterizes the variance of the predicted training noise, Characterize the mean of the prediction training noise. The mean of the prediction training noise is expressed as: ; Wherein, Characterize the prediction training noise, wherein is the noise attenuation coefficient; Mprob Represents the segmentation mask probability training image, β t is a preset coefficient.

[0046] It should be noted that through the above steps S410 to S420, after the denoising unit of the diffusion network module obtains the prediction training noise through prediction, based on the prediction training noise, the denoising process is performed on the noise-added training image, and the denoising process is the image reconstruction process, so as to obtain the reconstructed training image. The noise prediction process and the image reconstruction process are based on the structure of the diffusion network module, and the present application does not specifically limit the structure of the diffusion network module. The diffusion network module adopts the diffusion model DDPM.

[0047] In some embodiments, referring to Figure 5 , Figure 5 is a specific process schematic diagram of step S160 of the embodiment of the present application. Step S160 may include but is not limited to steps S510 to S560.

[0048] Step S510, calculate the image reconstruction loss based on the reconstructed training image and the initial training image; Step S520, calculate the denoising loss based on the prediction training noise; Step S530, calculate the mask consistency loss based on the prediction training noise and the segmentation mask probability training image; Step S540, calculate the segmentation loss based on the segmentation mask probability training image and the true label image; Step S550, calculate the sparsity loss based on the segmentation mask probability training image; Step S560, obtain the total loss based on the image reconstruction loss, the denoising loss, the mask consistency loss, the segmentation loss and the sparsity loss.

[0049] It should be noted that in the embodiments of the present application, by steps S510 to S560, the image reconstruction loss, denoising loss, mask consistency loss, segmentation loss, and sparsity loss are calculated respectively, so as to calculate the total loss, and based on the total loss, iterative update is performed to obtain the trained target image reconstruction model. In this way, during the iterative update process, the sparsity of the intermediate representation and output of the image reconstruction model can be constrained by the sparsity loss; by the denoising loss, the prediction accuracy of the image reconstruction model for noise can be improved; by the image reconstruction loss, the accuracy of the image reconstruction model can be improved; by the segmentation loss, the accuracy of the image reconstruction model in identifying occluded regions can be improved. In this way, by iteratively updating the image reconstruction model with multiple losses, the reconstruction quality of the image reconstruction model can be improved.

[0050] In one embodiment, based on the reconstructed training image and the initial training image, the image reconstruction loss is calculated, including: Based on the reconstructed training image, the initial training image, and the reconstruction loss function, the image reconstruction loss is calculated; the reconstruction loss function is: ; Wherein, L recon represents the image reconstruction loss, I output represents the reconstructed training image, I org represents the initial training image.

[0051] In one embodiment, based on the predicted training noise, the denoising loss is calculated, including: Based on the predicted training noise and the denoising loss function, the denoising loss is calculated; the denoising loss function is: ; Wherein, L denoise represents the denoising loss; E represents the mathematical expectation; is Gaussian noise; represents the predicted training noise, and y0 represents the noise-added training image; y t represents the noise-added training image obtained by the diffusion network module at time step t; Mprob represents the segmentation mask probability training image. The denoising loss can guide the model to learn to gradually remove noise from the noise-added image, and can improve the prediction accuracy of the image reconstruction model for noise.

[0052] In one embodiment, based on the predicted training noise and the segmentation mask probability training image, the mask consistency loss is calculated, including: Based on the predicted training noise and the segmentation mask probability, train the image and mask consistency loss function to calculate the mask consistency loss; the mask consistency loss function is: ; Wherein, L mask is the mask consistency loss, and E represents the mathematical expectation; is Gaussian noise, represents the predicted training noise, y t represents the noisy training image obtained by the diffusion network module at time step t; Mprob represents the segmentation mask probability training image, and y0 represents the noisy training image at time step 0.

[0053] In one embodiment, based on the segmentation mask probability training image and the true label image, calculate the segmentation loss, including: Based on the segmentation mask probability training image, the true label image and the segmentation loss function, calculate the segmentation loss; the segmentation loss function is: ; Wherein, L seg represents the segmentation loss, represents the probability of pixel i in the segmentation mask probability training image; represents the probability of pixel i in the true label image; N is the total number of pixels in the segmentation mask probability training image.

[0054] In one embodiment, calculate the sparsity loss based on the segmentation mask probability training image, including: Based on the segmentation mask probability training image and the sparsity loss function, calculate the sparsity loss; the sparsity loss function is: ; Wherein, L sparisty represents the sparsity loss, and Mprob represents the segmentation mask probability training image.

[0055] In one embodiment, based on the image reconstruction loss, the denoising loss, the mask consistency loss, the segmentation loss and the sparsity loss, obtain the total loss, which is specifically expressed as: ; Wherein, L is the total loss, L sparisty represents the sparsity loss, L seg represents the segmentation loss, L mask is the mask consistency loss, L denoiseCharacterize the denoising loss; L recon Characterize the image reconstruction loss. is the weight of the denoising loss, is the weight of the mask consistency loss, is the weight of the image reconstruction loss, is the weight of the sparsity loss, is the weight of the segmentation loss. Those skilled in the art can set , , , , values according to the actual situation.

[0056] In some embodiments, the target image reconstruction model of the embodiments of the present application is applied to assist in fire accident investigation and is used for the restoration and reconstruction of fire images. Generally, flames and smoke coexist, so fire images usually have both flame regions and smoke regions. In this embodiment, an RGB image with flame regions and smoke regions is selected as the initial training image. Specifically, before step S120, preprocessing is performed on the initial training image, and the preprocessing steps include: Step S111, convert the initial training image from the three-primary color space to the hexagonal pyramid model color space; It should be noted that the three-primary color space refers to a color space composed of three primary colors. Among them, the primary colors are usually red, green, and blue, corresponding to the three channels in the RGB color model respectively. In the RGB color space, the intensity of each color can be represented by the mixture of different degrees of red, green, and blue. The hexagonal pyramid model color (HSV) space is a color space created based on the intuitiveness of colors. Among them, each color is represented by hue, saturation, and value.

[0057] Step S112, regard the region in the initial training image where the hue is in the first preset region, the saturation is greater than the first preset value, and the value is greater than the second preset value as the first region; It should be noted that the first region represents the flame region in the initial training image. In one embodiment, the first preset interval is [0, 30], the first preset value is 150, and the second preset value is 200. The characteristics of the flame are that the color is red or dark red, and the saturation and value are relatively high. Therefore, through the first preset interval, the first preset value, and the second preset value, the flame region can be determined. It should be noted that those skilled in the art can adjust the first preset interval, the first preset value, and the second preset value based on the actual situation.

[0058] Step S113: Take the area in the initial training image where the hue is within the second preset range, the saturation is less than the third preset value, and the lightness is less than the fourth preset value as the second area; It should be noted that the second area represents the smoke area in the initial training image. In one embodiment, the second preset range is [90, 120], the third preset value is 50, and the fourth preset value is 50. The characteristics of the smoke area are gray, low saturation, and medium - low brightness. Therefore, based on the second preset range, the third preset value, and the fourth preset value, the smoke area can be determined. It should be noted that those skilled in the art can adjust the second preset range, the third preset value, and the fourth preset value according to the actual situation.

[0059] Step S114: Subtract the first area and the second area from the initial training image to obtain the third area; It should be noted that the third area is the background area outside the flame area and the smoke area.

[0060] Step S115: Increase the saturation and lightness of the first area to obtain the first updated area; Specifically, increasing the saturation and lightness of the first area can make the flame area present a more vivid red / orange color, and the saturation is significantly improved.

[0061] Exemplarily, the saturation of the pixel points in the first updated area = the saturation of the corresponding pixel points in the first area + 30. The lightness of the pixel points in the first updated area = the lightness of the corresponding pixel points in the first area + 10.

[0062] Step S116: Decrease the hue and saturation of the second area and increase the lightness of the second area to obtain the second updated area; Specifically, decreasing the hue and saturation of the second area can weaken the hue interference in the smoke area and reduce the color interference, while increasing the lightness of the second area can highlight the smoke texture.

[0063] Exemplarily, the hue of the pixel points in the second updated area = the hue of the corresponding pixel points in the second area - 10; the saturation of the pixel points in the second updated area = the saturation of the corresponding pixel points in the second area * 0.3. The lightness of the pixel points in the second updated area = the lightness of the corresponding pixel points in the second area + 10.

[0064] Step S117: Perform a merging process on the third area, the first updated area, and the second updated area to obtain a merged image; Step S118: Convert the merged image from the hexacone model color space to the three - primary - color space to obtain a new initial training image.

[0065] During the preprocessing process of the embodiments of the present application, through steps S111 to S118, in the initial training image, the flame area presents a brighter red / orange color, and the saturation is significantly improved; the color interference in the occlusion area is reduced, and the texture is clearer. Inputting the new initial training image into the image reconstruction model helps to improve the recognition ability of the segmentation network module in the image reconstruction model for the smoke area and the flame area, thereby improving the quality of the reconstructed image output by the target image reconstruction model.

[0066] The second aspect of the embodiments of the present application provides an image reconstruction method. Refer to Figure 6 , Figure 6 which is the schematic flowchart of the steps of the image reconstruction method of the embodiments of the present application. The image reconstruction method includes but is not limited to steps S610 and S620.

[0067] Step S610, obtaining a to-be-reconstructed image with an occlusion area; It should be noted that a camera can be used to capture an object with flames or smoke, thereby obtaining a to-be-reconstructed image with an occlusion area, and the occlusion area is the flame part or the smoke part.

[0068] Step S620, inputting the to-be-reconstructed image into the target image reconstruction model to obtain a target reconstructed image; wherein, the target image reconstruction model is trained according to the image reconstruction model training method of any one of the first aspect embodiments.

[0069] It should be noted that the image reconstruction model can be trained through Figure 1 the steps S110 to S170 shown. Specifically, after inputting the to-be-reconstructed image into the trained target image reconstruction model, the to-be-reconstructed image is subjected to image segmentation processing through the segmentation network module of the image reconstruction model to obtain a segmentation mask image; the segmentation mask image and the to-be-reconstructed image are subjected to pixel-by-pixel weighting processing to obtain a weighted image; the weighted image is subjected to noise addition processing through the noise addition unit of the diffusion network module of the image reconstruction model to obtain a noisy image; the noisy image is subjected to noise prediction processing through the denoising unit of the diffusion network module to obtain a predicted noise, and then based on the predicted noise and the noisy image, the noisy image is subjected to denoising processing to obtain a target reconstructed image.

[0070] It should be noted that since the image reconstruction model in the image reconstruction method of the second application embodiment of the present application is obtained by the image reconstruction model training method of the first aspect embodiment, therefore, the specific implementation manner of the image reconstruction method of the second embodiment of the present application is basically the same as the specific embodiment of the above image reconstruction model training method and has the same beneficial effects, which will not be elaborated here.

[0071] Some other embodiments of the present application also provide an image reconstruction device. Refer toFigure 7 , Figure 7 is a schematic structural diagram of an image reconstruction device according to an embodiment of the present application. The image reconstruction device includes: A to-be-reconstructed image acquisition module 710, configured to acquire a to-be-reconstructed image with an occlusion area; A reconstruction module 720, configured to input the to-be-reconstructed image into a target image reconstruction model to obtain a target reconstructed image; wherein, the target image reconstruction model is trained according to the image reconstruction model training method of any one of the embodiments of the first aspect.

[0072] The image reconstruction device according to the embodiment of the present application is used to execute the image reconstruction method according to the embodiment of the second aspect of the present application. When executing the method, after inputting the to-be-reconstructed image into the trained target image reconstruction model, the to-be-reconstructed image is subjected to image segmentation processing through the segmentation network module of the image reconstruction model to obtain a segmentation mask image; the segmentation mask image and the to-be-reconstructed image are subjected to pixel-by-pixel weighting processing to obtain a weighted image; the weighted image is subjected to noise addition processing through the noise addition unit of the diffusion network module of the image reconstruction model to obtain a noise-added image; the noise-added image is subjected to noise prediction processing through the denoising unit of the diffusion network module to obtain predicted noise, and then based on the predicted noise and the noise-added image, the noise-added image is subjected to denoising processing to obtain a target reconstructed image.

[0073] An embodiment of the third aspect of the present application provides an image reconstruction model training device. Refer to Figure 7 , Figure 7 is a schematic structural diagram of the image reconstruction model training device according to the embodiment of the present application. The image reconstruction model training device includes: An acquisition module 810, configured to acquire an initial training image with an occlusion area and a true label image corresponding to the initial training image; A first input module 820, configured to input the initial training image into an initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a segmentation network module that has completed pre-training; A segmentation module 830, configured to perform image segmentation processing on the initial training image through the segmentation network module to obtain a segmentation mask probability training image; wherein, the first pixel of the segmentation mask probability training image corresponds to the second pixel of the initial training image one by one, and the value of the first pixel represents the probability that the corresponding second pixel is an occlusion area; A second input module 840, configured to input the segmentation mask probability training image and the initial training image into the noise addition unit of the diffusion network module to obtain a noise-added training image; A third input module 850, configured to input the noise-added training image into the denoising unit of the diffusion network module to obtain predicted training noise and a reconstructed training image; A calculation module 860 is configured to calculate a total loss based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the true label image; The iterative update module 880 is used to iteratively update the initial image reconstruction model based on the total loss to obtain a trained target image reconstruction model.

[0074] It is worth noting that the image reconstruction model training device is used to execute the image reconstruction model training method of the first aspect of the embodiment of the present application. When executing the method, the initial training image is first input into the initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a pre-trained segmentation network module; image segmentation processing is performed by the pre-trained segmentation network module to obtain a segmentation mask probability training image, and the segmentation mask probability training image and the initial training image are input into the denoising unit of the diffusion network module to obtain a noisy training image; the noisy training image is input into the denoising unit of the diffusion network module to obtain predicted training noise and reconstructed training image; based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the true label image, the total loss is calculated, and the initial image reconstruction model is iteratively updated based on the total loss to obtain a trained target image reconstruction model. In this process, since the segmentation network module is pre-trained, the accuracy of the segmentation network module is relatively high. In the segmentation mask probability training image output by the segmentation network module, the first pixel of the segmentation mask probability training image corresponds one-to-one to the second pixel of the initial training image, and the value of the first pixel represents the probability that the corresponding second pixel is an occluded area. Therefore, the information represented by the segmentation mask probability training image can be used as prior knowledge to help the diffusion network module learn the intrinsic mapping mechanism between the transparency of the occluded area and the information represented by the segmentation mask probability training image based on prior knowledge to accurately distinguish the edge and background pixels of the occluded area, thereby improving the ability to recognize occluded areas and overcome texture interference. Under the guidance of prior knowledge, the diffusion network module can deeply explore the background structure information that may exist in the occluded area, and while removing blurring objects such as flames and smoke, reasonably integrate this background information into the image reconstruction process. In this way, the diffusion network module can better preserve the background characteristics and logical structure of the original scene of the initial training image, help the reconstruction model solve the edge blur dilemma, improve the accuracy and effect of edge processing of the image reconstruction model, and thus significantly enhance the quality of image reconstruction.

[0075] In some embodiments, the second input module 840 includes a pixel-by-pixel weighting submodule and a noise adding submodule.

[0076] The pixel-by-pixel weighting submodule is used to perform pixel-by-pixel weighting operations on the initial training image and the segmentation mask probability training image to obtain a weighted training image; The noise adding module is used to perform noise adding processing on the weighted training image through the noise adding unit to obtain the noisy training image.

[0077] In some embodiments, the third input module 850 includes a prediction sub-module and a reconstruction sub-module.

[0078] The prediction sub-module is used to perform noise prediction processing on the noisy training image based on the denoising unit to obtain the predicted training noise; The reconstruction sub-module is used to perform image reconstruction processing based on the predicted training noise and the noisy training image to obtain the reconstructed training image.

[0079] In some embodiments, the calculation module 860 includes a first calculation sub-module, a second calculation sub-module, a third calculation sub-module, a fourth calculation sub-module, a fifth calculation sub-module, and a sixth calculation sub-module.

[0080] The first calculation sub-module is used to calculate the image reconstruction loss based on the reconstructed training image and the initial training image; The second calculation sub-module is used to calculate the denoising loss based on the predicted training noise; The third calculation sub-module is used to calculate the mask consistency loss based on the predicted training noise and the segmentation mask probability training image; The fourth calculation sub-module is used to calculate the segmentation loss based on the segmentation mask probability training image and the true label image; [[ID=2,1]] The fifth calculation sub-module is used to calculate the sparsity loss based on the segmentation mask probability training image; The sixth calculation sub-module is used to obtain the total loss based on the image reconstruction loss, the denoising loss, the mask consistency loss, the segmentation loss, and the sparsity loss.

[0081] In some embodiments, the first calculation sub-module is specifically used to calculate the image reconstruction loss based on the reconstructed training image, the initial training image, and the image reconstruction loss function.

[0082] In some embodiments, the second calculation sub-module is specifically used to calculate the denoising loss based on the predicted training noise and the denoising loss function.

[0083] In some embodiments, the third calculation sub-module is specifically used to calculate the mask consistency loss based on the predicted training noise, the segmentation mask probability training image, and the mask consistency loss function.

[0084] In some embodiments, the fourth calculation sub-module is specifically used to calculate the segmentation loss based on the segmentation mask probability training image, the true label image, and the segmentation loss function.

[0085] In a fourth aspect embodiment of the present application, an electronic device is provided. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the image reconstruction model training method according to any one of the embodiments of the first aspect or the image reconstruction method according to the embodiment of the second aspect. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0086] Referring to Figure 9 , Figure 9 FIG. [FIGURE NUMBER] is a schematic structural diagram of an electronic device according to an embodiment. The electronic device includes: A processor 901, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application; A memory 902, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and the processor 901 is called to execute the image reconstruction model training method or the image reconstruction method of the embodiments of the present application; An input / output interface 903, which is used to implement information input and output; A communication interface 904, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904); Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.

[0087] In a fifth aspect embodiment of the present application, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the image reconstruction model training method according to any one of the embodiments of the first aspect or the image reconstruction method according to the embodiment of the second aspect.

[0088] Please note that the figure number in "FIG. [FIGURE NUMBER] is a schematic structural diagram of an electronic device according to an embodiment." should be replaced with the actual figure number.As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0089] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0090] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0091] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0092] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0093] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0094] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the mapping relationship of mapped objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the front and back mapped objects. "At least one (item) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0095] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0096] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0097] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0098] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0099] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall fall within the scope of rights of the embodiments of this application.

Claims

1. A method for training an image reconstruction model, characterized in that, Including: Obtain an initial training image with an occlusion area and a corresponding ground truth label image of the initial training image; Input the initial training image into an initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a pre-trained segmentation network module; Perform image segmentation processing on the initial training image through the segmentation network module to obtain a segmentation mask probability training image; wherein, the first pixel of the segmentation mask probability training image corresponds one-to-one with the second pixel of the initial training image, and the value of the first pixel represents the probability that the corresponding second pixel is an occlusion area; Input the segmentation mask probability training image and the initial training image into the noise addition unit of the diffusion network module to obtain a noise-added training image; Input the noise-added training image into the denoising unit of the diffusion network module to obtain a predicted training noise and a reconstructed training image; Calculate a total loss based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the ground truth label image; Iteratively update the initial image reconstruction model based on the total loss to obtain a trained target image reconstruction model.

2. The method for training an image reconstruction model according to claim 1, wherein The step of inputting the segmentation mask probability training image and the initial training image into the noise addition unit of the diffusion network module to obtain a noise-added training image includes: Perform a pixel-wise weighting operation on the initial training image and the segmentation mask probability training image to obtain a weighted training image; Perform noise addition processing on the weighted training image through the noise addition unit to obtain the noise-added training image.

3. The image reconstruction model training method according to claim 1, wherein The step of inputting the noise-added training image into the denoising unit of the diffusion network module to obtain a predicted training noise and a reconstructed training image includes: Perform noise prediction processing on the noise-added training image based on the denoising unit to obtain a predicted training noise; Perform image reconstruction processing on the noise-added training image based on the predicted training noise to obtain the reconstructed training image.

4. The method for training an image reconstruction model according to claim 1, wherein The step of calculating the total loss based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the ground truth label image includes: Calculate an image reconstruction loss based on the reconstructed training image and the initial training image; Calculate a denoising loss based on the predicted training noise; Calculate a mask consistency loss based on the predicted training noise and the segmentation mask probability training image; Calculate a segmentation loss based on the segmentation mask probability training image and the ground truth label image; Calculate a sparsity loss based on the segmentation mask probability training image; Obtain the total loss based on the image reconstruction loss, the denoising loss, the mask consistency loss, the segmentation loss, and the sparsity loss.

5. The image reconstruction model training method according to claim 4, wherein The step of calculating the image reconstruction loss based on the reconstructed training image and the initial training image includes: Calculate the image reconstruction loss based on the reconstructed training image, the initial training image, and a reconstruction loss function; the reconstruction loss function is: ; Among them, L recon characterizes the image reconstruction loss, I output characterizes the reconstructed training image, I org characterizes the initial training image.

6. The image reconstruction model training method according to claim 4, wherein The denoising loss calculated based on the predicted training noise includes: Calculating the denoising loss based on the predicted training noise and a denoising loss function; the denoising loss function is: ; Among them, L denoise represents the denoising loss; E represents the mathematical expectation; is Gaussian noise; represents the predicted training noise, and y0 represents the noisy training image; y t represents the noisy training image obtained by the diffusion network module at time step t; Mprob represents the segmentation mask probability training image.

7. The method for training an image reconstruction model according to claim 4, wherein The mask consistency loss calculated based on the predicted training noise and the segmentation mask probability training image includes: Calculating the mask consistency loss based on the predicted training noise, the segmentation mask probability training image, and a mask consistency loss function; the mask consistency loss function is: ; wherein, L mask is the mask consistency loss, and E represents the mathematical expectation; is Gaussian noise, represents the predicted training noise, y t represents the noisy training image obtained by the diffusion network module at time step t; Mprob represents the segmentation mask probability training image, and y0 represents the noisy training image at time step 0.

8. The method for training an image reconstruction model according to claim 4, wherein The segmentation loss calculated based on the segmentation mask probability training image and the ground truth label image includes: Calculating the segmentation loss based on the segmentation mask probability training image, the ground truth label image, and a segmentation loss function; the segmentation loss function is: ; Among them, L seg represents the segmentation loss, represents the probability of pixel i in the segmentation mask probability training image; represents the probability of pixel i in the ground truth label image; N is the total number of pixels in the segmentation mask probability training image.

9. An image reconstruction method, characterized in that, Including: Obtaining a to-be-reconstructed image with an occluded region; Inputting the to-be-reconstructed image into a target image reconstruction model to obtain a target reconstructed image; wherein, the target image reconstruction model is trained according to the image reconstruction model training method according to any one of claims 1 to 8.

10. An image reconstruction model training device, characterized in that, Including: An acquisition module, configured to acquire an initial training image with an occluded region and a ground truth label image corresponding to the initial training image; A first input module, configured to input the initial training image into an initial image reconstruction model; the initial image reconstruction model includes an initial diffusion network module and a pre-trained segmentation network module; A segmentation module, configured to perform image segmentation processing on the initial training image through the segmentation network module to obtain a segmentation mask probability training image; wherein, a first pixel of the segmentation mask probability training image corresponds to a second pixel of the initial training image one by one, and a value of the first pixel represents a probability that the corresponding second pixel is an occluded region; A second input module, configured to input the segmentation mask probability training image and the initial training image into a noise addition unit of the diffusion network module to obtain a noise-added training image; A third input module, configured to input the noise-added training image into a denoising unit of the diffusion network module to obtain a predicted training noise and a reconstructed training image; A calculation module, configured to calculate a total loss based on the reconstructed training image, the initial training image, the predicted training noise, the segmentation mask probability training image, and the ground truth label image; An iterative update module, configured to iteratively update the initial image reconstruction model based on the total loss to obtain a trained target image reconstruction model.

11. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the image reconstruction model training method according to any one of claims 1 to 8, or the image reconstruction method according to claim 9.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the image reconstruction model training method according to any one of claims 1 to 8, or the image reconstruction method according to claim 9.

Citation Information

Patent Citations

  • Image reconstruction model training method and device, equipment and medium

    CN118096560A

  • Image defogging method based on physical prior

    CN119006339A

  • Skin lesion segmentation method based on auxiliary task and mutual information loss

    CN119445104A

  • Modifying digital images by moving objects to create depth effect

    DE102023131117A1

Cited By

  • Image data reconstruction method and system based on spatial region perception and mask guidance

    CN121639959A

  • Image Data Reconstruction Method and System Based on Spatial Region Awareness and Mask-Guided Design

    CN121639959B