Image processing method and device, readable storage medium and electronic equipment

CN122820482APending Publication Date: 2026-09-25XG TECHNOLOGIES PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610964719.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2026-06-25
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

这样的处理方式,由于需要采用两套独立的神经网络,两套神经网络的参数众多、结构复杂性高,增加了硬件的计算开销和内存占用量,且两种数据分开处理导致计算效率低、实时性差

Benefits of technology

[0008]本公开实施例的第四个方面,提供了一种电子设备,该电子设备包括:存储器,用于存储处理器可执行指令;处理器,用于从存储器中读取可执行指令,并执行指令以实现本公开任一实施例的图像处理方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820482A_ABST
    Figure CN122820482A_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing method, device, readable storage medium and electronic equipment, which comprises the following steps: obtaining a scene image and a photon statistics graph corresponding to a target scene; using a shared encoding network included in a pre-trained image processing model to encode the scene image and the photon statistics graph respectively to obtain image features and photon statistics features; using a deep decoder included in the image processing model to decode the photon statistics features, and using the decoded photon statistics features to generate a depth map corresponding to the target scene; and using an image decoder included in the image processing model to decode the image features to obtain a denoised image. The shared encoding network of the embodiment of the present disclosure can effectively utilize the complementarity between two-dimensional features and three-dimensional features, improve the accuracy of the depth map decoded by the model and the quality of the denoised image, reduce the parameter redundancy of the image processing model, and improve the processing efficiency of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to computer vision technology, deep learning technology, and in particular to an image processing method, apparatus, readable storage medium, and electronic device. Background Technology

[0002] TOF (Time of Flight) data and image data are commonly used in the field of machine vision. The two can be combined for multimodal perception systems to capture the three-dimensional geometric structure and two-dimensional visual features of objects, thus overcoming the limitations of single sensors in complex environments.

[0003] In related technologies, two independent neural network models are typically used to process the Time-of-Flight (TOF) data and image data separately to obtain a depth map and a denoised image. This approach requires two independent neural networks, which have numerous parameters and high structural complexity, increasing hardware computational overhead and memory usage. Furthermore, processing the two types of data separately leads to low computational efficiency and poor real-time performance. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides an image processing method, apparatus, readable storage medium, and electronic device that can utilize a shared coding network to extract features from scene images and photon statistical maps. The decoded data can more accurately represent the features of the actual scene and effectively reduce the model's parameters, thereby improving model processing efficiency.

[0005] A first aspect of this disclosure provides an image processing method, comprising: acquiring a scene image and a photon statistical map corresponding to a target scene, wherein for any pixel in the photon statistical map, the pixel corresponds to a photon statistical data sequence, the photon statistical data in the photon statistical data sequence is used to characterize the number of reflected photons received from a target location at a corresponding time, the target location being the position in the target scene corresponding to the pixel; encoding the scene image and the photon statistical map respectively using a shared coding network included in a pre-trained image processing model to obtain image features and photon statistical features; decoding the photon statistical features using a depth decoder included in the image processing model, and generating a depth map corresponding to the target scene using the decoded photon statistical features; and decoding the image features using an image decoder included in the image processing model to obtain a denoised image.

[0006] A second aspect of this disclosure provides an image processing apparatus, comprising: an acquisition module, configured to acquire a scene image and a photon statistical map corresponding to a target scene, wherein for any pixel in the photon statistical map, the pixel corresponds to a photon statistical data sequence, the photon statistical data in the photon statistical data sequence is used to characterize the number of reflected photons received from a target location at a corresponding time, the target location being the position in the target scene corresponding to the pixel; an encoding module, configured to encode the scene image and the photon statistical map respectively using a shared encoding network included in a pre-trained image processing model to obtain image features and photon statistical features; a first decoding module, configured to decode the photon statistical features using a depth decoder included in the image processing model, and generate a depth map corresponding to the target scene using the decoded photon statistical features; and a second decoding module, configured to decode the image features using an image decoder included in the image processing model to obtain a denoised image.

[0007] A third aspect of this disclosure is to provide a computer-readable storage medium storing a computer program that, when executed, implements the image processing method of any embodiment of this disclosure.

[0008] A fourth aspect of the present disclosure provides an electronic device comprising: a memory for storing processor-executable instructions; and a processor for reading the executable instructions from the memory and executing the instructions to implement the image processing method of any embodiment of the present disclosure.

[0009] A fifth aspect of this disclosure provides a computer program product including computer program instructions that, when executed by an instruction processor, perform an image processing method according to any embodiment of this disclosure.

[0010] Based on the embodiments of this disclosure, by acquiring the scene image and photon statistical map corresponding to the target scene, the scene image and photon statistical map are encoded separately using a shared coding network included in a pre-trained image processing model to obtain image features and photon statistical features. A depth decoder is used to decode the photon statistical features, and a depth map corresponding to the target scene is generated using the decoded photon statistical features. Finally, an image decoder is used to decode the image features to obtain a denoised image. Because a shared coding network is used to encode the scene image and photon statistical map, this shared coding network can extract both two-dimensional and three-dimensional features contained in the input data. The features output by the shared coding network can reflect the complementarity between two-dimensional and three-dimensional features. Utilizing this complementarity, the photon statistical features and image features are decoded more accurately, improving the accuracy of the depth map after model decoding and the quality of the denoised image. Furthermore, using a shared coding network avoids using separate coding networks to encode the scene image and photon statistical map separately, reducing parameter redundancy in the image processing model and improving the model's processing efficiency. Attached Figure Description

[0011] Figure 1 This is a schematic flowchart of an image processing method provided in an exemplary embodiment of this disclosure; Figure 2 This is a schematic diagram of the structure of an image processing model provided in an exemplary embodiment of the present disclosure; Figure 3 This is a schematic flowchart of an image processing method provided in another exemplary embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of an image processing model provided in another exemplary embodiment of this disclosure; Figure 5 This is a schematic flowchart of an image processing method provided in yet another exemplary embodiment of this disclosure; Figure 6 This is a schematic flowchart of an image processing method provided in yet another exemplary embodiment of this disclosure; Figure 7 This is a schematic diagram of the process of training an image processing model provided in an exemplary embodiment of this disclosure; Figure 8 This is a schematic diagram of the process of training an image processing model provided in another exemplary embodiment of this disclosure; Figure 9 This is a schematic flowchart of training an image processing model provided in yet another exemplary embodiment of this disclosure; Figure 10 This is a schematic flowchart of training an image processing model provided in yet another exemplary embodiment of this disclosure; Figure 11This is a structural diagram of an image processing apparatus provided in an exemplary embodiment of the present disclosure; Figure 12 This is a structural diagram of an image processing apparatus provided in another exemplary embodiment of the present disclosure; Figure 13 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0012] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.

[0013] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0014] Application Overview In related technologies, TOF data processing and image processing are usually implemented using separate neural networks. This approach has the following technical problems: Waste of computational resources: Independent networks need to perform forward inference separately, which increases computational overhead and memory usage; Intermodal information is not fully utilized: TOF data and image data contain complementary information, that is, both TOF data and image data contain two-dimensional and three-dimensional features. Independent networks cannot extract two-dimensional and three-dimensional features from TOF data or image data at the same time, and cannot fully utilize the complementarity of two-dimensional and three-dimensional features, resulting in insufficient accuracy of model processing. Poor real-time performance: On resource-constrained edge devices, it is difficult to meet real-time requirements by processing two types of data separately; Parameter redundancy: The two independent networks have a large number of duplicate parameters and computational structures, which also increases computational overhead and memory usage.

[0015] To address the aforementioned issues, embodiments of this disclosure employ a shared coding network to encode scene images and photon statistical maps. This shared coding network can extract both two-dimensional and three-dimensional features from the input data. The features output by the shared coding network reflect the complementarity between the two-dimensional and three-dimensional features. Utilizing this complementarity, the photon statistical features and image features are decoded more accurately, improving the accuracy of the depth map after model decoding and the quality of the denoised image. Furthermore, using a shared coding network avoids the need to encode scene images and photon statistical maps using separate coding networks, reducing parameter redundancy in the image processing model and improving its processing efficiency.

[0016] Exemplary methods Figure 1 This is a schematic flowchart illustrating an exemplary embodiment of the image processing method provided in this disclosure. This embodiment can be applied to various types of electronic devices, such as vehicle controllers, personal computers, and backend servers. Figure 1 As shown, it includes the following steps 101-103.

[0017] Step 101: Obtain the scene image and photon statistics map corresponding to the target scene.

[0018] The target scene can be any type of scene, such as a vehicle driving scene. The scene image can be a two-dimensional image captured by a camera in the target scene. The photon statistics map can be obtained by a lidar system such as dToF (direct time-of-flight). The photon statistics map reflects the photon flight time corresponding to each pixel in the two-dimensional image mapped from the target scene. The photon flight time is the time it takes for a photon to be emitted from the laser emitter, reflected by an object in the target scene, and returned to the photoelectric sensor.

[0019] The dToF lidar system operates as follows: A laser emitter emits an extremely short light pulse. A single-photon avalanche diode (SPAD) sensor detects photons reflected back from various objects in the scene. For each detected photon, a time-to-digital converter (TDC) precisely records the time between its emission and reception. The scene is mapped onto a two-dimensional image, and the number of received photons is counted for each pixel at each scanned location, thus obtaining a photon statistics map.

[0020] In this context, for any pixel in the photon statistics map, the pixel corresponds to a photon statistics data sequence. The photon statistics data in the photon statistics data sequence are used to characterize the number of reflected photons received from the target location at the corresponding time. The target location is the position in the target scene corresponding to that pixel.

[0021] Typically, for any pixel in the photon statistics map to be denoised, the corresponding data sequence can form a time-photon count histogram. For example, the histogram corresponding to one pixel has 512 time points, covering a measurement range of 0 to 100 nanoseconds (corresponding to a distance of 0 to 15 meters). After repeatedly counting the detected photons, the data in the resulting histogram is as follows: Time points: [1, 2, 3, 4, 5, ..., 510, 511, 512]; Photon count: [2, 5, 21, 105, 203, 450, 301, 98, ..., 1, 0, 3].

[0022] The peak value of the histogram corresponds to the photon flight time at the location of that pixel.

[0023] In some alternative implementations, step 101 can also be executed by an NPU (neural processing unit), or by the processor calling the corresponding instructions stored in memory, or by the acquisition module 1101 run by the processor.

[0024] Step 102: Using the shared coding network included in the pre-trained image processing model, the scene image and photon statistical map are encoded respectively to obtain image features and photon statistical features.

[0025] In this embodiment, the image processing model can process multimodal image data. That is, scene images and photon statistical maps can be simultaneously input into the image processing model, which then processes both. In some implementation examples, the image processing model can be a neural network model with a Unet structure, which encodes and then decodes the input image data to obtain the processed image data.

[0026] like Figure 2 As shown, the image processing model can include a shared coding network, a deep decoder, and an image decoder. The parameters of the image processing model can be obtained through pre-training. The shared coding network can use the same parameters to encode the scene image and the photon statistical map separately. The shared coding network can include at least one convolutional layer, at least one pooling layer, etc., to extract features layer by layer from the input scene image and photon statistical map, and output image features and photon statistical features.

[0027] In some alternative implementations, step 102 can also be executed by the NPU, or by the processor calling the corresponding instructions stored in memory, or by the encoding module 1102 run by the processor.

[0028] Step 103: Use the depth decoder included in the image processing model to decode the photon statistical features, and use the decoded photon statistical features to generate a depth map corresponding to the target scene.

[0029] The depth decoder may include at least one convolutional layer, an upsampling layer, and other structures. The depth decoder decodes photon statistical features layer by layer. The resolution of the decoded photon statistical features output by the last layer can be restored to the same level as the input photon statistical map. Then, the data corresponding to each pixel in the decoded photon statistical features can be mapped to a depth value, resulting in a depth map. Each pixel in the depth map corresponds to a depth value, which represents the distance between the target and the sensor.

[0030] In some alternative implementations, step 103 can also be executed by the NPU, or by the processor calling the corresponding instructions stored in memory, or by the first decoding module 1103 run by the processor.

[0031] Step 104: Use the image decoder included in the image processing model to decode the image features and obtain the denoised image.

[0032] The image decoder may include at least one convolutional layer, an upsampling layer, or other structures. The image decoder decodes photon statistical features layer by layer. The resolution of the decoded image features output by the last layer can be restored to the same as the input scene image. Then, image reconstruction can be performed on each pixel contained in the decoded image features to obtain the denoised image.

[0033] In some alternative implementations, step 104 can also be executed by the NPU, or by the processor calling the corresponding instructions stored in memory, or by the second decoding module 1104 run by the processor.

[0034] The image processing method provided in this disclosure acquires a scene image and a photon statistical map corresponding to a target scene. It then uses a shared coding network included in a pre-trained image processing model to encode the scene image and photon statistical map separately, obtaining image features and photon statistical features. A depth decoder is used to decode the photon statistical features, and the decoded photon statistical features are used to generate a depth map corresponding to the target scene. Finally, an image decoder is used to decode the image features, obtaining a denoised image. Because a shared coding network is used to encode the scene image and photon statistical map, this network can extract both two-dimensional and three-dimensional features from the input data. The features output by the shared coding network can reflect the complementarity between two-dimensional and three-dimensional features. Utilizing this complementarity, the photon statistical features and image features are decoded more accurately, improving the accuracy of the decoded depth map and the quality of the denoised image. Furthermore, using a shared coding network avoids using separate coding networks for encoding the scene image and photon statistical map, reducing parameter redundancy in the image processing model and improving its processing efficiency.

[0035] In some alternative implementations, such as Figure 3 As shown, step 102 may include steps 1021-1023.

[0036] Step 1021: Project the photon statistical map using the projection units included in the shared coding network to obtain projection data.

[0037] In this embodiment, as Figure 4 As shown, the shared coding network includes a projection unit, an image precoding unit, and a shared encoder. The projection unit is used to transform the dimensions of the photon statistical graph so that the dimensions of the transformed projected data are the same as the dimensions of the image precoding data. For example, if the initial tensor size of the photon statistical graph is 96×240×256, after projection, the tensor size of the resulting projected data is 96×240×128. The formula used by the projection unit can be shown in the following equation (1): (1) in, and These are the parameters of the projection unit. This is the activation function.

[0038] Step 1022: Use the image precoding unit included in the shared coding network to precode the scene image to obtain image precoded data.

[0039] In this model, the projection data and the image precoding data have the same dimension. The image precoding unit can encode the input scene image, and the encoding process can be implemented using methods such as multiple convolutions and pooling.

[0040] For example, the input scene image size is 1080×1920×3; Convolutional layer 1 structure: 3×3 convolutional kernel, 64 channels, ReLU activation function; Convolutional layer 2 structure: 3×3 convolutional kernel, 64 channels, ReLU activation function; Max pooling structure: 2×2 pooling kernel, output tensor size 540×960×64; Convolutional layer 3 structure: 3×3 convolutional kernel, 128 channels, ReLU activation function; Convolutional layer 4 structure: 3×3 convolutional kernel, 128 channels, ReLU activation function; Adaptive pooling: output tensor size 96×240×128. The dimensions of the output image pre-coded data are the same as the projected data in the above example.

[0041] Step 1023: Using the shared encoder included in the shared coding network, the projection data and image precoding data are encoded respectively to obtain photon statistical features and image features.

[0042] The shared encoder can include multiple stacked shared sub-networks. The input projection data and image pre-encoded data are encoded by the multiple shared sub-networks to obtain downsampled photon statistical features and image features.

[0043] For example, in some implementation examples, the hierarchical configuration of the shared encoder is as follows: Layer 1: 128 input channels, 128 output channels, downsampled to 48×120; Layer 2: 128 input channels, 256 output channels, downsampled to 24×60; Layer 3: 256 input channels, 512 output channels, downsampled to 12×30; Layer 4: 512 input channels, 1024 output channels, downsampled to 6×15.

[0044] Each shared subnetwork layer can contain at least one convolutional layer and a downsampling layer. After each convolution, the resulting data is processed by the ReLU activation function and then input into the lower layer structure for computation. Downsampling can be performed using a 2×2 pooling kernel.

[0045] In some implementation examples, for each shared sub-network layer, the input data of the two modalities can be merged first, and the shared sub-network layer encodes the photon statistical features and image features. The two types of encoded data are then separated, and the separated encoded data is input into the next shared sub-network layer. The same steps are repeated until the last layer outputs the photon statistical features and image features.

[0046] In some alternative implementations, steps 1021-1023 may also be executed by the first conversion unit 11021, the second conversion unit 11022, and the encoding unit 11023 respectively, which are run by the processor.

[0047] This embodiment achieves the transformation of scene images and photon statistical maps to the same dimension before encoding data of two modalities using the shared encoder by setting up a projection unit, an image pre-coding unit, and a shared encoder in the shared coding network. At the same time, preliminary features are extracted from the scene images, so that the subsequent shared encoder can encode multimodal data more accurately, thereby helping to improve the accuracy of subsequent decoding.

[0048] In some alternative implementations, step 1023 may include: By using a shared encoder, the projection data is encoded to obtain photon statistical features containing two-dimensional planar features and three-dimensional depth features; by using a shared encoder, the image pre-encoded data is encoded to obtain image features containing two-dimensional planar features and three-dimensional depth features.

[0049] In this embodiment, since the shared encoder is pre-trained based on sample photon statistical maps and sample scene images, and the sample photon statistical maps contain more three-dimensional features while the sample scene images contain more two-dimensional planar features, after supervised learning by two-dimensional planar annotation information and three-dimensional annotation information, the trained shared encoder can simultaneously extract two-dimensional and three-dimensional features from the input data. This allows the obtained photon statistical features and image features to reflect the complementarity of two-dimensional and three-dimensional features, respectively. The photon statistical features and image features can more accurately represent the positions of objects in the actual scene, thereby improving the accuracy of subsequent decoding.

[0050] In some alternative implementations, the shared encoder comprises at least two layers of shared subnetworks arranged sequentially. Accordingly, step 1023 can be implemented as follows: For any one of the at least two shared subnetworks, the data input to that shared subnetwork is encoded and downsampled using the shared parameters of that shared subnetwork to obtain downsampled image features and photon statistical features, and the encoded image features and photon statistical features before downsampling are backed up.

[0051] In this system, the image features and photon statistical features output from each shared sub-network can be input into the next shared sub-network. The image features and photon statistical features output from the last shared sub-network are input into the image decoder and depth decoder, respectively. The backup image features and photon statistical features are used to perform skip connections (i.e., skip connections between the decoder and encoder of the Unet network) with the corresponding features during decoding.

[0052] As an example, each shared sub-network layer can include two convolutional layers. After the first convolution, the output features are processed by an activation function, followed by a second convolution and another activation function calculation to obtain the image features and photon statistical features output by that shared sub-network layer, denoted as... and The superscript 'l' indicates the layer number of the current shared subnetwork. and Backup as , Then, according to the pooling kernel of size 2×2. and Downsampling calculations are performed separately to obtain the image features and photon statistical features output by the shared sub-network of this layer, denoted as follows: and .

[0053] This embodiment utilizes at least two shared sub-networks, each layer of which sequentially encodes and downsamples the two different modalities of input data. Simultaneously, it backs up the encoded image features and photon statistical features. This allows for multi-level encoding and dimensional transformation of projection data and image pre-coded data using the shared parameters of the shared sub-networks. During decoding, the backed-up image features and photon statistical features can be used as a reference, helping to avoid the loss of spatial details in image features and photon statistical features and improving the accuracy of the decoded data.

[0054] In some alternative implementations, the depth decoder includes at least two depth decoding sub-networks. Accordingly, such as... Figure 5 As shown, based on the above embodiments, step 103 may include steps 1031-1033.

[0055] Step 1031: For any one of the at least two deep decoding subnetworks, upsample the photon statistical features input to that deep decoding subnetwork, and merge the upsampled photon statistical features with the backup photon statistical features corresponding to that deep decoding subnetwork; decode the merged photon statistical features to obtain the decoded photon statistical features.

[0056] Among them, at least two deep decoding subnetworks can adopt a stacked structure, that is, the data output by the upper subnetwork is processed by the lower subnetwork.

[0057] For example, in some implementation examples, the depth decoder is configured with the following hierarchy: Layer 4: 1024 input channels, 512 output channels, upsampled to 12×30; Layer 3: 512 input channels, 256 output channels, upsampled to 24×60; Layer 2: 256 input channels, 128 output channels, upsampled to 48×120; Layer 1: 128 input channels, 64 output channels, upsampled to 96×240.

[0058] The layer numbers correspond to the layer numbers of the shared encoder mentioned above. For example, the 4th layer of the depth decoder corresponds to the 4th layer of the shared encoder.

[0059] Each deep decoding subnetwork can include an upsampling layer, a skip connection layer, and at least one convolutional layer. After each upsampling, the obtained data is merged with the corresponding backup photon statistical features (i.e., skip connection). The merged photon statistical features are then processed by at least one convolution and ReLU activation function to obtain the decoded photon statistical features. The decoded photon statistical features are then input into the lower layer structure for calculation.

[0060] The expression for merging the above-upsampled photon statistical features with the corresponding backup photon statistical features can be shown in the following equation (2): (2) in, The statistical characteristics of photons after upsampling. To back up photon statistical characteristics.

[0061] As an example, each deep decoding subnetwork may include two convolutional layers. After the first convolution, the output features are calculated by an activation function, then a second convolution is performed, and after another activation function calculation, the photon statistical features output by that deep decoding subnetwork are obtained.

[0062] Step 1032: If the depth decoding subnetwork is not the last layer, input the decoded photon statistical features into the next depth decoding subnetwork.

[0063] Step 1033: If the depth decoding subnetwork is the last layer, perform depth calculation on the decoded photon statistical features to obtain a depth map.

[0064] The resolution of the photon statistical features output by the final layer of the deep decoding subnetwork is the same as the resolution of the initial photon statistical map. During depth calculation, the data set corresponding to each pixel in the photon statistical features of the photon statistical map can be mapped to a depth value. The specific calculation formulas are shown in equations (3) and (4) below: (3) (4) in, The photon statistical features output by the last deep decoding sub-network. and These are the parameters obtained during training. It is the preset maximum measurement distance. This is the sigma activation function.

[0065] This embodiment sets up multiple depth decoding subnetworks in the depth decoder, and merges the corresponding backup photon statistical features with each depth decoding subnetwork before decoding. This allows the backup photon statistical features to be used as a reference in the decoding process, which helps to avoid the loss of spatial details in the photon statistical features and improves the accuracy of depth calculation after decoding.

[0066] In some alternative implementations, the image decoder includes at least two layers of image decoding sub-networks. Accordingly, such as... Figure 6 As shown, based on the above embodiments, step 104 may include steps 1041-1043.

[0067] Step 1041: For any one of the at least two image decoding subnetworks, upsample the image features input to that image decoding subnetwork and merge the upsampled image features with the backup image features corresponding to that image decoding subnetwork; decode the merged image features to obtain the decoded image features.

[0068] At least two layers of image decoding subnetworks can be stacked sequentially, meaning that the data output from the upper subnetwork is processed by the lower subnetwork.

[0069] Each image decoding subnetwork layer may include an upsampling layer, a skip connection layer, and at least one convolutional layer. After each upsampling, the obtained data is merged with the corresponding backup image features (i.e., skip connection). The merged image features are then processed by at least one convolution and ReLU activation function to obtain the decoded image features. The decoded image features are then input into the lower layer structure for computation.

[0070] The expression for merging the upsampled image features with the corresponding backup image features can be shown in the following equation (5): (5) in, The statistical characteristics of photons after upsampling. To back up photon statistical characteristics.

[0071] Step 1042: If the image decoding subnetwork of this layer is not the last layer, input the decoded image features into the next layer of the image decoding subnetwork.

[0072] Step 1043: If the image decoding subnetwork of this layer is the last layer, perform image reconstruction on the decoded image features to obtain the denoised image.

[0073] The resolution of the image features output by the final image decoding sub-network is the same as the resolution of the original scene image. The calculation formulas used for image reconstruction are shown in equations (6) and (7) below: (6) (7) in, The image features output by the last layer of the image decoding subnetwork, and These are the parameters obtained during training. This is the sigma activation function.

[0074] This embodiment sets up a multi-layer image decoding sub-network in the depth decoder, and merges it with the corresponding backup image features before decoding each layer. This allows the backup image features to be used as a reference in the decoding process, which helps to avoid the loss of spatial details in image features and improves the accuracy of image noise reduction.

[0075] In some alternative implementations, such as Figure 7 As shown, the image processing model in any of the above embodiments can be trained in advance according to the following steps 701-705.

[0076] Step 701: Obtain the sample scene image and the corresponding sample denoised image, as well as the sample photon statistics map and the corresponding sample depth map.

[0077] Among them, the sample denoised image is the image after denoising the sample scene image, which serves as the benchmark for image denoising during training. The sample depth map is the depth map obtained by measuring the distance to the spatial position indicated by each pixel in the sample photon statistical map, which serves as the benchmark for calculating the depth value during training.

[0078] Step 702: Input the sample scene image and sample photon statistical map into the initial image processing model to obtain the predicted depth map and the predicted denoised image.

[0079] In this embodiment of the disclosure, the initial image processing model is an image processing model to be trained, which may be an image processing model that has not yet started training or is in the process of training.

[0080] The structure of the initial image processing model and the process of performing image processing can be referred to the above embodiments, and will not be repeated here. The predicted depth map and the predicted denoised image are the depth map and the denoised image output by the initial image processing model, respectively.

[0081] Step 703: Based on a preset depth loss function, determine the depth loss between the sample depth map and the predicted depth map, and based on a preset image denoising loss function, determine the denoising loss between the sample denoised image and the predicted denoised image.

[0082] The depth loss function and the image denoising loss function can be constructed using various types of loss functions. For example, the depth loss function can be constructed using the L1 loss function, and the image denoising loss function can be constructed using both the L1 loss function and the perceptual loss function.

[0083] Step 704: Based on depth loss and noise reduction loss, train the initial image processing model, including the shared coding network, depth decoder, and image decoder.

[0084] In some optional examples, the training objective can be to minimize the depth loss and noise reduction loss, using backpropagation and gradient descent to iteratively adjust the parameters of the modules included in the initial image processing model.

[0085] Step 705: In response to the initial image processing model meeting the preset training termination condition, the current initial image processing model is determined as the image processing model that has been trained.

[0086] In this embodiment of the disclosure, steps 701-705 can be executed iteratively to train the initial image processing model until the preset training termination condition is met.

[0087] In some optional examples, the training termination condition may include, but is not limited to, at least one of the following: convergence of depth loss and noise reduction loss, reaching a preset number of training iterations, reaching a preset training duration, etc.

[0088] This embodiment trains the image processing model using multimodal sample data, enabling the trained model to perform more accurate depth map prediction and image denoising for real-world application scenarios.

[0089] In some alternative implementations, such as Figure 8 As shown, in Figure 7Based on the illustrated embodiment, step 704 may include steps 7041-7042.

[0090] Step 7041: Following the phased training strategy, in the first phase, a shared coding network is trained based on depth loss and noise reduction loss.

[0091] In some optional examples, the initial image processing model can be trained using photon statistics and image data separately at each training iteration. For example, the parameters of the depth decoder and image decoder can be fixed, a sample photon statistics map can be input into the initial image processing model, the predicted depth map output by the initial image processing model can be obtained, the depth loss representing the error between the predicted depth map and the sample depth map can be calculated, and the parameters of the shared coding network can be adjusted to gradually reduce the depth loss; and, a sample scene image can be input, a predicted denoised image can be output, the denoising loss representing the error between the predicted denoised image and the sample denoised image can be calculated, and the parameters of the shared coding network can be adjusted to gradually reduce the denoising loss.

[0092] It is understandable that if the shared coding network includes a projection unit, an image precoding unit, and a shared encoder, then when training the shared coding network using sample photon statistical maps, it is not necessary to adjust the parameters of the image precoding unit; and when training the shared coding network using sample scene images, it is not necessary to adjust the parameters of the projection unit.

[0093] Step 7042, in the second stage, redetermine the depth loss and denoising loss, and train the depth decoder and image decoder based on the redetermined depth loss and denoising loss.

[0094] In the second stage, the parameters of the trained shared coding network can be fixed, and the depth decoder and image decoder can be trained separately using photon statistics and image data. For example, a sample photon statistics map can be input, a predicted depth map can be output, the depth loss representing the error between the predicted depth map and the sample depth map can be calculated, and the parameters of the depth decoder can be adjusted to gradually reduce the depth loss; similarly, a sample scene image can be input, a predicted denoised image can be output, the denoising loss representing the error between the predicted denoised image and the sample denoised image can be calculated, and the parameters of the image decoder can be adjusted to gradually reduce the denoising loss.

[0095] This embodiment employs a phased training model, which facilitates targeted training of each module within the model, improving training efficiency and data processing accuracy. Furthermore, in the first training phase, training the same shared coding network using data from two different modalities allows the network to learn the ability to extract 3D features from sample photon statistical maps and 2D features from sample scene images. The shared parameters within the network enable simultaneous extraction of both 2D and 3D features from a single modality of data, achieving complementarity between the two. The photon statistical features and image features output by the shared coding network more accurately represent the shape and position of objects in actual 3D space, improving the accuracy of depth estimation and image denoising.

[0096] In some alternative implementations, Figure 8 Based on the illustrated embodiment, step 7041 above can be implemented in the following manner: First, with the goal of minimizing the depth loss, the projection units included in the shared encoding network are trained.

[0097] Then, with the goal of minimizing the noise reduction loss, the image precoding units included in the shared coding network are trained.

[0098] Finally, with the goal of minimizing the weighted sum of depth loss and noise reduction loss, the shared encoder of the shared coding network is trained.

[0099] The shared coding network in this embodiment includes a projection unit, an image precoding unit, and a shared encoder. The parameters of the projection unit and the image precoding unit are independent, while the parameters of the shared encoder can be shared during the extraction of photon statistical features and image features. That is, the shared encoder can encode both the projection data and the image precoding data.

[0100] During the training of the projection unit, its parameters can be iteratively adjusted to gradually reduce the depth loss. Similarly, during the training of the image precoding unit, its parameters can be iteratively adjusted to gradually reduce the noise reduction loss. It should be understood that during the training of the projection unit, there is no need to input sample scene images into the image precoding unit, and during the training of the image precoding unit, there is no need to input sample photon statistics into the projection unit.

[0101] The weighted sum of the depth loss and noise reduction loss can be calculated as shown in equation (8): (8) in, This represents the total loss value. For deep loss, To reduce noise loss, , To preset weights, , The values ​​of are all real numbers greater than 0.

[0102] When training the shared encoder, it is possible to use sample photon statistics and sample scene images for training, that is, by calculating the total loss value mentioned above, the parameters of the shared encoder are iteratively adjusted so that the total loss value gradually decreases.

[0103] This embodiment uses sample data to train the projection unit, image precoding unit, and shared encoder separately, realizing staged training of each unit contained in the shared encoder network, which helps to simplify the training process and improve training efficiency.

[0104] In some alternative implementations, Figure 8 Based on the illustrated embodiment, step 7042 above can be implemented in the following manner: The depth decoder and image decoder are trained with the goal of minimizing the weighted sum of the redetermined depth loss and denoising loss.

[0105] In this embodiment, the weighted summation calculation method for depth loss and noise reduction loss can refer to the above equation (8). When training the depth decoder and image decoder, the sample photon statistical map and sample scene image can be input, and the predicted depth map and the predicted noise-reduced image can be output. The total loss value shown in equation (8) can be calculated, and then the parameters of the depth decoder and image decoder can be adjusted to minimize the total loss value until the training termination condition is met.

[0106] This embodiment trains the depth decoder and image decoder by weighted summation of depth loss and denoising loss. This allows the training process to reflect the correlation between multimodal data and further enables the decoder parameters to complement the two-dimensional and three-dimensional features between photon statistical features and image features, thereby improving the accuracy of depth estimation and image denoising.

[0107] In some alternative implementations, such as Figure 9 As shown, step 703 above may include steps 7031-7033.

[0108] Step 7031: Determine the signal-to-noise ratio of the photon statistical data sequence corresponding to each pixel in the sample photon statistical map.

[0109] The signal-to-noise ratio can be obtained by determining the peak value from the photon statistics sequence corresponding to each pixel and the ratio between the peak value and the basic noise intensity (e.g., the mean of other data besides the peak data).

[0110] Step 7032: Based on the signal-to-noise ratio, determine the loss weight corresponding to each pixel in the sample photon statistical map.

[0111] Specifically, the loss weight for each pixel can be calculated using the following formula (9): (9) in, This represents the signal-to-noise ratio corresponding to the pixel with coordinates (i,j). This represents the sum of the signal-to-noise ratios of all pixels in the sample photon statistical image. , The height and width of the sample photon statistics graph.

[0112] Step 7033: Use the loss weights as parameters of the depth loss function to determine the depth loss between the sample depth map and the predicted depth map.

[0113] In some optional examples, the depth loss function can be constructed using a weighted L1 loss function. For example, the depth loss function is shown in equation (10) below: (10) in, To predict the depth value of pixel (i,j) in the depth map, This corresponds to the annotation depth value.

[0114] This embodiment calculates the loss weight corresponding to each pixel in the sample photon statistical map, and uses the loss weight to calculate the depth loss. This enables the depth loss to fully reflect the regions with high signal quality in the sample photon statistical map. By improving the accuracy of the depth loss based on the signal quality, the accuracy of model training is improved.

[0115] In some alternative implementations, such as Figure 10 As shown, step 703 above may include steps 7034-7036.

[0116] Step 7034: Based on the absolute error loss function included in the image denoising loss function, determine the absolute error loss between the sample denoised image and the predicted denoised image.

[0117] The absolute error loss function is the L1 loss function, and its calculation formula is shown in equation (11) below: (11) in, , , The height, width, and number of channels of the sample scene image. To predict the denoised image, The image after denoising the sample.

[0118] Step 7035: Based on the perceptual loss function included in the image denoising loss function, determine the perceptual loss between the sample denoised image and the predicted denoised image.

[0119] The formula for calculating perceived loss is shown in equation (12) below: (12) in, , , The feature map height, width, and number of channels of the l-th layer network are given. This represents the feature extraction function of the l-th layer.

[0120] Step 7036: Weighted summation of absolute error loss and perception loss to obtain noise reduction loss.

[0121] The weighted summation formula can be shown in equation (13) below: (13) in, For weights, for example, Accordingly, in the above equation (8) , .

[0122] This embodiment calculates the denoising loss by setting absolute error loss and perceptual loss for image data, thereby characterizing the error between the predicted denoised image and the sample denoised image from multiple dimensions, resulting in higher accuracy of the denoising loss.

[0123] In some alternative implementations, after step 7042, step 704 may further include: In the third stage, the depth loss and denoising loss are redefined, and the learning rate is reset. Based on the depth loss and denoising loss redefined in the third stage, and the reset learning rate, the shared encoder network, the depth decoder, and the image decoder are jointly trained.

[0124] Jointly training the shared encoder network, depth decoder, and image decoder refers to simultaneously adjusting the parameters of these three components using the input sample photon statistical map and sample scene image. The learning rate, a hyperparameter in deep learning, controls the step size of parameter updates during model training. In this embodiment, after the first and second stages, a third stage can be initiated, setting the learning rate to a smaller value to fine-tune the model and further improve the accuracy of depth estimation and image denoising. The implementation methods for depth loss and denoising loss, as well as the specific process of training the shared encoder network, depth decoder, and image decoder, can be found in the above embodiments and will not be repeated here.

[0125] Exemplary device Figure 11 This is a schematic diagram of the structure of an image processing apparatus provided in an exemplary embodiment of this disclosure. The image processing apparatus provided in this embodiment can be used to implement the image processing method of this embodiment. This embodiment can be applied to electronic devices, such as... Figure 11 As shown, the image processing device includes: an acquisition module 1101, an encoding module 1102, a first decoding module 1103, and a second decoding module 1104. The acquisition module 1101 is used to acquire a scene image and a photon statistical map corresponding to the target scene. For any pixel in the photon statistical map, the pixel corresponds to a photon statistical data sequence. The photon statistical data in the photon statistical data sequence is used to characterize the number of reflected photons received from the target location at the corresponding time. The target location is the position in the target scene corresponding to that pixel. The encoding module 1102 is used to encode the scene image and the photon statistical map using a shared encoding network included in a pre-trained image processing model, respectively, to obtain image features and photon statistical features. The first decoding module 1103 is used to decode the photon statistical features using a depth decoder included in the image processing model, and uses the decoded photon statistical features to generate a depth map corresponding to the target scene. The second decoding module 1104 is used to decode the image features using an image decoder included in the image processing model, to obtain a denoised image.

[0126] Reference Figure 12 , Figure 12 This is a schematic diagram of the structure of an image processing apparatus provided in another exemplary embodiment of the present disclosure.

[0127] In some optional implementations, the encoding module 1102 may include: a first conversion unit 11021, a second conversion unit 11022, and an encoding unit 11023. The first conversion unit 11021 is used to project the photon statistical map using the projection unit included in the shared encoding network to obtain projection data; the second conversion unit 11022 is used to pre-encode the scene image using the image pre-coding unit included in the shared encoding network to obtain image pre-coded data, wherein the projection data and the image pre-coded data have the same dimension; the encoding unit 11023 is used to encode the projection data and the image pre-coded data respectively using the shared encoder included in the shared encoding network to obtain photon statistical features and image features.

[0128] In some optional implementations, the encoding unit 11023 may include: a first encoding subunit 110231 and a second encoding subunit 110232. The first encoding subunit 110231 is used to encode the projection data using a shared encoder to obtain photon statistical features containing two-dimensional planar features and three-dimensional depth features; the second encoding subunit 110232 is used to encode the image pre-encoded data using a shared encoder to obtain image features containing two-dimensional planar features and three-dimensional depth features.

[0129] In some alternative implementations, the shared encoder includes at least two shared subnetworks arranged sequentially. Accordingly, the encoding unit 11023 may be further configured to: for any one of the at least two shared subnetworks, use the shared parameters of that shared subnetwork to encode and downsample the data input to that subnetwork to obtain downsampled image features and photon statistical features, and back up the encoded image features and photon statistical features before downsampling.

[0130] In some optional implementations, the depth decoder includes at least two depth decoding subnetworks. Accordingly, the first decoding module 1103 may be further configured to: for any one of the at least two depth decoding subnetworks, upsample the photon statistical features input to that depth decoding subnetwork, and merge the upsampled photon statistical features with the corresponding backup photon statistical features of that depth decoding subnetwork; decode the merged photon statistical features to obtain decoded photon statistical features; if that depth decoding subnetwork is not the last layer, input the decoded photon statistical features into the next depth decoding subnetwork; if that depth decoding subnetwork is the last layer, perform depth calculation on the decoded photon statistical features to obtain a depth map.

[0131] In some optional implementations, the image decoder includes at least two layers of image decoding subnetworks. Accordingly, the second decoding module 1104 may be further configured to: for any one of the at least two layers of image decoding subnetworks, upsample the image features input to that layer of image decoding subnetwork, and merge the upsampled image features with the corresponding backup image features of that layer of image decoding subnetwork; decode the merged image features to obtain the decoded image features; if that layer of image decoding subnetwork is not the last layer, input the decoded image features into the next layer of image decoding subnetwork; if that layer of image decoding subnetwork is the last layer, reconstruct the image from the decoded image features to obtain the denoised image.

[0132] In some optional implementations, the image processing model can be pre-trained according to the following steps: acquiring sample scene images and corresponding denoised sample images, as well as sample photon statistical maps and corresponding sample depth maps; inputting the sample scene images and sample photon statistical maps into the initial image processing model to obtain predicted depth maps and predicted denoised images; determining the depth loss between the sample depth maps and predicted depth maps based on a preset depth loss function, and determining the denoising loss between the sample denoised images and predicted denoised images based on a preset image denoising loss function; training the shared coding network, depth decoder, and image decoder included in the initial image processing model based on the depth loss and denoising loss; and determining the current initial image processing model as the trained image processing model in response to the initial image processing model meeting the preset training termination condition.

[0133] In some alternative implementations, training the initial image processing model, including the shared coding network, depth decoder, and image decoder, based on depth loss and denoising loss, may include: in the first stage, training the shared coding network based on depth loss and denoising loss according to a phased training strategy; and in the second stage, redetermining the depth loss and denoising loss, and training the depth decoder and image decoder based on the redetermined depth loss and denoising loss.

[0134] In some alternative implementations, training the shared coding network based on depth loss and denoising loss may include: training the projection units included in the shared coding network with the goal of minimizing depth loss; training the image precoding units included in the shared coding network with the goal of minimizing denoising loss; or training the shared encoder included in the shared coding network with the goal of minimizing the weighted sum of depth loss and denoising loss.

[0135] In some alternative implementations, the depth decoder and image decoder are trained based on the redefined depth loss and denoising loss, including: training the depth decoder and image decoder with the objective of minimizing the weighted sum of the redefined depth loss and denoising loss.

[0136] In some optional implementations, the depth loss between the sample depth map and the predicted depth map is determined based on a preset depth loss function, including: determining the signal-to-noise ratio of the photon statistical data sequence corresponding to each pixel in the sample photon statistical map; determining the loss weight corresponding to each pixel in the sample photon statistical map based on the signal-to-noise ratio; and using the loss weight as a parameter of the depth loss function to determine the depth loss between the sample depth map and the predicted depth map.

[0137] In some optional implementations, determining the denoising loss between the sample denoised image and the predicted denoised image based on a preset image denoising loss function may include: determining the absolute error loss between the sample denoised image and the predicted denoised image based on the absolute error loss function included in the image denoising loss function; determining the perceptual loss between the sample denoised image and the predicted denoised image based on the perceptual loss function included in the image denoising loss function; and obtaining the denoising loss by weighted summation of the absolute error loss and the perceptual loss.

[0138] In some alternative implementations, after redetermining the depth loss and denoising loss in the second stage, and training the depth decoder and image decoder based on the redetermined depth loss and denoising loss, the apparatus is further configured to: redetermine the depth loss and denoising loss in the third stage, and reset the learning rate; and jointly train the shared coding network, the depth decoder, and the image decoder based on the depth loss and denoising loss redetermined in the third stage.

[0139] The exemplary embodiments of this device correspond to the exemplary method section described above in terms of implementation. The corresponding content between the two can be referenced, combined, and cited, and will not be repeated here. The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects of the exemplary method section described above, and will not be repeated here.

[0140] Exemplary electronic devices Figure 13 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 1301 and a memory 1302.

[0141] The processor 1301 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0142] The memory 1302 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1301 may execute one or more computer program instructions to implement the image processing methods and / or other desired functions of the various embodiments of this disclosure described above.

[0143] In one example, the electronic device may also include an input device 1303 and an output device 1304, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0144] The input device 1303 may include, for example, a keyboard, a mouse, etc.

[0145] The output device 1304 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0146] Of course, for the sake of simplicity, Figure 13 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.

[0147] Exemplary computer program products and computer-readable storage media In addition to the methods and apparatus described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the image processing methods of the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0148] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0149] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the image processing methods of the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0150] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0151] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0152] Various modifications and variations can be made to this disclosure without departing from its spirit and scope. Therefore, this disclosure is also intended to include such modifications and variations if they fall within the scope of the claims of this disclosure and their equivalents.

Claims

1. An image processing method, comprising: Obtain the scene image and photon statistics map corresponding to the target scene. For any pixel in the photon statistics map, the pixel corresponds to a photon statistics data sequence. The photon statistics data in the photon statistics data sequence is used to characterize the number of reflected photons received from the target location at the corresponding time. The target location is the position in the target scene corresponding to the pixel. Using a pre-trained image processing model that includes a shared coding network, the scene image and the photon statistical map are encoded respectively to obtain image features and photon statistical features; The image processing model includes a depth decoder to decode the photon statistical features, and the decoded photon statistical features are used to generate a depth map corresponding to the target scene. The image features are decoded using the image decoder included in the image processing model to obtain a denoised image.

2. The method according to claim 1, wherein, The pre-trained image processing model, including a shared coding network, encodes the scene image and the photon statistical map respectively to obtain image features and photon statistical features, including: The photon statistical map is projected using the projection units included in the shared coding network to obtain projection data; The scene image is pre-coded using the image precoding unit included in the shared coding network to obtain image pre-coded data, wherein the projection data and the image pre-coded data have the same dimension; The projection data and the image pre-coded data are encoded using the shared encoder included in the shared coding network to obtain the photon statistical features and the image features, respectively.

3. The method according to claim 2, wherein, The step of using the shared encoder included in the shared coding network to encode the projection data and the image pre-coded data respectively to obtain the photon statistical features and the image features includes: The shared encoder is used to encode the projection data to obtain the photon statistical features that include two-dimensional planar features and three-dimensional depth features; The image precoding data is encoded using the shared encoder to obtain the image features containing two-dimensional planar features and three-dimensional depth features.

4. The method according to claim 2, wherein, The shared encoder comprises at least two layers of shared subnetworks arranged in sequence; The step of using the shared encoder included in the shared coding network to encode the projection data and the image pre-coded data respectively to obtain the photon statistical features and the image features includes: For any one of the at least two shared sub-networks, the data input to the shared sub-network is encoded and downsampled using the shared parameters of that shared sub-network to obtain the downsampled image features and photon statistical features, and the encoded image features and photon statistical features before downsampling are backed up.

5. The method according to claim 4, wherein, The depth decoder includes at least two layers of depth decoding sub-networks; The step of using the depth decoder included in the image processing model to decode the photon statistical features and using the decoded photon statistical features to generate a depth map corresponding to the target scene includes: For any one of the at least two deep decoding subnetworks, the photon statistical features input to that deep decoding subnetwork are upsampled, and the upsampled photon statistical features are merged with the backup photon statistical features corresponding to that deep decoding subnetwork; the merged photon statistical features are then decoded to obtain the decoded photon statistical features. If the current deep decoding subnetwork is not the last layer, the decoded photon statistical features are input into the next deep decoding subnetwork. If the depth decoding subnetwork is the last layer, depth calculation is performed on the decoded photon statistical features to obtain the depth map.

6. The method according to claim 4, wherein, The image decoder includes at least two layers of image decoding sub-networks; The step of using the image decoder included in the image processing model to decode the image features and obtain the denoised image includes: For any one of the at least two image decoding sub-networks, the image features input to that image decoding sub-network are upsampled, and the upsampled image features are merged with the backup image features corresponding to that image decoding sub-network; the merged image features are then decoded to obtain the decoded image features. If the image decoding subnetwork is not the last layer, the decoded image features are input into the next layer of the image decoding subnetwork; If this image decoding subnetwork is the last layer, image reconstruction is performed on the decoded image features to obtain the denoised image.

7. The method according to claim 1, wherein, The image processing model was pre-trained according to the following steps: Acquire the sample scene image and the corresponding sample denoised image, as well as the sample photon statistics map and the corresponding sample depth map; The sample scene image and the sample photon statistical map are input into the initial image processing model to obtain the predicted depth map and the predicted denoised image. Based on a preset depth loss function, the depth loss between the sample depth map and the predicted depth map is determined, and based on a preset image denoising loss function, the denoising loss between the sample denoised image and the predicted denoised image is determined. Based on the depth loss and the noise reduction loss, the initial image processing model is trained, including the shared coding network, the depth decoder, and the image decoder. In response to the initial image processing model meeting the preset training termination condition, the current initial image processing model is determined as the image processing model that has been trained.

8. The method according to claim 7, wherein, The initial image processing model, trained based on the depth loss and the noise reduction loss, includes a shared encoder network, a depth decoder, and an image decoder, comprising: According to the phased training strategy, in the first phase, the shared coding network is trained based on the depth loss and the noise reduction loss; In the second stage, the depth loss and the noise reduction loss are redefined, and the depth decoder and the image decoder are trained based on the redefined depth loss and the noise reduction loss.

9. The method according to claim 8, wherein, The step of training the shared coding network based on the depth loss and the noise reduction loss includes: The projection units included in the shared encoding network are trained with the goal of minimizing the depth loss. The shared coding network is trained with the goal of minimizing the noise reduction loss by including the image precoding units. The shared encoder, which is part of the shared coding network, is trained with the goal of minimizing the weighted sum of the depth loss and the noise reduction loss.

10. The method according to claim 8, wherein, The training of the depth decoder and the image decoder based on the redefined depth loss and the noise reduction loss includes: The depth decoder and the image decoder are trained with the goal of minimizing the weighted sum of the redefined depth loss and the noise reduction loss.

11. The method according to claim 7, wherein, The determination of the depth loss between the sample depth map and the predicted depth map based on a preset depth loss function includes: Determine the signal-to-noise ratio of the photon statistical data sequence corresponding to each pixel in the sample photon statistical map; Based on the signal-to-noise ratio, determine the loss weight corresponding to each pixel in the sample photon statistical map; The loss weights are used as parameters of the depth loss function to determine the depth loss between the sample depth map and the predicted depth map.

12. The method according to claim 7, wherein, The method of determining the denoising loss between the sample denoised image and the predicted denoised image based on a preset image denoising loss function includes: Based on the absolute error loss function included in the image denoising loss function, the absolute error loss between the sample denoised image and the predicted denoised image is determined; Based on the perceptual loss function included in the image denoising loss function, the perceptual loss between the sample denoised image and the predicted denoised image is determined; The noise reduction loss is obtained by weighted summation of the absolute error loss and the perception loss.

13. The method according to claim 8, wherein, In the second stage, after redetermining the depth loss and the noise reduction loss, and training the depth decoder and the image decoder based on the redetermined depth loss and the noise reduction loss, the method further includes: In the third stage, the depth loss and the noise reduction loss are redefined, and the learning rate is reset. Based on the depth loss and the noise reduction loss redefined in the third stage, and the learning rate, the shared encoder network, the depth decoder, and the image decoder are jointly trained.

14. An image processing apparatus, comprising: The acquisition module is used to acquire a scene image and a photon statistics map corresponding to the target scene. For any pixel in the photon statistics map, the pixel corresponds to a photon statistics data sequence. The photon statistics data in the photon statistics data sequence is used to characterize the number of reflected photons received from the target location at the corresponding time. The target location is the position in the target scene corresponding to the pixel. The encoding module is used to encode the scene image and the photon statistical map respectively using a shared encoding network included in a pre-trained image processing model to obtain image features and photon statistical features; The first decoding module is used to decode the photon statistical features using the depth decoder included in the image processing model, and to generate a depth map corresponding to the target scene using the decoded photon statistical features. The second decoding module is used to decode the image features using the image decoder included in the image processing model to obtain a denoised image.

15. A computer-readable storage medium storing a computer program, which, when executed, implements the image processing method according to any one of claims 1-13.

16. An electronic device, the electronic device comprising: Memory is used to store processor-executable instructions; A processor is configured to read the executable instructions from the memory and execute the instructions to implement the image processing method according to any one of claims 1-13.