Multi-task inverse imaging method, system, terminal and readable storage medium based on expert mixture cooperative diffusion operator learning
The multi-task inverse imaging method based on expert mixed collaborative diffusion operator learning solves the problems of low image reconstruction accuracy and geometric structure ambiguity under high-dimensional sparse observation and multimodal fusion, achieves high-quality image reconstruction and geometric structure preservation, and is suitable for a variety of image processing tasks and high-dimensional scenarios.
Patent Information
- Application Number
- CN202510681032.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Existing technologies have problems such as low accuracy, blurred geometric structure and insufficient prior modeling in image inverse problems under high-dimensional sparse observations, multimodal fusion and physical consistency constraints.
A multi-task inverse imaging method based on expert mixed collaborative diffusion operator learning is adopted. By obtaining multiple images to be processed input by the user, normalizing them and inputting them into the branch network for adaptive average pooling and feature fusion, the feature fusion model and the expert backbone network are combined to construct the reconstruction loss, geometric loss and expert diversity loss functions for optimization, and output high-quality target images.
It significantly improves the accuracy of image reconstruction and the clarity of geometric structures, reduces gradient explosion and sensitivity to initial parameters, and is suitable for tasks such as image denoising, image restoration, and super-resolution, and can be extended to computational physics and remote sensing reconstruction.
Smart Images

Figure CN120198764B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and analysis, and in particular to a multi-task inverse imaging method, system, terminal and computer-readable storage medium based on expert mixed collaborative diffusion operator learning. Background Art
[0002] In fields such as computational imaging, scientific computing, and artificial intelligence image processing, image reconstruction tasks are a fundamental and critical problem. The main goal is to restore high-quality images or physical field distributions when the observation data is limited, incomplete, or noisy.
[0003] However, existing methods still have limitations. They lack effective modeling of geometric information such as edges and curvature in images, resulting in excessive texture smoothing or insufficient detail recovery. For situations with multiple data modalities (such as multispectral images), existing models often lack modeling and fusion mechanisms for uncertainty between observations, which easily leads to information conflicts.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide a multi-task inverse imaging method, system, terminal and computer-readable storage medium based on expert mixed cooperative diffusion operator learning, aiming to solve the problems of low accuracy, blurred geometric structure and insufficient prior modeling in the image inverse problem under high-dimensional sparse observation, multimodal fusion and physical consistency constraint scenarios in the existing technology.
[0006] To achieve the above object, the present invention provides a multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning, the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning comprising the following steps:
[0007] Acquire multiple images to be processed input by a user, and perform normalization processing on all the images to be processed according to their types to obtain an image tensor;
[0008] Inputting the image tensor into a branch network, the branch network performs adaptive average pooling on the image tensor to generate multiple pooling features of the image tensor, and fuses all the pooling features to output a branch feature representation;
[0009] The branch feature representation is input into a feature fusion model, a high-dimensional feature representation is output, and the high-dimensional feature representation is mapped to a target image control through an output head to obtain a target reconstructed image.
[0010] Optionally, the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning, wherein the step of obtaining multiple images to be processed input by a user and performing normalization processing on all the images to be processed according to their types to obtain an image tensor, specifically includes:
[0011] Obtain all images to be processed input by the user, and determine the number, color and size of all the images to be processed;
[0012] If the color of the image to be processed is color, the number of channels of the corresponding image to be processed is set to multiple; if the color of the image to be processed is black and white, the number of channels of the corresponding image to be processed is set to single;
[0013] According to the size of each of the images to be processed, adjusting the resolution of all the images to be processed to a unified standard;
[0014] According to the quantity, all the adjusted resolutions and all the channel numbers, all the images to be processed are converted into a unified tensor format and fused to obtain an image tensor, wherein the image tensor includes quantity information of the images to be processed.
[0015] Optionally, the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning, wherein the image tensor is input into a branch network, the branch network performs adaptive average pooling on the image tensor to generate multiple pooling features of the image tensor, and all the pooling features are fused to output a branch feature representation, specifically including:
[0016] Inputting the image tensor into a pyramid module of a branch network, wherein the pyramid module performs multi-scale adaptive average pooling processing on the image tensor to obtain multiple pooled feature representations;
[0017] After upsampling all the pooled feature representations to a preset spatial dimension, all the pooled feature representations and all the images to be processed are spliced together to output a branch feature representation, wherein the preset spatial dimension is the same as the dimension of the image tensor.
[0018] Optionally, the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning, wherein the pyramid module performs multi-scale adaptive average pooling processing on the image tensor to obtain multiple pooling feature representations, specifically includes:
[0019] Acquire multiple pooling scales specified by a user, and adjust multiple pooling windows of the pyramid module according to all the pooling scales;
[0020] Pooling is performed on the image tensor, and multiple pooled feature representations are output through all the pooling windows.
[0021] Optionally, the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning, wherein the branch feature representation is input into a feature fusion model, a high-dimensional feature representation is output, and the high-dimensional feature representation is mapped to a target image control through an output head to obtain a target reconstructed image, specifically includes:
[0022] Inputting the branch feature representation into the feature fusion model to perform dimension adjustment and feature refinement, and outputting a high-dimensional feature representation;
[0023] Inputting the high-dimensional feature representation into an output head, the output head transforms the high-dimensional feature representation into features that match the target image space, and performs mapping to obtain a target reconstructed image;
[0024] The number of channels represented by the transformed high-dimensional feature matches the number of channels in the target image space.
[0025] Optionally, the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning, wherein the branch feature representation is input into an optimized feature fusion model for dimensionality adjustment and feature refinement, a high-dimensional feature representation is output, and the high-dimensional feature representation is mapped to a target image control through an output head to obtain a target reconstructed image, further comprising:
[0026] Extracting the spatial coordinates of the image tensor, and inputting the spatial coordinates into different types of expert backbone networks, respectively, to output corresponding expert feature representations;
[0027] Inputting the spatial coordinates into a gating network, the gating network outputting a weight tensor corresponding to each of the expert backbone networks;
[0028] A reconstruction loss function and a geometric loss function are constructed, and an expert diversity loss function is constructed based on all the weight tensors and all the expert feature representations. The constructed feature fusion model is optimized based on the reconstruction loss function, the geometric loss function and the expert diversity loss function.
[0029] Optionally, the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning, wherein the reconstruction loss function and the geometric loss function are constructed, and the expert diversity loss function is constructed according to all the weight tensors and all the expert feature representations, and the constructed feature fusion model is optimized according to the reconstruction loss function, the geometric loss function and the expert diversity loss function, specifically including:
[0030] Obtain the pixel-level difference between the target reconstructed image and the image to be processed, and construct a reconstruction loss function based on the pixel-level difference:
[0031] ;
[0032] in, represents the reconstruction loss function, represents the target reconstructed image, Represents the image to be processed, Indicates the total number of pixels, represents the pixel index, Indicates the target reconstructed image, Indicates the images to be processed;
[0033] Obtain the curvature matching loss and gradient matching loss in the curvature driven diffusion model, and construct a geometric loss function based on the curvature matching loss and the gradient matching loss:
[0034] ;
[0035] in, represents the geometric loss function, represents the weight of the curvature matching loss, represents the curvature matching loss, represents the weight of the gradient matching loss, represents the gradient matching loss;
[0036] Constructing the weights of the expert backbone network based on all the weight tensors, and constructing the expert diversity loss function based on all the expert feature representations;
[0037] According to the reconstruction loss function, the geometric loss function and the expert diversity loss function, a total loss function is constructed, and gradient optimization is performed on the branch network, the feature fusion network and all the expert backbone networks according to the total loss function:
[0038] ;
[0039] in, represents the total loss function, Represents the model parameters of the branch network, feature fusion network or expert backbone network, represents the weight of the geometric loss function, represents the weight of the expert backbone network, represents the expert diversity loss function, represents the number of expert backbone networks, Indicates the The expert feature representation output by the expert backbone network is Indicates the number index of the expert backbone network.
[0040] In addition, to achieve the above-mentioned object, the present invention further provides a multi-task inverse imaging system based on expert mixture cooperative diffusion operator learning, wherein the multi-task inverse imaging system based on expert mixture cooperative diffusion operator learning includes:
[0041] A preprocessing module is used to obtain multiple images to be processed input by a user, and perform normalization processing on all the images to be processed according to their types to obtain image tensors;
[0042] a feature extraction module, configured to input the image tensor into a branch network, wherein the branch network performs adaptive average pooling on the image tensor to generate multiple pooled features of the image tensor, fuse all the pooled features, and output a branch feature representation;
[0043] The image reconstruction module is used to adjust the dimension and refine the features of the branch feature representation input feature fusion model, output a high-dimensional feature representation, and map the high-dimensional feature representation to the target image control through the output head to obtain a target reconstructed image.
[0044] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning stored in the memory and runnable on the processor, wherein the multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning, when executed by the processor, implements the steps of the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning as described above.
[0045] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning, and when the multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning is executed by a processor, the steps of the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning as described above are implemented.
[0046] In the present invention, multiple images to be processed are obtained as input by the user, and all the images to be processed are normalized according to their types to obtain an image tensor; the image tensor is input into a branch network, and the branch network performs adaptive average pooling on the image tensor to generate multiple pooling features of the image tensor, and all the pooling features are fused to output a branch feature representation; the branch feature representation is input into a feature fusion model for dimension adjustment and feature refinement, and a high-dimensional feature representation is output, and the high-dimensional feature representation is mapped to a target image control through an output head to obtain a target reconstructed image. The present invention performs calculations through a separable backbone network and combines a hybrid expert mechanism to improve the convergence speed under high-dimensional tasks, significantly reduce gradient explosion and initial parameter sensitivity, and can be widely applied to image inverse problem tasks such as image denoising, image restoration, super-resolution, and motion blur recovery, and can be extended to high-dimensional scenarios such as computational physics and remote sensing reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a flow chart of a preferred embodiment of the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning of the present invention;
[0048] Figure 2 1 is a system framework diagram of a preferred embodiment of the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning of the present invention;
[0049] Figure 3 1 is a structural diagram of a preferred embodiment of a multi-task inverse imaging system based on expert mixture cooperative diffusion operator learning according to the present invention;
[0050] Figure 4 FIG. 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0052] The multi-task inverse imaging method based on expert mixed cooperative diffusion operator learning described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning includes the following steps:
[0053] Step S10: Acquire multiple images to be processed input by the user, and perform normalization processing on all the images to be processed according to their types to obtain image tensors.
[0054] Among them, the images to be processed input by the user can be in any common format or original size. After entering the system, these images to be processed will be uniformly converted into a standard tensor format and then used as the input of the branch network, so that any image has a standardized processing process, which can reduce the impact of input features of different scales on the model learning process, prevent certain features from dominating the update process due to excessive values, and make the model less sensitive to small changes in the input data, thereby improving the stability of the model and its generalization ability to new data.
[0055] Specifically, all images to be processed input by the user are obtained, and the number, color and size of all the images to be processed are determined; if the color of the image to be processed is color, the number of channels of the corresponding image to be processed is set to multiple, and if the color of the image to be processed is black and white, the number of channels of the corresponding image to be processed is set to single; according to the size of each image to be processed, the resolution of all the images to be processed is adjusted to a unified standard; according to the number, all the adjusted resolutions and all the number of channels, all the images to be processed are converted into a unified tensor format and fused to obtain an image tensor, wherein the image tensor includes the quantity information of the images to be processed.
[0056] The standard tensor format obtained after the unified conversion and fusion of the images is [B, C, H, W], where B represents the batch size, that is, the number of images to be processed simultaneously in this batch. Batch processing can effectively improve efficiency; C represents the number of channels. For example, for color images, it includes three channels, RGB (representing red, green, and blue color channels respectively), and for grayscale images, there is only one channel; H and W represent the height and width of the image to be processed, respectively. When these dimensions are input, the image will be uniformly adjusted to the preset resolution (for example, 256×256 pixels) to ensure the consistency of subsequent network processing.
[0057] Furthermore, each pixel value of the image is normalized. The pixel values of the original image are usually in the range of 0-255. Normalization linearly maps these pixel values (that is, the brightness or color intensity values of all pixels in the image) to a new numerical range, usually [-1, 1] or [0, 1]. The general formula for this process can be expressed as:
[0058] ;
[0059] in, represents the normalized pixel value, represents the original pixel value, represents the mean pixel value, Represents the standard deviation of pixel values. Normalization helps optimization algorithms (such as gradient descent) find the optimal solution more quickly by adjusting the input feature values to a similar range, because the gradient of the loss function will be more balanced. By ensuring that the input data is in a well-defined numerical space, the activation function of the neural network works in its most effective area, avoiding the problem of gradient vanishing or exploding, so that the network can more effectively learn and extract meaningful semantic features. For example, if the pixel value is directly input without normalization, the larger pixel value may cause the activation function to saturate, making the gradient close to zero, thereby hindering learning. The normalized value can better stimulate the response of each layer of the network and promote the effective extraction of deep features.
[0060] Step S20: input the image tensor into the branch network, the branch network performs adaptive average pooling on the image tensor, generates multiple pooling features of the image tensor, fuses all the pooling features, and outputs a branch feature representation.
[0061] Among them, the branch network is mainly composed of multiple stacked convolutional layers (Convolutional Layers), batch normalization layers (BatchNorm), and activation functions, which can gradually extract multi-level features from low-level textures to high-level abstract concepts. After the image tensor passes through the core convolution structure, the pyramid pooling module is introduced to further enhance the feature representation capability and extract deep semantic features.
[0062] Specifically, the image tensor is input into the pyramid module of the branch network, and the pyramid module performs multi-scale adaptive average pooling processing on the image tensor to obtain multiple pooled feature representations; after all the pooled feature representations are upsampled to a preset spatial dimension, all the pooled feature representations and all the images to be processed are spliced to output a branch feature representation, wherein the preset spatial dimension is the same as the dimension of the image tensor.
[0063] Furthermore, multiple pooling scales specified by a user are obtained, and multiple pooling windows of the pyramid module are adjusted according to all the pooling scales; pooling processing is performed on the image tensor, and multiple pooling feature representations are output through all the pooling windows.
[0064] Among them, such as Figure 2As shown in the figure, the branch network performs adaptive average pooling on the input image tensor at multiple different scales (levels). Unlike fixed-size pooling windows, adaptive pooling automatically calculates the size and stride of the pooling window based on the specified output size. For example, for a feature map, the pyramid pooling module may pool the image tensor into feature representations of different sizes, such as 1×1, 2×2, and 4×4. This can capture contextual information at different scales, making the network more robust to changes in object size and position. Each pooled result represents global or semi-global contextual features at different granularities.
[0065] Furthermore, after obtaining multiple pooled feature representations at different scales, these features are typically dimensionalized or transformed using 1×1 convolutions. These processed multi-scale pooled feature representations are then upsampled to the same spatial dimensions as the image tensor input to the pyramid pooling module (or a unified target dimension). Finally, these upsampled multi-scale pooled feature representations are concatenated with the original image (or the image after processing by the backbone convolutional network) along the channel dimension. This concatenation operation aggregates contextual information from different scales and receptive fields to form a richer and more comprehensive branch feature representation. The final output feature dimension is [B, branch_channels, H, W], where branch_channels is the number of channels added after concatenation. This branch feature representation, which incorporates multi-scale contextual information, provides high-quality input for subsequent tasks such as image reconstruction and segmentation.
[0066] Step S30: input the branch feature representation into a feature fusion model, output a high-dimensional feature representation, and map the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image.
[0067] The feature fusion model is primarily responsible for processing and fusing features from upstream networks (primarily branch networks). In the specific implementation of this framework, its structure is relatively simple, typically consisting of a convolutional layer (Convf), followed by batch normalization (BNf) and an activation function.
[0068] Specifically, the branch feature representation is input into a feature fusion model for dimension adjustment and feature refinement, and a high-dimensional feature representation is output; the high-dimensional feature representation is input into an output head, and the output head transforms the high-dimensional feature representation into features that match the target image space, and performs mapping to obtain a target reconstructed image; wherein the number of channels of the transformed high-dimensional feature representation matches the number of channels of the target image space.
[0069] Among them, adjusting the number of channels of branch feature representation through the convolution layer of the feature fusion model, or further abstracting and refining features through convolution operations and introducing activation functions can enhance the expressive power of the module, while the subsequent batch normalization layer can improve the stability of the training process and accelerate convergence.
[0070] In this model, the feature fusion model serves as the key link connecting the encoder (branch network) and the final output (through the output head). It performs the final processing and preparation of the deep semantic features extracted from the input image. It is suitable for various image-to-image conversion tasks such as image restoration and generation, in which the encoded features need to be mapped back to the target image space.
[0071] The function of the output head is to map the high-dimensional feature representation processed by the feature fusion module back to the target image space to generate a preliminary reconstructed image. Specifically, the output head consists of one or two convolutional layers, including batch normalization and activation functions. The first convolutional layer can further transform the features. The batch normalization and activation functions are similar to those in the fusion module. The number of output channels of the final convolutional layer usually matches the number of channels of the target image (for example, 3 for color images and 1 for grayscale images). The convolution kernel and bias of this layer learn how to reconstruct pixel values from high-level features. Batch normalization and activation functions are usually not used (or linear activation is used) to allow the output of pixel values in any range (or clipping, Sigmoid (S-type function) / Tanh (hyperbolic tangent) activation layers to constrain to a specific range, such as [0, 1] or [-1, 1]).
[0072] Furthermore, the spatial coordinates of the image tensor are extracted, and the spatial coordinates are respectively input into different types of expert backbone networks, and the corresponding expert feature representations are output respectively; the gating network inputs the spatial coordinates into the gating network, and outputs the weight tensor corresponding to each of the expert backbone networks; a reconstruction loss function and a geometric loss function are constructed, and an expert diversity loss function is constructed based on all the weight tensors and all the expert feature representations, and the constructed feature fusion model is optimized based on the reconstruction loss function, the geometric loss function and the expert diversity loss function.
[0073] The backbone network plays a key role in the framework, processing spatial coordinate information and assisting in generating location-aware features. Working in conjunction with the branch network, the branch network processes the image content itself, while the backbone network focuses on understanding and encoding the relative or absolute positions of pixels in the image.
[0074] Among them, the backbone network receives the spatial coordinates of the normalized image tensor, which are usually constructed into a tensor that matches the input image or feature map in the spatial dimension; by specifically processing the coordinate information, the network can learn position-related priors, which is crucial for image restoration tasks that require fine spatial control (such as precise filling of missing areas in image repair, or accurate placement of details in super-resolution).
[0075] In order to improve the modeling ability and robustness, this framework adopts an integration strategy and combines the Mixture-of-Experts (MoE) mechanism; for each spatial axis (For example, corresponding to the width and height of the image), the overall output of the backbone network is obtained by averaging the outputs of multiple independently trained individual backbone networks. This helps reduce the bias of individual models and improve generalization ability.
[0076] This process gives the model the ability to understand and utilize spatial context, which is crucial for structure preservation and detail generation; by separating coordinate encoding and content encoding (branch networks), more focused and potentially lightweight network modules can be designed; and the introduction of the MoE mechanism allows the model to dynamically select or weight different "expert" networks based on the input, thereby adapting to more diverse data and tasks. Guided by precise geometric information, it helps to generate images that are more consistent with the structure of the real scene, and can effectively handle a variety of image inverse problems, such as denoising, restoration, super-resolution, deblurring, etc., especially in scenarios that require high-fidelity structure restoration; at the same time, medical image analysis can also be performed. For example, in MRI or CT image reconstruction, accurate spatial correspondence is crucial for diagnosis; further, for fields of computer graphics, such as texture synthesis and view synthesis, it is also necessary to precisely control the spatial layout of the generated content.
[0077] The pixel-level difference between the target reconstructed image and the image to be processed is obtained, and a reconstruction loss function is constructed according to the pixel-level difference:
[0078] ;
[0079] in, represents the reconstruction loss function, represents the target reconstructed image, Represents the image to be processed, Indicates the total number of pixels, represents the pixel index, Indicates the target reconstructed image, Indicates the images to be processed; obtain the curvature matching loss and gradient matching loss in the curvature driven diffusion model, and construct a geometric loss function based on the curvature matching loss and the gradient matching loss:
[0080] ;
[0081] in, represents the geometric loss function, represents the weight of the curvature matching loss, represents the curvature matching loss, represents the weight of the gradient matching loss, Represents the gradient matching loss; constructs the weight of the expert backbone network based on all the weight tensors, and constructs the expert diversity loss function based on all the expert feature representations; constructs the total loss function based on the reconstruction loss function, the geometric loss function and the expert diversity loss function, and performs gradient optimization on the branch network, the feature fusion network and all the expert backbone networks based on the total loss function:
[0082] ;
[0083] in, represents the total loss function, Represents the model parameters of the branch network, feature fusion network or expert backbone network, represents the weight of the geometric loss function, represents the weight of the expert backbone network, represents the expert diversity loss function, represents the number of expert backbone networks, Indicates the The expert feature representation output by the expert backbone network is Indicates the number index of the expert backbone network.
[0084] Among them, the overall architecture of the feature fusion module includes the backbone network, but the output of the backbone network does not directly participate in the feature fusion step in the forward propagation path. Instead, it affects the overall optimization process of the model through the loss function (especially the Hölder divergence regularization when using the MoE backbone network, or other possible coordinate-based loss terms), thereby indirectly affecting the quality of the branch features received by the fusion module.
[0085] This approach utilizes multiple independent "Expert Networks," each typically employing an architecture similar to a standard backbone network to independently process the spatial coordinates of the image tensor. In this way, different expert networks can learn to focus on different aspects or subregions of the input space, thereby improving the overall expressiveness and adaptability of the model.
[0086] Specifically, the input projection layer of an expert network is usually a convolutional layer, which is used to map the input coordinates to a higher feature dimension, stack multiple separable operator blocks (SOBs), each of which performs efficient feature transformation, and finally the output projection layer maps the transformed features back to the required output dimension.
[0087] Among them, each separable operator block includes depth-wise separable convolution to improve computational efficiency; two batch normalization layers are used to stabilize the training process; activation functions and multi-scale attention modules are used to capture feature dependencies at different scales; and a squeeze-and-excitation module is also included for inter-channel feature recalibration.
[0088] Furthermore, each expert is expected to learn a specific aspect of the data or task; for example, in image processing, one expert may excel at processing the coordinates of edge regions, while another expert may excel at processing the coordinates of smooth regions.
[0089] Furthermore, a gating network can be set up to dynamically assign "importance" weights (weight tensor) to each expert network based on the image tensor. The input of the gating network can be spatial coordinates or other important features of the image tensor. It mainly includes one or two convolutional layers followed by a Softmax function. Softmax ensures that for each input position (or sample), the sum of the weights of all experts is 1. These weights can be interpreted as the probability or relative importance of each expert's contribution to the final output.
[0090] Furthermore, the hybrid expert network can finally output the final expert feature representation obtained by weighted summation of the outputs of multiple expert networks:
[0091] ;
[0092] in, represents the final expert feature representation, represents the spatial coordinates, represents the number of expert networks, represents the index of the expert network, Indicates the A weight tensor, Indicates the The output of the expert network, Represents element-by-element multiplication; the input of the original expert network will also be retained for subsequent calculation of the loss function of the hybrid expert network.
[0093] Furthermore, during training, the model aims to minimize the value of this objective function by adjusting its internal parameters (weights and biases). Through optimization algorithms such as gradient descent, the model iteratively updates its parameters so that the predictions are closer and closer to the real situation. In this framework, the optimization objective function guides the entire network (including the branch network, backbone network, fusion module, output head, and optional GCDD layer (Gaussian Curvature-Driven Diffusion) and gating network parameters) to learn how to recover high-quality images from damaged inputs, making them as close as possible to the original clear images.
[0094] The total loss is obtained by weighted summation of multiple sub-loss items, including reconstruction loss, geometric loss, and expert diversity loss. During the optimization process, the Adam (Adaptive Moment Estimation) optimizer or other gradient descent algorithms are used to calculate the gradient of the model parameters according to the total loss function and perform iterative updates:
[0095] ;
[0096] in, Indicates the The model parameters of the branch network, feature fusion network or expert backbone network are updated in the iteration. Indicates the The model parameters of the branch network, feature fusion network or expert backbone network are updated in the iteration. represents the Adam optimizer, Indicates the The learning rate updated in the iteration, express, and represents the hyperparameters of the Adam optimizer, Indicates gradient calculation on the parameters of the branch network, feature fusion network, or expert backbone network model.
[0097] In the CelebA dataset (CelebFaces Attributes Dataset), the present invention compares the image restoration effects of multiple methods in the Denoising (denoising) task, Inpainting (image restoration) task, and Super-Resolution (super-resolution) task. The quantitative comparison of the three groups of tasks in terms of PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index Measure) is shown in Table 1 below:
[0098] Table 1: Quantitative comparison table
[0099]
[0100] Among them, VDVAE stands for Very Deep Variational Autoencoders, DPS stands for Data Processing System, PnP-HVAE stands for Perspective-n-Point Hierarchical Variational Autoencoder (a framework that combines the geometric modeling capabilities of PnP and the hierarchical probability of HVAE), and BPnP stands for Backpropagation Perspective-n-Point (an improved method for the traditional Perspective-n-Point (PnP) algorithm in the field of computer vision, supporting end-to-end differentiable computing).
[0101] The image results output by each network model disclosed in the present invention are improved by 5-8 decibels in terms of PSNR and SSIM indicators compared with traditional methods, especially when the noise level is greater than 30%; by constraining the expert output distribution difference through Holder divergence, the weight of Holder divergence can be gradually increased during the first 10 passes through all data to avoid model collapse caused by initial training instability and improve the generalization ability of the model. The test error on the BSD300 dataset (Berkeley Segmentation Dataset and Benchmark) is reduced by 12%; and the curvature function (in addition to the GCDD proposed in this embodiment, other curvature functions can also be used instead, such as TSC (Tensor-based Surface Curvature, tensor-based surface curvature), TRV (Tangential Relative Variation, tangential relative change curvature), TAC (Total Absolute Curvature, Total Absolute Curvature), MS (Mean Squared Curvature (mean square curvature), Geman (Generalized Euler-Manifold, generalized Euler manifold curvature), Log-det (Logarithm of Determinant Curvature Function, logarithmic determinant curvature function), Laplace (Laplace Operator, Laplace operator) and EE (Euler-Elastica Curvature Function, Euler elastic curvature function)) dynamically adjust the diffusion intensity, suppress and smooth high curvature areas (edges), improve edge preservation rate by 30%, and suppress artifact generation; the coordinate separation backbone network decomposes calculations along the coordinate axis, reducing dimensional complexity and memory usage by 40%, making it suitable for a variety of high-resolution image processing.
[0102] Furthermore, the quantitative results of blur recovery on the BSD300 dataset under different noise intensities are shown in Table 2 below:
[0103] Table 2: Quantitative results of fuzzy recovery
[0104]
[0105] Among them, GS-PNP stands for Gaussian Splatting-Perspective-n-Point JointOptimization, EPLL stands for Expected Patch Log-Likelihood (an optimization method based on a probability model), and PnP-MM stands for Plug-and-Play Majorization-Minimization.
[0106] The present invention uses a separable backbone network for calculation and combines it with a hybrid expert mechanism to improve the convergence speed under high-dimensional tasks and significantly reduce gradient explosion and initial parameter sensitivity. It can be widely applied to image inverse problem tasks such as image denoising, image restoration, super-resolution, and motion blur recovery, and can be extended to high-dimensional scenarios such as computational physics and remote sensing reconstruction.
[0107] Furthermore, if Figure 3 As shown, based on the above-mentioned multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning, the present invention also provides a multi-task inverse imaging system based on expert mixture collaborative diffusion operator learning, wherein the multi-task inverse imaging system based on expert mixture collaborative diffusion operator learning includes:
[0108] A pre-processing module 51 is configured to obtain a plurality of images to be processed input by a user, and perform normalization processing on all the images to be processed according to their types to obtain image tensors;
[0109] a feature extraction module 52 for inputting the image tensor into a branch network, wherein the branch network performs adaptive average pooling on the image tensor to generate multiple pooled features of the image tensor, fuses all the pooled features, and outputs a branch feature representation;
[0110] The image reconstruction module 53 is used to adjust the dimension and refine the features of the branch feature representation input into the feature fusion model, output a high-dimensional feature representation, and map the high-dimensional feature representation to the target image control through the output head to obtain a target reconstructed image.
[0111] Furthermore, if Figure 4 As shown, based on the above-mentioned multi-task inverse imaging method and system based on expert mixture cooperative diffusion operator learning, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.
[0112] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard drive or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash card, etc. equipped on the terminal. Furthermore, the memory 20 may include both the internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software installed on the terminal and various types of data, such as program code of the installed terminal. The memory 20 may also be used to temporarily store data that has been output or is about to be output. In one embodiment, the memory 20 stores a multi-task inverse imaging program 40 based on expert mixture collaborative diffusion operator learning. The multi-task inverse imaging program 40 based on expert mixture collaborative diffusion operator learning can be executed by the processor 10, thereby implementing the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning in the present application.
[0113] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program code or process data stored in the memory 20, such as executing the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning.
[0114] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0115] In one embodiment, when the processor 10 executes the multi-task inverse imaging program 40 based on expert mixture cooperative diffusion operator learning in the memory 20 , the steps of the multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning are implemented.
[0116] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning, and when the multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning is executed by a processor, the steps of the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning as described above are implemented.
[0117] In summary, the present invention provides a multi-task inverse imaging method and related equipment based on expert mixed collaborative diffusion operator learning, the method comprising: obtaining multiple images to be processed input by the user, normalizing all the images to be processed according to the types of all the images to be processed to obtain an image tensor; inputting the image tensor into a branch network, the branch network performing adaptive average pooling on the image tensor to generate multiple pooled features of the image tensor, and fusing all the pooled features to output a branch feature representation; inputting the branch feature representation into a feature fusion model for dimensionality adjustment and feature refinement, outputting a high-dimensional feature representation, and mapping the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image. The present invention uses a separable backbone network for calculation and combines a mixed expert mechanism to improve the convergence speed under high-dimensional tasks, significantly reduce gradient explosion and initial parameter sensitivity, and can be widely applied to image inverse problem tasks such as image denoising, image restoration, super-resolution, motion blur recovery, and can be extended to high-dimensional scenarios such as computational physics and remote sensing reconstruction.
[0118] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.
[0119] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When executed, the program can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0120] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning, characterized in that: The multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning includes: Acquire multiple images to be processed input by a user, and perform normalization processing on all the images to be processed according to their types to obtain an image tensor; Inputting the image tensor into a branch network, the branch network performs adaptive average pooling on the image tensor to generate multiple pooling features of the image tensor, and fuses all the pooling features to output a branch feature representation; The branch feature representation is input into a feature fusion model to adjust the dimension and refine the features, output a high-dimensional feature representation, and map the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image; Extracting the spatial coordinates of the image tensor, and inputting the spatial coordinates into different types of expert backbone networks, respectively, to output corresponding expert feature representations; Inputting the spatial coordinates into a gating network, the gating network outputting a weight tensor corresponding to each of the expert backbone networks; Constructing a reconstruction loss function and a geometric loss function, and constructing an expert diversity loss function based on all the weight tensors and all the expert feature representations, and optimizing the constructed feature fusion model based on the reconstruction loss function, the geometric loss function, and the expert diversity loss function; The constructing of the reconstruction loss function and the geometric loss function, and constructing the expert diversity loss function according to all the weight tensors and all the expert feature representations, and optimizing the constructed feature fusion model according to the reconstruction loss function, the geometric loss function, and the expert diversity loss function, specifically includes: Obtain the pixel-level difference between the target reconstructed image and the image to be processed, and construct a reconstruction loss function based on the pixel-level difference: ; in, represents the reconstruction loss function, represents the target reconstructed image, Represents the image to be processed, Indicates the total number of pixels, represents the pixel index, Indicates the target reconstructed image, Indicates the images to be processed; Obtain the curvature matching loss and gradient matching loss in the curvature driven diffusion model, and construct a geometric loss function based on the curvature matching loss and the gradient matching loss: ; in, represents the geometric loss function, represents the weight of the curvature matching loss, represents the curvature matching loss, represents the weight of the gradient matching loss, represents the gradient matching loss; Constructing the weights of the expert backbone network based on all the weight tensors, and constructing the expert diversity loss function based on all the expert feature representations; According to the reconstruction loss function, the geometric loss function and the expert diversity loss function, a total loss function is constructed, and gradient optimization is performed on the branch network, the feature fusion model and all the expert backbone networks according to the total loss function: ; in, represents the total loss function, represents the parameters of the branch network, feature fusion model or expert backbone network model, represents the weight of the geometric loss function, represents the weight of the expert backbone network, represents the expert diversity loss function, represents the number of expert backbone networks, Indicates the The expert feature representation output by the expert backbone network is Indicates the number index of the expert backbone network.
2. The multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning according to claim 1, characterized in that: The step of obtaining a plurality of images to be processed input by a user and performing normalization processing on all the images to be processed according to their types to obtain an image tensor specifically includes: Obtain all images to be processed input by the user, and determine the number, color and size of all the images to be processed; If the color of the image to be processed is color, the number of channels of the corresponding image to be processed is set to multiple; if the color of the image to be processed is black and white, the number of channels of the corresponding image to be processed is set to single; According to the size of each of the images to be processed, adjusting the resolution of all the images to be processed to a unified standard; According to the quantity, all the adjusted resolutions and all the channel numbers, all the images to be processed are converted into a unified tensor format and fused to obtain an image tensor, wherein the image tensor includes quantity information of the images to be processed.
3. The multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning according to claim 1, characterized in that: The image tensor is input into the branch network, the branch network performs adaptive average pooling on the image tensor, generates multiple pooling features of the image tensor, and fuses all the pooling features to output a branch feature representation, specifically including: Inputting the image tensor into a pyramid module of a branch network, wherein the pyramid module performs multi-scale adaptive average pooling processing on the image tensor to obtain multiple pooled feature representations; After upsampling all the pooled feature representations to a preset spatial dimension, all the pooled feature representations and all the images to be processed are spliced together to output a branch feature representation, wherein the preset spatial dimension is the same as the dimension of the image tensor.
4. The multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning according to claim 3, characterized in that: The pyramid module performs multi-scale adaptive average pooling processing on the image tensor to obtain multiple pooling feature representations, specifically including: Acquire multiple pooling scales specified by a user, and adjust multiple pooling windows of the pyramid module according to all the pooling scales; Pooling is performed on the image tensor, and multiple pooled feature representations are output through all the pooling windows.
5. The multi-task inverse imaging method based on expert mixture cooperative diffusion operator learning according to claim 1, characterized in that: The step of inputting the branch feature representation into a feature fusion model, outputting a high-dimensional feature representation, and mapping the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image specifically includes: Inputting the branch feature representation into the feature fusion model to perform dimension adjustment and feature refinement, and outputting a high-dimensional feature representation; Inputting the high-dimensional feature representation into an output head, the output head transforms the high-dimensional feature representation into features that match the target image space, and performs mapping to obtain a target reconstructed image; The number of channels represented by the transformed high-dimensional feature matches the number of channels in the target image space.
6. A multi-task inverse imaging system based on expert mixture cooperative diffusion operator learning, characterized in that: The multi-task inverse imaging system based on expert mixture collaborative diffusion operator learning is applied to the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning according to any one of claims 1 to 5, and the multi-task inverse imaging system based on expert mixture collaborative diffusion operator learning includes: A preprocessing module is used to obtain multiple images to be processed input by a user, and perform normalization processing on all the images to be processed according to their types to obtain image tensors; a feature extraction module, configured to input the image tensor into a branch network, wherein the branch network performs adaptive average pooling on the image tensor to generate multiple pooled features of the image tensor, fuse all the pooled features, and output a branch feature representation; The image reconstruction module is used to adjust the dimension and refine the features of the branch feature representation input feature fusion model, output a high-dimensional feature representation, and map the high-dimensional feature representation to the target image control through the output head to obtain a target reconstructed image.
7. A terminal, characterized in that: The terminal includes: a memory, a processor, and a multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning stored in the memory and executable on the processor. When the multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning is executed by the processor, the steps of the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning as described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning. When the multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning is executed by a processor, the steps of the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Image denoising method based on double-branch feature fusion
CN119313583A
Image reconstruction method and system based on hybrid network framework, terminal and storage medium
CN119963682A