Multitask inverse imaging method and system based on expert mixed collaborative diffusion operator learning, terminal and readable storage medium
Through the multi-task inverse imaging method based on expert hybrid collaborative diffusion operator learning, the problems of low accuracy, blurred geometric structure and insufficient prior modeling in the image inverse problem in the prior art are solved, and efficient image reconstruction and multi-task adaptability are achieved.
Patent Information
- Application Number
- CN202510681032.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
In the prior art, in the face of image inverse problems in high-dimensional sparse observations, multimodal fusion and physical consistency constraint scenarios, there are problems of low accuracy, blurred geometric structures and insufficient prior modeling.
Using a multi-task inverse imaging method based on expert hybrid collaborative diffusion operator learning, multiple to-process images input by the user are obtained, normalized processing is performed, and input branch networks are input for adaptive average pooling, multiple pooling features are generated and fused, and branch feature representation is output. Then, the branch feature representation input feature fusion model is dimensionally adjusted and feature refinement, and the high-dimensional feature representation is mapped to the target image control through the output head to obtain the target reconstruction image.
It improves the convergence speed under high-dimensional tasks, significantly reduces the sensitivity of gradient explosion and initial parameter, and can be widely used in image inverse problems such as image denoising, image repair, super resolution, motion blur recovery, and can be expanded to high-dimensional scenes such as computational physics and remote sensing reconstruction.
Smart Images

Figure CN120198764A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and analysis, and in particular, to a multi-task inverse imaging method, system, terminal and computer-readable storage medium based on the learning of an expert mixture collaborative diffusion operator. Background Art
[0002] In the fields of computational imaging, scientific computing, and artificial intelligence image processing, the image reconstruction task is a fundamental and crucial problem. The main goal is to recover a high-quality image or physical field distribution under the conditions of limited, incomplete, or noisy observation data.
[0003] However, existing methods still have limitations. There is a lack of effective modeling of geometric information such as edges and curvatures in images, resulting in over-smoothed textures or insufficient detail recovery. For cases with multiple data modalities (such as multi-spectral images, etc.), existing models often lack a mechanism for modeling and fusing the uncertainties between observations, and are prone to information conflicts.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main objective of the present invention is to provide a multi-task inverse imaging method, system, terminal and computer-readable storage medium based on the learning of an expert mixture collaborative diffusion operator, aiming to solve the problems of low accuracy, blurred geometric structure, and insufficient prior modeling in the image inverse problem in the scenarios of high-dimensional sparse observations, multi-modal fusion, and physical consistency constraints in the existing technology.
[0006] To achieve the above objective, the present invention provides a multi-task inverse imaging method based on the learning of an expert mixture collaborative diffusion operator. The multi-task inverse imaging method based on the learning of an expert mixture collaborative diffusion operator includes the following steps: Obtain a plurality of images to be processed input by a user, and perform normalization processing on all the images to be processed according to the types of all the images to be processed to obtain an image tensor; Input the image tensor into a branch network. The branch network performs adaptive average pooling on the image tensor to generate a plurality of pooled features of the image tensor, and fuses all the pooled features to output a branch feature representation; Input the branch feature representation into a feature fusion model, output a high-dimensional feature representation, and map the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image.
[0007] Optionally, in the multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator, the step of obtaining a plurality of images to be processed input by the user and normalizing all the images to be processed according to the types of all the images to be processed to obtain an image tensor specifically includes: Obtain all the images to be processed input by the user, and determine the quantity, color, and size of all the images to be processed; If the color of the image to be processed is color, set the number of channels of the corresponding image to be processed to multiple; if the color of the image to be processed is black and white, set the number of channels of the corresponding image to be processed to single. According to the size of each image to be processed, adjust the resolution of all the images to be processed to a unified standard; According to the quantity, all the adjusted resolutions, and all the numbers of channels, convert all the images to be processed into a unified tensor format and fuse them to obtain an image tensor, where the image tensor includes the quantity information of the images to be processed.
[0008] Optionally, in the multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator, the step of inputting the image tensor into a branch network, where the branch network performs adaptive average pooling on the image tensor, generates a plurality of pooled features of the image tensor, and fuses all the pooled features to output a branch feature representation specifically includes: Input the image tensor into the pyramid module of the branch network, and the pyramid module performs multi-scale adaptive average pooling processing on the image tensor to obtain a plurality of pooled feature representations; After upsampling all the pooled feature representations to a preset spatial dimension, perform splicing processing on all the pooled feature representations and all the images to be processed, and output a branch feature representation, where the preset spatial dimension is the same as the dimension of the image tensor.
[0009] Optionally, in the multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator, the step of the pyramid module performing multi-scale adaptive average pooling processing on the image tensor to obtain a plurality of pooled feature representations specifically includes: Obtain a plurality of pooling scales specified by the user, and adjust a plurality of pooling windows of the pyramid module according to all the pooling scales; Perform pooling processing on the image tensor, and output a plurality of pooled feature representations through all the pooling windows.
[0010] Optionally, for the multi-task inverse imaging method based on the learning of the mixture-of-experts collaborative diffusion operator, the step of inputting the branch feature representation into the feature fusion model, outputting a high-dimensional feature representation, and mapping the high-dimensional feature representation to the target image control through an output head to obtain the target reconstructed image specifically includes: Input the branch feature representation into the feature fusion model for dimension adjustment and feature refinement, and output a high-dimensional feature representation; Input the high-dimensional feature representation into the output head, and the output head transforms the high-dimensional feature representation into a feature matching the target image space and performs mapping to obtain the target reconstructed image; Among them, the number of channels of the transformed high-dimensional feature representation matches the number of channels of the target image space.
[0011] Optionally, for the multi-task inverse imaging method based on the learning of the mixture-of-experts collaborative diffusion operator, after the step of inputting the branch feature representation into the optimized feature fusion model for dimension adjustment and feature refinement, outputting a high-dimensional feature representation, and mapping the high-dimensional feature representation to the target image control through an output head to obtain the target reconstructed image, it further includes: Extract the spatial coordinates of the image tensor, and input the spatial coordinates into different types of expert backbone networks respectively, and output the corresponding expert feature representations; Input the spatial coordinates into the gating network, and the gating network outputs the weight tensors corresponding to each expert backbone network; Construct a reconstruction loss function and a geometric loss function, and construct an expert diversity loss function according to all the weight tensors and all the expert feature representations, and optimize the constructed feature fusion model according to the reconstruction loss function, the geometric loss function, and the expert diversity loss function.
[0012] Optionally, for the multi-task inverse imaging method based on the learning of the mixture-of-experts collaborative diffusion operator, the step of constructing a reconstruction loss function and a geometric loss function, constructing an expert diversity loss function according to all the weight tensors and all the expert feature representations, and optimizing the constructed feature fusion model according to the reconstruction loss function, the geometric loss function, and the expert diversity loss function specifically includes: Obtain the pixel-level difference between the target reconstructed image and the image to be processed, and construct a reconstruction loss function according to the pixel-level difference: ; Among them, represents the reconstruction loss function, represents the target reconstructed image, represents the image to be processed, represents the total number of pixels, represents the pixel index, represents the th target reconstructed image, represents the th image to be processed; Obtain the curvature matching loss and the gradient matching loss in the curvature-driven diffusion model, and construct a geometric loss function according to the curvature matching loss and the gradient matching loss: ; wherein, represents the geometric loss function, represents the weight of the curvature matching loss, represents the curvature matching loss, represents the weight of the gradient matching loss, represents the gradient matching loss; Construct the weights of the expert backbone network according to all the weight tensors, and construct an expert diversity loss function according to all the expert feature representations; Construct a total loss function according to the reconstruction loss function, the geometric loss function and the expert diversity loss function, and perform gradient optimization on the branch network, the feature fusion network and all the expert backbone networks according to the total loss function: ; wherein, represents the total loss function, represents the model parameters of the branch network, the feature fusion network or the expert backbone network, represents the weight of the geometric loss function, represents the weight of the expert backbone network, represents the expert diversity loss function, represents the number of expert backbone networks, represents the th expert feature representation output by the th expert backbone network,
[0013] In addition, to achieve the above object, the present invention also provides a multi-task inverse imaging system based on the learning of an expert hybrid collaborative diffusion operator, wherein the multi-task inverse imaging system based on the learning of an expert hybrid collaborative diffusion operator includes: A preprocessing module, configured to obtain a plurality of images to be processed input by a user, and perform normalization processing on all the images to be processed according to the types of all the images to be processed, so as to obtain an image tensor; A feature extraction module, configured to input the image tensor into a branch network, where the branch network performs adaptive average pooling on the image tensor to generate multiple pooled features of the image tensor, and fuse all the pooled features to output a branch feature representation; An image reconstruction module, configured to input the branch feature representation into a feature fusion model for dimension adjustment and feature refinement, output a high-dimensional feature representation, and map the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image.
[0014] In addition, to achieve the above object, the present invention also provides a terminal, where the terminal includes: a memory, a processor, and a multi-task inverse imaging program based on the learning of a mixture of experts collaborative diffusion operator stored in the memory and executable on the processor. When the multi-task inverse imaging program based on the learning of a mixture of experts collaborative diffusion operator is executed by the processor, the steps of the multi-task inverse imaging method based on the learning of a mixture of experts collaborative diffusion operator as described above are implemented.
[0015] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, where the computer-readable storage medium stores a multi-task inverse imaging program based on the learning of a mixture of experts collaborative diffusion operator. When the multi-task inverse imaging program based on the learning of a mixture of experts collaborative diffusion operator is executed by a processor, the steps of the multi-task inverse imaging method based on the learning of a mixture of experts collaborative diffusion operator as described above are implemented.
[0016] In the present invention, a plurality of images to be processed input by a user are obtained, and all the images to be processed are normalized according to the types of all the images to be processed to obtain an image tensor; the image tensor is input into a branch network, and the branch network performs adaptive average pooling on the image tensor to generate multiple pooled features of the image tensor, and fuse all the pooled features to output a branch feature representation; the branch feature representation is input into a feature fusion model for dimension adjustment and feature refinement, output a high-dimensional feature representation, and map the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image. The present invention performs calculations through a separable backbone network, combines a mixture of experts mechanism, improves the convergence speed in high-dimensional tasks, significantly reduces gradient explosion and initial parameter sensitivity, can be widely applied to image inverse problem tasks such as image denoising, image restoration, super-resolution, motion blur recovery, etc., and can be extended to high-dimensional scenarios such as computational physics and remote sensing reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of a preferred embodiment of the multi-task inverse imaging method based on the learning of a mixture of experts collaborative diffusion operator of the present invention; Figure 2It is the system framework diagram of the preferred embodiment of the multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator of the present invention; Figure 3 It is the first result diagram of image reconstruction of the preferred embodiment of the multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator of the present invention; Figure 4 It is the structure diagram of the preferred embodiment of the multi-task inverse imaging system based on the learning of the expert mixture collaborative diffusion operator of the present invention; Figure 5 It is the structure diagram of the preferred embodiment of the terminal of the present invention. Detailed implementation manners
[0018] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0019] The multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator according to the preferred embodiment of the present invention, as Figure 1 shown, the multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator includes the following steps: Step S10: Obtain a plurality of images to be processed input by a user, and perform normalization processing on all the images to be processed according to the types of all the images to be processed, so as to obtain an image tensor.
[0020] Among them, the images to be processed input by the user can be in any common format or the original size. After being input into the system, these images to be processed will be uniformly converted into a standard tensor format and then used as the input of the branch network, so that any image has a standardized processing process, which can reduce the influence of input features of different scales on the model learning process, prevent some features from dominating the update process due to excessive numerical values, make the model less sensitive to small changes in the input data, and thus improve the stability of the model and the generalization ability to new data.
[0021] Specifically, obtain all the images to be processed input by the user, and determine the quantity, color and size of all the images to be processed; if the color of the image to be processed is color, set the number of channels of the corresponding image to be processed to be multiple, if the color of the image to be processed is black and white, set the number of channels of the corresponding image to be processed to be single; adjust the resolution of all the images to be processed to a unified standard according to the size of each image to be processed; convert all the images to be processed into a unified tensor format and fuse them according to the quantity, all the adjusted resolutions and all the numbers of channels, so as to obtain an image tensor, where the image tensor includes the quantity information of the images to be processed.
[0022] Among them, the standard tensor format obtained by fusing the uniformly transformed images is [B, C, H, W]. Here, B represents the batch size, that is, the number information of the images to be processed in this batch. By batch processing, the efficiency can be effectively improved; C represents the number of channels. For example, for color images, it includes three channels of RGB (representing the red, green, and blue color channels respectively), and for grayscale images, there is only one channel number; H and W respectively represent the height and width of the image to be processed. When inputting these dimensions, the image will be uniformly adjusted to a preset resolution (for example, 256×256 pixels), which can ensure the consistency of subsequent network processing.
[0023] Furthermore, each pixel value of the image is normalized. The pixel values of the original image are usually in the range of 0 - 255. Normalization linearly maps these pixel values (i.e., the brightness or color intensity values of all pixels in the image) to a new numerical range, usually [-1, 1] or [0, 1]. The general formula for this process can be expressed as: ; Among them, represents the normalized pixel value, represents the original pixel value, represents the pixel value mean, represents the pixel value standard deviation; Normalization helps the optimization algorithm (such as gradient descent) to find the optimal solution faster by adjusting the input eigenvalue to a similar range because the gradient of the loss function will be more balanced; By ensuring that the input data is in a well-defined numerical space, the activation function of the neural network works in its most effective region, avoiding the problems of gradient vanishing or explosion, so that the network can learn and extract meaningful semantic features more effectively. For example, if the pixel values are directly input without normalization, larger pixel values may cause the activation function to saturate, making the gradient close to zero, thus hindering learning. The normalized values can better stimulate the response of each layer of the network and promote the effective extraction of deep features.
[0024] Step S20: Input the image tensor into the branch network. The branch network performs adaptive average pooling on the image tensor, generates multiple pooling features of the image tensor, and fuses all the pooling features to output a branch feature representation.
[0025] Among them, the branch network mainly consists of multiple stacked convolutional layers, batch normalization layers (Batch Normalization, BatchNorm), and activation functions, which can gradually extract multi-level features from low-level textures to high-level abstract concepts. After passing the image tensor through the core convolutional structure, a pyramid pooling module is introduced to further enhance the feature representation ability and extract deep semantic features.
[0026] Specifically, the image tensor is input into the pyramid module of the branch network. The pyramid module performs multi-scale adaptive average pooling on the image tensor to obtain multiple pooled feature representations. After upsampling all the pooled feature representations to a preset spatial dimension, all the pooled feature representations are concatenated with all the images to be processed, and a branch feature representation is output, where the preset spatial dimension is the same as the dimension of the image tensor.
[0027] Furthermore, multiple pooling scales specified by the user are obtained, and multiple pooling windows of the pyramid module are adjusted according to all the pooling scales. Pooling processing is performed on the image tensor, and multiple pooled feature representations are output through all the pooling windows.
[0028] Among them, as Figure 2 shown, the branch network performs adaptive average pooling on the input image tensor at multiple different scales (levels). Different from a pooling window of a fixed size, adaptive pooling automatically calculates the size and stride of the pooling window according to the specified output size. For example, for a feature map, the pyramid pooling module may pool the image tensor into feature representations of different sizes such as 1×1, 2×2, 4×4, etc., which can capture context information at different scales, making the network more robust to changes in the size and position of objects. Each result after pooling represents global or semi-global context features at different granularities.
[0029] Further, after obtaining the pooling feature representations of multiple different scales, 1×1 convolutions are usually applied to these features for dimension adjustment or feature transformation. Then, these processed multi-scale pooling feature representations are upsampled to the same spatial dimension (or a certain unified target dimension) as the image tensor input to the pyramid pooling module. Finally, these upsampled multi-scale pooling feature representations are concatenated with the original (or the image to be processed after passing through the backbone convolutional network) in the channel dimension. This concatenation operation aggregates the context information from different scales and receptive fields, forming a richer and more comprehensive branch feature representation. The final output feature dimension is [B, branch_channels, H, W], where branch_channels is the number of channels increased after concatenation. This branch feature representation that integrates multi-scale context information provides high-quality input for subsequent tasks such as image reconstruction and segmentation.
[0030] Step S30: Input the branch feature representation into the feature fusion model, output a high-dimensional feature representation, and map the high-dimensional feature representation to the target image control through the output head to obtain the target reconstructed image.
[0031] Among them, the feature fusion model is mainly responsible for processing and fusing the features from the upstream network (mainly the branch network). In the specific implementation of this framework, its structure is relatively simple, usually including a convolutional layer (Convf), followed by batch normalization (BNf) and an activation function.
[0032] Specifically, input the branch feature representation into the feature fusion model for dimension adjustment and feature refinement, and output a high-dimensional feature representation. Input the high-dimensional feature representation into the output head, and the output head transforms the high-dimensional feature representation into a feature that matches the target image space and performs mapping to obtain the target reconstructed image. Among them, the number of channels of the transformed high-dimensional feature representation matches the number of channels of the target image space.
[0033] Among them, through the convolutional layer of the feature fusion model, the number of channels of the branch feature representation is adjusted, or the features are further abstracted and refined through convolutional operations. Introducing an activation function can enhance the expressive ability of the module, while the subsequent connected batch normalization layer can improve the stability of the training process and accelerate convergence.
[0034] In this model, the feature fusion model, as a key link connecting the encoder (branch network) and the final output (through the output head), performs the final processing and preparation of the deep semantic features extracted from the input image, and is applicable to various image-to-image conversion tasks such as image restoration and generation, where the encoded features need to be mapped back to the target image space.
[0035] Among them, the function of the output head is to map the high-dimensional feature representation processed by the feature fusion module back to the target image space to generate a preliminary reconstructed image. Specifically, the output head consists of one or two convolutional layers, including batch normalization and activation functions; among them, the first convolutional layer can further transform the features; the batch normalization and activation functions are similar to those in the fusion module; the number of output channels of the last convolutional layer usually matches the number of channels of the target image (for example, 3 for color images and 1 for grayscale images). The convolutional kernel and bias of this layer will learn how to reconstruct pixel values from high-level features. Usually, batch normalization and activation functions are not used (or linear activation is used) to allow the output of pixel values in any range (or use a cropping, Sigmoid / Tanh activation layer to constrain to a specific range, such as [0, 1] or [-1, 1]).
[0036] Furthermore, extract the spatial coordinates of the image tensor and input the spatial coordinates into different types of expert backbone networks respectively to output corresponding expert feature representations; the gating network inputs the spatial coordinates into the gating network and outputs a weight tensor corresponding to each expert backbone network; construct a reconstruction loss function and a geometric loss function, and construct an expert diversity loss function according to all the weight tensors and all the expert feature representations, and optimize the constructed feature fusion model according to the reconstruction loss function, the geometric loss function and the expert diversity loss function.
[0037] Among them, the backbone network plays a key role in the framework in processing spatial coordinate information and assisting in generating position-aware features. It works in cooperation with the branch network. The branch network processes the image content itself, while the backbone network focuses on understanding and encoding the relative or absolute positions of pixels in the image.
[0038] Among them, the backbone network receives the spatial coordinates of the normalized image tensor, and these coordinates are usually constructed into a tensor that matches the input image or feature map in the spatial dimension; by specifically processing the coordinate information, the network can learn position-related priors, which are crucial for image restoration tasks that require fine spatial control (such as accurately filling in missing areas in image inpainting or accurately placing details in super-resolution).
[0039] To improve the modeling ability and robustness, this framework adopts an integration strategy and combines the Mixture-of-Experts (MoE) mechanism; for each spatial axis (for example, corresponding to the width and height of the image), the overall output of the backbone network is obtained by averaging the outputs of multiple independently trained single backbone networks. This helps to reduce the bias of a single model and improve the generalization ability.
[0040] This process endows the model with the ability to understand and utilize spatial context, which is crucial for structure preservation and detail generation. By separating coordinate encoding and content encoding (branch networks), more focused and potentially lighter network modules can be designed. The introduction of the MoE mechanism allows the model to dynamically select or weight different "expert" networks according to the input, thus adapting to more diverse data and tasks. Guided by precise geometric information, it helps generate images that better conform to the structure of real scenes and can effectively handle various image inverse problems, such as denoising, inpainting, super-resolution, deblurring, etc., especially in scenarios that require high-fidelity structure recovery. At the same time, it can also perform medical image analysis. For example, in MRI (Magnetic Resonance Imaging) or CT (Computed Tomography) image reconstruction, precise spatial correspondence is crucial for diagnosis. Further, in the field of computer graphics, such as texture synthesis and view synthesis, precise control of the spatial layout of the generated content is also required.
[0041] Among them, obtain the pixel-level difference between the target reconstructed image and the image to be processed, and construct a reconstruction loss function according to the pixel-level difference: ; Among them, represents the reconstruction loss function, represents the target reconstructed image, represents the image to be processed, represents the total number of pixels, represents the pixel index, represents the th target reconstructed image, represents the th image to be processed; Obtain the curvature matching loss and the gradient matching loss in the curvature-driven diffusion model, and construct a geometric loss function according to the curvature matching loss and the gradient matching loss: ; Among them, represents the geometric loss function, represents the weight of the curvature matching loss, represents the curvature matching loss, represents the weight of the gradient matching loss, denote the gradient matching loss; construct the weights of the expert backbone network according to all the weight tensors, and construct an expert diversity loss function according to all the expert feature representations; construct a total loss function according to the reconstruction loss function, the geometric loss function and the expert diversity loss function, and perform gradient optimization on the branch network, the feature fusion network and all the expert backbone networks according to the total loss function: ; wherein, denotes the total loss function, denotes the model parameters of the branch network, the feature fusion network or the expert backbone network, denotes the weight of the geometric loss function, denotes the weight of the expert backbone network, denotes the expert diversity loss function, denotes the number of expert backbone networks, denotes the -th expert feature representation output by the -th expert backbone network,
[0042] wherein, the overall architecture of the feature consensus module includes a backbone network, but the output of the backbone network does not directly participate in the feature fusion step in the forward propagation path. Instead, it affects the overall optimization process of the model through a loss function (especially the Hölder divergence regularization when using the MoE backbone network, or other possible coordinate-based loss terms), thereby indirectly affecting the quality of the branch features received by the fusion module.
[0043] Among them, multiple independent "Expert Networks" are used, and each expert network usually adopts an architecture similar to the standard backbone network to independently process the spatial coordinates of the image tensor. In this way, different expert networks can learn to focus on different aspects or sub-regions of the input space, thereby improving the overall expressive ability and adaptability of the model.
[0044] Specifically, the input projection layer of an expert network is usually a convolutional layer for mapping the input coordinates to a higher feature dimension, stacking multiple Separable Operator Blocks (SOBs), each SOB block performs efficient feature transformation, and finally the output projection layer maps the transformed features back to the required output dimension.
[0045] Among them, each separable operator block includes depthwise separable convolution to improve computational efficiency; two batch normalization layers are used to stabilize the training process; the activation function and the multi-scale attention module are used to capture feature dependencies at different scales; and a squeeze-and-excitation module is also included for feature recalibration between channels.
[0046] Furthermore, each expert is expected to learn a specific aspect of the data or task; for example, in image processing, one expert may be good at processing the coordinates of the edge region, and another expert may be good at processing the coordinates of the smooth region.
[0047] Furthermore, a gating network can be set up to dynamically assign "importance" weights (weight tensors) to each expert network according to the image tensor. The input of the gating network can be spatial coordinates or other important features of the image tensor, mainly including one or two convolutional layers followed by a Softmax function. Softmax ensures that for each input position (or sample), the sum of the weights of all experts is 1, and these weights can be interpreted as the probability or relative importance of each expert's contribution to the final output.
[0048] Furthermore, the mixture-of-experts network can finally output the final expert feature representation obtained by weighted summation of the outputs of multiple expert networks: ; where, represents the final expert feature representation, represents the spatial coordinates, represents the number of expert networks, represents the index of the expert network, represents the th weight tensor, represents the th output of the expert network, represents element-wise multiplication; where the input of the original expert network will also be retained for subsequent calculation of the loss function of the mixture-of-experts network.
[0049] Furthermore, during the training process, the goal of the model is to minimize the value of this objective function by adjusting its internal parameters (weights and biases); through optimization algorithms such as gradient descent, the model iteratively updates the parameters, making the prediction results closer and closer to the real situation; in this framework, the optimized objective function guides the entire network (including the branch network, backbone network, fusion module, output head, and optionally the GCDD layer (Gaussian Curvature-Driven Diffusion) and gating network parameters) to learn how to recover high-quality images from the corrupted input and make them as close as possible to the original clear images.
[0050] Among them, the total loss is obtained by weighted summation of multiple sub-loss terms. The sub-loss terms include reconstruction loss, geometric loss, and expert diversity loss. During the optimization process, through the Adam (Adaptive Moment Estimation) optimizer or other gradient descent algorithms, the gradients of the model parameters are calculated according to the total loss function and iteratively updated: ; where represents the model parameters of the branch network, feature fusion network, or expert backbone network updated at the -th iteration, represents the model parameters of the branch network, feature fusion network, or expert backbone network updated at the -th iteration, represents the Adam optimizer, represents the learning rate updated at the -th iteration, represents the total loss function updated at the -th iteration, and represent the hyperparameters of the Adam optimizer, represents the gradient operation on the model parameters of the branch network, feature fusion network, or expert backbone network.
[0051] Among them, as Figure 3As shown, it presents a comparison of the image restoration effects of multiple methods in the Denoising task, Inpainting task, and Super-Resolution task in the CelebA dataset (CelebFaces Attributes Dataset). The quantitative comparison of the three groups of tasks in terms of PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index Measure) is shown in Table 1 below: Table 1: Quantitative Comparison Table
[0052] Among them, VDVAE represents Very Deep Variational Autoencoders, DPS represents Data Processing System, PnP-HVAE represents Perspective-n-Point Hierarchical Variational Autoencoder (a framework that combines the geometric modeling ability of PnP and the hierarchical probability of HVAE), and BPnP represents Backpropagation Perspective-n-Point (an improved method for the traditional Perspective-n-Point (PnP) algorithm in the field of computer vision, supporting end-to-end differentiable computing).
[0053] The image results output by each network model disclosed by the present invention have, compared with traditional methods, improved by 5 - 8 dB in terms of PSNR and SSIM metrics, especially when the noise level is greater than 30%; by constraining the distribution difference of expert outputs through Holder divergence, the weight of Holder divergence can be gradually increased when traversing all data in the first 10 times, avoiding model collapse caused by unstable initial training and improving the generalization ability of the model, and the test error is reduced by 12% on the BSD300 dataset (Berkeley Segmentation Dataset and Benchmark); while the curvature function (in addition to the GCDD proposed in this embodiment, other curvature functions can be used instead, such as TSC (Tensor-based Surface Curvature), TRV (Tangential Relative Variation), TAC (Total Absolute Curvature), MS (Mean Squared Curvature), Geman (Generalized Euler-Manifold), Log-det (Logarithm of Determinant Curvature Function), Laplace (Laplace Operator), and EE (Euler-Elastica Curvature Function)) dynamically adjusts the diffusion intensity, suppressing smoothing in high-curvature regions (edges), resulting in a 30% increase in the edge retention rate and suppressing the generation of artifacts; the coordinate-separated backbone network decomposes calculations along the coordinate axes, reducing the dimensional complexity and resulting in a 40% reduction in memory occupancy, making it suitable for processing various high-resolution images.
[0054] Furthermore, the quantitative results of blur recovery on the BSD300 dataset under different noise intensities are shown in Table 2 below: Table 2: Quantitative Results Table of Blur Recovery
[0055] Among them, GS-PNP represents Gaussian Splatting-Perspective-n-Point Joint Optimization, EPLL represents Expected Patch Log-Likelihood (an optimization method based on a probability model), and PnP-MM represents Plug-and-Play Majorization-Minimization.
[0056] The present invention performs calculations through a separable backbone network, combines a mixture-of-experts mechanism, improves the convergence speed in high-dimensional tasks, significantly reduces gradient explosion and initial parameter sensitivity, and can be widely applied to image inverse problem tasks such as image denoising, image inpainting, super-resolution, and motion blur restoration, and can be extended to high-dimensional scenarios such as computational physics and remote sensing reconstruction.
[0057] Furthermore, as Figure 4 shown, based on the above multi-task inverse imaging method based on the learning of the mixture-of-experts collaborative diffusion operator, the present invention also correspondingly provides a multi-task inverse imaging system based on the learning of the mixture-of-experts collaborative diffusion operator, wherein the multi-task inverse imaging system based on the learning of the mixture-of-experts collaborative diffusion operator includes: A preprocessing module 51, configured to obtain a plurality of images to be processed input by a user, and perform normalization processing on all the images to be processed according to the types of all the images to be processed, so as to obtain an image tensor; A feature extraction module 52, configured to input the image tensor into a branch network, the branch network performs adaptive average pooling on the image tensor, generates a plurality of pooled features of the image tensor, and fuses all the pooled features, and outputs a branch feature representation; An image reconstruction module 53, configured to input the branch feature representation into a feature fusion model for dimension adjustment and feature refinement, output a high-dimensional feature representation, and map the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image.
[0058] Furthermore, as Figure 5 shown, based on the above multi-task inverse imaging method and system based on the learning of the mixture-of-experts collaborative diffusion operator, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 5 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0059] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit of the terminal and the external storage device. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code of the installed terminal, etc. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a multi-task inverse imaging program 40 based on the learning of an expert mixture collaborative diffusion operator is stored on the memory 20, and the multi-task inverse imaging program 40 based on the learning of an expert mixture collaborative diffusion operator can be executed by the processor 10, thereby implementing the multi-task inverse imaging method based on the learning of an expert mixture collaborative diffusion operator in this application.
[0060] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing the multi-task inverse imaging method based on the learning of an expert mixture collaborative diffusion operator, etc.
[0061] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other through a system bus.
[0062] In one embodiment, when the processor 10 executes the multi-task inverse imaging program 40 based on the learning of an expert mixture collaborative diffusion operator in the memory 20, the steps of the multi-task inverse imaging method based on the learning of an expert mixture collaborative diffusion operator as described above are implemented.
[0063] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multi-task inverse imaging program based on the learning of an expert mixture collaborative diffusion operator, and when the multi-task inverse imaging program based on the learning of an expert mixture collaborative diffusion operator is executed by a processor, the steps of the multi-task inverse imaging method based on the learning of an expert mixture collaborative diffusion operator as described above are implemented.
[0064] In summary, the present invention provides a multi-task inverse imaging method and related devices based on the learning of an expert mixture collaborative diffusion operator. The method includes: obtaining a plurality of images to be processed input by a user, normalizing all the images to be processed according to the types of all the images to be processed to obtain an image tensor; inputting the image tensor into a branch network, where the branch network performs adaptive average pooling on the image tensor to generate a plurality of pooled features of the image tensor, and fuses all the pooled features to output a branch feature representation; inputting the branch feature representation into a feature fusion model for dimension adjustment and feature refinement to output a high-dimensional feature representation, and mapping the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image. The present invention performs calculations through a separable backbone network, combines a mixture of experts mechanism, improves the convergence speed in high-dimensional tasks, significantly reduces gradient explosion and initial parameter sensitivity, can be widely applied to image inverse problem tasks such as image denoising, image restoration, super-resolution, motion blur recovery, etc., and can be extended to high-dimensional scenarios such as computational physics and remote sensing reconstruction.
[0065] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal including that element.
[0066] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0067] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All these improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A multi-task inverse imaging method based on the learning of an expert mixture collaborative diffusion operator, characterized in that, The multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator includes: Obtain multiple images to be processed input by the user, and perform normalization processing on all the images to be processed according to the types of all the images to be processed, so as to obtain an image tensor; Input the image tensor into a branch network, the branch network performs adaptive average pooling on the image tensor, generates multiple pooling features of the image tensor, and fuses all the pooling features to output a branch feature representation; Input the branch feature representation into a feature fusion model, output a high-dimensional feature representation, and map the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image.
2. The multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator according to claim 1, wherein The step of obtaining multiple images to be processed input by the user, and performing normalization processing on all the images to be processed according to the types of all the images to be processed, so as to obtain an image tensor specifically includes: Obtain all the images to be processed input by the user, and determine the quantity, color, and size of all the images to be processed; If the color of the image to be processed is color, set the number of channels of the corresponding image to be processed to multiple, and if the color of the image to be processed is black and white, set the number of channels of the corresponding image to be processed to a single one; Adjust the resolution of all the images to be processed to a unified standard according to the size of each image to be processed; According to the quantity, all the adjusted resolutions, and all the channel numbers, convert all the images to be processed into a unified tensor format and fuse them to obtain an image tensor, where the image tensor includes the quantity information of the images to be processed.
3. The multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator according to claim 1, wherein, The step of inputting the image tensor into a branch network, the branch network performs adaptive average pooling on the image tensor, generates multiple pooling features of the image tensor, and fuses all the pooling features to output a branch feature representation specifically includes: Input the image tensor into the pyramid module of the branch network, and the pyramid module performs multi-scale adaptive average pooling processing on the image tensor to obtain multiple pooling feature representations; After upsampling all the pooling feature representations to a preset spatial dimension, perform splicing processing on all the pooling feature representations and all the images to be processed, and output a branch feature representation, where the preset spatial dimension is the same as the dimension of the image tensor.
4. The multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator according to claim 3, characterized in that, The step that the pyramid module performs multi-scale adaptive average pooling processing on the image tensor to obtain multiple pooling feature representations specifically includes: Obtain multiple pooling scales specified by the user, and adjust multiple pooling windows of the pyramid module according to all the pooling scales; Perform pooling processing on the image tensor, and output multiple pooling feature representations through all the pooling windows.
5. The multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator according to claim 1, characterized in that, The step of inputting the branch feature representation into a feature fusion model, outputting a high-dimensional feature representation, and mapping the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image specifically includes: Input the branch feature representation into the feature fusion model for dimension adjustment and feature refinement, and output a high-dimensional feature representation; Input the high-dimensional feature representation into an output head, which transforms the high-dimensional feature representation into features matching the target image space and performs mapping to obtain a target reconstructed image; Among them, the number of channels of the transformed high-dimensional feature representation matches the number of channels of the target image space.
6. The multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator according to claim 1, characterized in that, After inputting the branch feature representation into an optimized feature fusion model for dimension adjustment and feature refinement, outputting a high-dimensional feature representation, and mapping the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image, it further includes: Extract the spatial coordinates of the image tensor, and input the spatial coordinates into different types of expert backbone networks respectively, and output corresponding expert feature representations respectively; Input the spatial coordinates into a gating network, and the gating network outputs a weight tensor corresponding to each expert backbone network; Construct a reconstruction loss function and a geometric loss function, and construct an expert diversity loss function according to all the weight tensors and all the expert feature representations, and optimize the constructed feature fusion model according to the reconstruction loss function, the geometric loss function and the expert diversity loss function.
7. The multi-task inverse imaging method based on the learning of the expert mixture collaborative diffusion operator according to claim 6, characterized in that The constructing the reconstruction loss function and the geometric loss function, and constructing the expert diversity loss function according to all the weight tensors and all the expert feature representations, and optimizing the constructed feature fusion model according to the reconstruction loss function, the geometric loss function and the expert diversity loss function specifically includes: Obtain the pixel-level difference between the target reconstructed image and the image to be processed, and construct a reconstruction loss function according to the pixel-level difference: ; Among them, represents the reconstruction loss function, represents the target reconstructed image, represents the image to be processed, represents the total number of pixels, represents the pixel index, represents the th target reconstructed image, represents the th image to be processed; Obtain the curvature matching loss and the gradient matching loss in the curvature-driven diffusion model, and construct a geometric loss function according to the curvature matching loss and the gradient matching loss: ; Among them, represents the geometric loss function, represents the weight of the curvature matching loss, represents the curvature matching loss, represents the weight of the gradient matching loss, represents the gradient matching loss; Construct the weights of the expert backbone networks according to all the weight tensors, and construct an expert diversity loss function according to all the expert feature representations; Construct a total loss function according to the reconstruction loss function, the geometric loss function and the expert diversity loss function, and perform gradient optimization on the branch network, the feature fusion network and all the expert backbone networks according to the total loss function: ; Among them, represents the total loss function, represents the model parameters of the branch network, feature fusion network or expert backbone network, represents the weight of the geometric loss function, represents the weight of the expert backbone network, represents the expert diversity loss function, represents the number of expert backbone networks, represents the expert feature representation output by the th expert backbone network, represents the number index of the expert backbone network.
8. A multi-task inverse imaging system based on the learning of an expert mixture collaborative diffusion operator, characterized in that, The multi-task inverse imaging system based on the learning of an expert mixture collaborative diffusion operator includes: A preprocessing module for obtaining a plurality of images to be processed input by a user, and performing normalization processing on all the images to be processed according to the types of all the images to be processed to obtain an image tensor; A feature extraction module for inputting the image tensor into a branch network, the branch network performing adaptive average pooling on the image tensor to generate a plurality of pooled features of the image tensor, and fusing all the pooled features to output a branch feature representation; An image reconstruction module for inputting the branch feature representation into a feature fusion model for dimension adjustment and feature refinement, outputting a high-dimensional feature representation, and mapping the high-dimensional feature representation to a target image control through an output head to obtain a target reconstructed image.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning stored on the memory and executable on the processor. When the multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning is executed by the processor, the steps of the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning. When the multi-task inverse imaging program based on expert mixture collaborative diffusion operator learning is executed by a processor, the steps of the multi-task inverse imaging method based on expert mixture collaborative diffusion operator learning according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Image denoising method based on double-branch feature fusion
CN119313583A
Image reconstruction method and system based on hybrid network framework, terminal and storage medium
CN119963682A
Cited By
Multi-modal data prediction method and device based on hybrid expert attention network
CN120974259A
Microscopic denoising method based on multi-expert discrimination
CN121724861A