Quantum image processing system and method based on mbconv-ssm cooperation

CN122821153APending Publication Date: 2026-09-25CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610910456.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

由于卷积神经网络主要通过局部卷积操作提取特征,对图像中远距离像素或区域之间的关联信息建模能力有限,在处理高分辨率医学图像时,难以兼顾全局语义信息与局部边界细节,容易造成细微结构遗漏、目标边界分割不清晰等情况;基于自注意力机制的模型在图像分辨率较高时需要计算大量特征位置之间的相关性,尤其在处理高分辨率3T MRI图像时,其计算复杂度和存储开销较大,导致训练成本高、推理速度慢、部署难度大

Benefits of technology

本发明通过在U-Net各级中将MBConv模块与Mamba的SSM层顺序连接,形成局部特征提取与全局依赖建模相结合的基础单元,并在瓶颈层设置SSM路径与量子注意力路径的双路协同全局建模结构,使模型在保留局部细节信息的同时增强对复杂结构区域和长距离关联信息的表征能力;其中,SSM层通过状态空间建模实现对全局依赖信息的有效提取,量子启发注意力机制通过提高特征关联表征效率,减少高分辨率图像处理中大规模位置相关性计算带来的开销,从而改善现有图像处理方法中细微结构易遗漏、模糊边界分割不清以及高分辨率图像处理效率较低的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821153A_ABST
    Figure CN122821153A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer vision and image processing, and proposes a quantum image processing system and method based on MBConv-SSM cooperation, in which, under a U-Net framework, an MBConv module and a selective state space model of Mamba are sequentially coupled in a hierarchical manner, a double-path cooperative global modeling structure of a quantum multi-head attention path and an SSM path is constructed in a bottleneck layer, through the multi-level fusion design, cooperative unification of local detail extraction, long-distance dependence modeling and efficient global correlation representation is realized, and therefore, image segmentation precision, boundary detail recovery capability and calculation efficiency are considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and image processing technology, specifically to an MBConv-SSM collaborative quantum image processing system and method. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Current image processing models primarily employ convolutional neural networks (CNNs) and Transformer architectures. U-Net, a CNN model, with its encoder-decoder symmetric structure and skip connections, effectively integrates multi-scale features and is widely used in medical image processing. Transformers, on the other hand, utilize self-attention mechanisms to enhance long-range dependency modeling capabilities. With the increasing clinical application of high-resolution, multimodal medical images such as 3T MRI, the demand for representing subtle lesions, complex boundaries, and long-distance tissue associations in images has significantly increased. Traditional single-architecture approaches are no longer sufficient to simultaneously address local detail extraction, global context modeling, computational efficiency, and deployment performance.

[0004] Existing image processing methods primarily employ convolutional neural network (CNN) models or models based on self-attention mechanisms. Since CNNs mainly extract features through local convolution operations, their ability to model the correlation information between distant pixels or regions in an image is limited. When processing high-resolution medical images, they struggle to balance global semantic information with local boundary details, easily leading to the omission of subtle structures and unclear target boundary segmentation. Self-attention models, on the other hand, require calculating the correlations between numerous feature locations at high image resolutions. Especially when processing high-resolution 3T MRI images, their computational complexity and storage overhead are significant, resulting in high training costs, slow inference speeds, and difficult deployment. Therefore, existing image processing methods suffer from insufficient long-distance dependency modeling, poor segmentation of complex boundaries and subtle structures, and low processing efficiency. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes an MBConv-SSM collaborative quantum image processing system and method. Within the U-Net framework, the MBConv module is hierarchically coupled with Mamba's selective state-space model, and a dual-path collaborative global modeling structure of quantum multi-head attention path and SSM path is constructed at the bottleneck layer. This multi-level fusion design achieves a synergistic unification of local detail extraction, long-distance dependency modeling, and efficient global correlation representation, thus balancing image segmentation accuracy, boundary detail recovery capability, and computational efficiency.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: One or more embodiments provide an MBConv-SSM-based collaborative quantum image processing system, comprising: The preprocessing module is used to preprocess the input image and then input the preprocessed image into the encoder. The encoder includes a multi-level cascaded downsampling unit. Each downsampling unit includes a sequentially connected MBConv module and a selective state space model layer (SSM layer). The MBConv module is used to extract local features from the input features, and the SSM layer is used to perform global dependency modeling on the features processed by the MBConv module and output deep features. The bottleneck layer includes a parallel quantum multi-head attention path and an SSM enhancement path. The outputs of the quantum multi-head attention path and the SSM enhancement path are fused to obtain the enhanced features. The decoder is used to upsample the enhanced features and fuse the skip connection features of the corresponding level of the encoder for feature recovery. The recovered features are then classified to obtain the image processing result.

[0007] A further technical solution is that the encoder includes multi-layer cascaded downsampling units, each downsampling unit including sequentially connected MBConv modules and a selective state space model layer; Each MBConv module includes a sequentially connected inverted residual structure, a depthwise separable convolutional layer, a Swish activation function layer, and a batch normalization layer, which are used to implement depthwise separable convolution, function activation, and batch normalization operations on the image to obtain a local feature map. The selective state space model layer receives the feature sequence formed by expanding the local feature map along the spatial dimension. Based on the current input feature sequence, it updates the hidden state to enhance important information in the feature sequence and suppress redundant information, thereby outputting sequence features containing long-distance dependency information.

[0008] Further technical solutions are available in the MBConv module: The inverse residual structure is used to perform dimensionality upscaling, convolutional feature extraction, and dimensionality reduction compression on the input image features, and retains the original feature information through residual connections to obtain the inverse residual output features. Depth-separable convolutional layers are used to perform spatial convolution on each channel of the inverse residual output feature to extract local spatial features. Then, pointwise convolution is used to achieve feature fusion between channels and adjustment of the number of channels to obtain the features after depth-separable convolution. Swish activation function layer is used to perform non-linear mapping on features after depthwise separable convolution; The batch normalization layer is used to normalize the data after nonlinear mapping to obtain the extracted local features.

[0009] A further technical solution involves a bottleneck layer positioned between the encoder and decoder. This layer is used for global context enhancement of the deep features output by the encoder. It includes a parallel quantum multi-head attention path and a SSM enhancement path, as well as a feature fusion unit positioned after the quantum multi-head attention path and the SSM enhancement path. Quantum multi-head attention path is used to perform global correlation modeling on deep features of the input; deep features are the output features of the last stage of the encoder; SSM enhancement path, used for state-space modeling of deep features of the same input; The feature fusion unit is used to fuse the features output by the quantum multi-head attention path and the SSM enhancement path to obtain the enhanced bottleneck layer features, which are then output to the decoder.

[0010] A further technical solution is that the quantum multi-head attention path includes a feature projection unit, a complex mapping unit, a quantum correlation calculation unit, a weighting unit, and an output projection unit connected in sequence. The feature projection unit is used to perform a linear transformation on the input deep features through a learnable linear layer to generate query features Q, key features K, and value features V. The complex mapping unit is used to split the query feature Q and the key feature K into real and imaginary parts, or to perform a preset transformation on the query feature Q and the key feature K to map the query feature Q and the key feature K to the complex vector space to obtain the complex feature representation; The quantum correlation computing unit is used to calculate the correlation between query feature Q and key feature K based on complex inner product, and to apply unitary matrix transformation to key feature K and / or query feature Q to enhance the high-dimensional correlation expression between features and obtain attention correlation results. The weighting unit is used to perform normalization processing on the feature representation in the form of constructed complex numbers output by the quantum correlation computing unit, generate attention weights, and perform weighted summation on the value feature V according to the attention weights to obtain the weighted attention feature; The output projection unit is used to map the weighted attention features from complex or high-dimensional representation back to the real-valued feature space through a linear projection layer, thereby obtaining the quantum multi-head attention path output features.

[0011] A further technical solution is a decoder, which is set after the bottleneck layer and includes a multi-level cascaded upsampling unit. Each upsampling unit includes an upsampling layer, a skip connection fusion unit, and a feature extraction module. The upsampling layer is used to restore the spatial size of the input feature map to obtain upsampled features. The skip connection fusion unit is used to fuse upsampled features with features output by the downsampled unit of the corresponding level of the encoder. The feature extraction module is used to perform convolution processing on the features fused by the skip connection fusion unit to obtain the current level decoding features, and output them to the next level upsampling unit or output layer.

[0012] A further technical solution involves connecting an output layer after the decoder, which includes a lightweight MBConv module, a channel mapping layer, and a probability output layer. The lightweight MBConv module is used to extract features from the recovered features output by the decoder; The channel mapping layer is used to map the number of feature channels extracted by the lightweight MBConv module to the number of target categories; The probability output layer is used to normalize the number of target categories after mapping in order to generate a pixel-level probability map.

[0013] One or more embodiments provide a method for quantum image processing based on MBConv-SSM cooperative quantum imaging, comprising the following steps: Preprocess the acquired input image; The preprocessed image is input into the encoder, and features are extracted step by step through a multi-level downsampling unit. Each downsampling unit performs the following steps in sequence: local feature extraction is performed on the input features using the MBConv module; global dependency modeling is performed on the features after local feature extraction using the selective state space model (SSM) layer. The deep features output by the encoder are input into the bottleneck layer, and global modeling is performed in parallel through the quantum multi-head attention path and the SSM enhancement path. The outputs of the quantum multi-head attention path and the SSM enhancement path are fused to obtain the enhanced features; The enhanced features are input to the decoder for upsampling and recovery, and then combined with the skip connection features of the corresponding level of the encoder for feature fusion to obtain the recovered features. These features are then mapped to categories to obtain the image processing result.

[0014] A further technical solution involves globally associating deep features of the input through a quantum multi-head attention path, including the following steps: The deep features of the input are linearly transformed by a learnable linear layer to generate query features Q, key features K, and value features V. The query feature Q and the key feature K are split into real and imaginary parts, or a preset transformation is performed on the query feature Q and the key feature K to map the query feature Q and the key feature K to the complex vector space to obtain the complex feature representation; Based on complex feature representation, the correlation between query feature Q and key feature K is calculated using complex inner product, and unitary matrix transformation is applied to key feature K and / or query feature Q to enhance the high-dimensional association expression between features, thus obtaining attention association results; The attention association results are normalized to generate attention weights, and the value features V are weighted and summed according to the attention weights to obtain the weighted attention features. The weighted attention features are mapped back to the real-valued feature space to obtain the quantum multi-head attention path output features.

[0015] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps described above in the MBConv-SSM-based collaborative quantum image processing method.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention connects the MBConv module and the SSM layer of Mamba sequentially in each stage of U-Net to form a basic unit that combines local feature extraction and global dependency modeling. At the bottleneck layer, a dual-path collaborative global modeling structure of SSM path and quantum attention path is set up, which enables the model to enhance its ability to represent complex structural regions and long-distance correlation information while retaining local detail information. Among them, the SSM layer realizes the effective extraction of global dependency information through state space modeling, and the quantum-inspired attention mechanism improves the efficiency of feature correlation representation and reduces the overhead of large-scale position correlation calculation in high-resolution image processing. This improves the problems of easy omission of fine structures, unclear segmentation of blurred boundaries and low efficiency of high-resolution image processing in existing image processing methods.

[0017] The advantages of the present invention, as well as its additional advantages, will be described in detail in the following specific embodiments. Attached Figure Description

[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.

[0019] Figure 1 This is a system block diagram of the MBConv-SSM collaborative quantum image processing system according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the bottleneck layer structure of Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the experimental results of the UNet++ model compared to Embodiment 1 of the present invention; Figure 4This is a schematic diagram of the experimental results of the MBConv-SSM cooperative quantum model in Embodiment 1 of the present invention; Figure 5 This is a schematic diagram of the experimental results of the UNet model compared to Embodiment 1 of the present invention. Detailed Implementation

[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0021] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0022] It should be noted that the terminology used herein is for describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features within those embodiments can be combined with each other. The embodiments will now be described in detail with reference to the accompanying drawings.

[0023] Example 1 In one or more of the technical solutions disclosed in the embodiments, such as Figures 1 to 5 As shown, an image network architecture based on the MBConv-SSM collaborative quantum image processing system is constructed using the U-Net encoder-decoder framework, which deeply integrates the MBConv module of Mamba and the quantum multi-head attention mechanism to realize image processing. The architecture includes a preprocessing module, an encoder, a bottleneck layer, and a decoder. The preprocessing module is used to preprocess the input image and then input the preprocessed image into the encoder. The encoder includes a multi-level cascaded downsampling unit. Each downsampling unit includes a sequentially connected MBConv module and a selective state space model layer (SSM layer). The MBConv module is used to extract local features from the input features, and the SSM layer is used to perform global dependency modeling on the features processed by the MBConv module and output deep features. The bottleneck layer includes a parallel quantum multi-head attention path and an SSM enhancement path. The outputs of the quantum multi-head attention path and the SSM enhancement path are fused to obtain the enhanced features. Among them, the quantum multi-head attention path is used to perform attention modeling on the deep features output by the encoder, and the SSM enhancement path is used to perform state-space modeling on the deep features. The decoder is used to upsample the enhanced features and fuse the skip connection features of the corresponding level of the encoder for feature recovery. The recovered features are then classified to obtain the image processing result.

[0024] In this embodiment, by sequentially connecting the MBConv module with the SSM layer of Mamba in each stage of U-Net, a basic unit combining local feature extraction and global dependency modeling is formed. A dual-path collaborative global modeling structure of SSM path and quantum attention path is set in the bottleneck layer, which enables the model to enhance the representation ability of complex structural regions and long-distance correlation information while retaining local detail information. Among them, the SSM layer realizes the effective extraction of global dependency information through state space modeling, and the quantum-inspired attention mechanism improves the efficiency of feature correlation representation and reduces the overhead of large-scale position correlation calculation in high-resolution image processing, thereby improving the problems of easy omission of fine structures, unclear segmentation of blurred boundaries and low efficiency of high-resolution image processing in existing image processing methods.

[0025] In some embodiments, the preprocessing module preprocesses the input image, including a normalization unit, a size adjustment unit, and a channel stitching unit; The normalization unit is used to scale the pixel values ​​of the input image to normalize the pixel values ​​to the [0,1] range; The size adjustment unit is used to adjust the resolution of the input image to uniformly adjust the input image to a preset size, preferably 256×256 pixels; The channel stitching unit is used to stitch different modal images along the channel dimension when the input image is a multimodal 3T MRI image to form multi-channel input data. The number of channels in the multi-channel input data is preferably 3 or 4.

[0026] In some embodiments, the encoder is a downsampling path that includes multiple cascaded downsampling units. Each downsampling unit includes sequentially connected MBConv modules and a selective state space model (SSM) layer. Each MBConv module includes sequentially connected inverse residual structure, depthwise separable convolutional layer, Swish activation function layer, and batch normalization layer, which are used to implement depthwise separable convolution, function activation, and batch normalization operations on the image to obtain a local feature map. The Selective State Space Model (SSM) layer can adopt Mamba's Selective State Space Model (SSM); it is used to receive the feature sequence formed by expanding the local feature map along the spatial dimension, update the hidden state based on the current input feature sequence to enhance the important information in the feature sequence and suppress the redundant information, thereby outputting the sequence feature containing long-distance dependency information; the output sequence feature is recombined to form a feature map and passed to the next level downsampling unit.

[0027] Mamba is a Selective State Space Model (SelectiveSSM) architecture, which is essentially a type of neural network method for sequence modeling, mainly used to model long-range dependency information with low computational complexity.

[0028] Specifically, in the MBConv module: The inverse residual structure is used to perform dimensionality upscaling, convolutional feature extraction, and dimensionality reduction compression on the input image features, and retains the original feature information through residual connections to obtain the inverse residual output features. Depthwise separable convolutional layers are used to perform spatial convolution on each channel of the inverse residual output feature to extract local spatial features. Then, pointwise convolution is used to achieve feature fusion between channels and adjustment of the number of channels, thereby completing feature extraction while reducing the number of parameters and computational complexity, and obtaining the features after depthwise separable convolution processing. The Swish activation function layer performs a non-linear mapping on the features processed by depthwise separable convolutions to enhance the network's ability to express complex feature distributions while preserving relatively continuous gradient change characteristics. The batch normalization layer is used to normalize the data after nonlinear mapping to obtain the extracted local features; this reduces the fluctuation of the input data distribution of each layer, speeds up network training convergence, and improves the stability of the model training process. Furthermore, a downsampling layer is set after each downsampling unit. The downsampling layer preferably uses a 3×3 convolution with a stride of 2 to achieve a 2x compression of the feature map space size. As the encoding process deepens, the number of feature channels increases sequentially, preferably to 64, 128, 256, 512, and 1024, to improve the expressive power of deep semantic features.

[0029] In this embodiment, the encoder extracts multi-level features from the input image and gradually compresses the spatial size through downsampling to capture semantic information. For the high detail requirements of 3T MRI images (such as the minute boundaries of brain tumors), this embodiment uses a 5-layer cascaded improved MBConv module as the downsampling unit. Each downsampling unit deeply integrates convolution and Mamba's Selective State Space Model (SSM). Through this sequential coupling structure, a processing method of local convolutional feature extraction and global sequence modeling is formed in each level of the encoding unit, thereby enabling the encoder to simultaneously possess local detail representation capabilities and long-distance dependency modeling capabilities.

[0030] In some embodiments, a bottleneck layer, positioned between the encoder and decoder, is used to perform global context enhancement on the deep features output by the encoder. This includes a parallel quantum multi-head attention path and a selective state-space model enhancement path (SSM enhancement path), and a feature fusion unit positioned after the quantum multi-head attention path and the SSM enhancement path, wherein: Quantum multi-head attention path is used to perform global correlation modeling on deep features of the input; deep features are the output features of the last stage of the encoder; SSM enhancement path, used for state-space modeling of deep features of the same input; The feature fusion unit is used to fuse the features output by the quantum multi-head attention path and the SSM enhancement path to obtain the enhanced bottleneck layer features, which are then output to the decoder.

[0031] A further technical solution is that the quantum multi-head attention path includes a feature projection unit, a complex mapping unit, a quantum correlation calculation unit, a weighting unit, and an output projection unit connected in sequence. The feature projection unit is used to perform a linear transformation on the input deep features through a learnable linear layer to generate query features Q, key features K, and value features V. The complex mapping unit is used to split the query feature Q and the key feature K into real and imaginary parts, or to perform a preset transformation on the query feature Q and the key feature K to map the query feature Q and the key feature K to the complex vector space to obtain the complex feature representation; Among them, the feature projection unit and the complex mapping unit realize state preparation and feature projection; The quantum correlation computing unit is used to calculate the correlation between query feature Q and key feature K based on complex inner product, and to apply unitary matrix U transformation to key feature K and / or query feature Q to enhance the high-dimensional correlation expression between features and obtain attention correlation results. The weighting unit is used to perform normalization processing on the feature representation in the form of constructed complex numbers output by the quantum correlation computing unit, generate attention weights, and perform weighted summation on the value feature V according to the attention weights to obtain the weighted attention feature; Optionally, the softmax function can be used to normalize the attention association results.

[0032] The output projection unit is used to map the weighted attention features from complex or high-dimensional representation back to the real-valued feature space through a linear projection layer, thereby obtaining the quantum multi-head attention path output features.

[0033] In the above embodiments, by employing complex space mapping and unitary matrix transformation, the quantum multi-head attention path no longer relies solely on dot product operations between real-valued features to characterize the correlation between positions. Instead, it can depict the complex nonlinear relationships between input features in a higher-dimensional feature representation space, thereby improving the ability to characterize the correlation information between distant pixels, cross-regional tissues, and complex structural boundaries. At the same time, since this path achieves global information interaction through projection, complex mapping, and correlation calculation after unitary transformation of input features, compared with the traditional dot product attention method of directly calculating pairwise large-scale positional relationships, it helps to reduce the computational burden in the global modeling process of high-resolution images. This improves the problems of insufficient global correlation modeling ability, low accuracy of complex boundary segmentation, and large computational overhead in existing image processing methods when processing high-resolution medical images.

[0034] A further technical solution, the SSM enhancement path, can adopt Mamba's selective state-space model; The SSM enhancement path is used to receive the same deep input features as the quantum multi-head attention path, perform state space modeling on the deep input features, update the hidden state based on the current input features, and extract long-range dependency information from the deep input features to obtain enhanced features containing long-range dependency information.

[0035] The feature fusion unit is used to fuse the output features of the quantum multi-head attention path and the output features of the SSM enhancement path to obtain fused features; and to perform layer normalization and residual connection processing on the fused features to obtain enhanced bottleneck layer features, and then pass the enhanced bottleneck layer features to the decoder.

[0036] Optionally, the output features of the quantum multi-head attention path and the output features of the SSM enhancement path can be fused by element-wise addition.

[0037] In this embodiment, the bottleneck layer receives the deepest features output by the encoder and performs dual-path parallel global modeling processing on these deep features. One path uses a quantum multi-head attention path to sequentially process the input deep features through feature projection, complex mapping, correlation calculation, weight allocation, and output projection to obtain attention features with global correlation information. The other path uses an SSM enhancement path to perform state-space modeling on the same input deep features to obtain enhanced features containing long-distance dependency information. Then, the output features from both paths are fused into a feature fusion unit, and after layer normalization and residual connection processing, the enhanced bottleneck layer features are output to the decoder. Through this dual-path parallel structure, the bottleneck layer can perform global context enhancement on the deep features output by the encoder from different modeling perspectives, thereby improving the representation ability of complex structural regions, long-distance correlated regions, and ambiguous boundary regions.

[0038] In some embodiments, the decoder, as an upsampling path, is set after the bottleneck layer and includes multi-level cascaded upsampling units, each upsampling unit including an upsampling layer, a skip connection fusion unit, and a feature extraction module. The upsampling layer is used to restore the spatial size of the input feature map to obtain upsampled features. The skip connection fusion unit is used to fuse upsampled features with features output by the downsampled unit of the corresponding level of the encoder. The feature extraction module is used to perform convolution processing on the features fused by the skip connection fusion unit to obtain the current level decoding features, and output them to the next level upsampling unit or output layer.

[0039] Specifically, the upsampling layer can be a transposed convolutional layer, used to upsample the feature map output from the bottleneck layer or the previous upsampling unit, so as to achieve the step-by-step recovery of the spatial resolution of the feature map. In one specific implementation, the upsampling layer preferably uses a transposed convolution with a kernel size of 4×4 and a stride of 2 to achieve a 2-fold magnification of the feature map spatial size; The skip connection fusion unit can use a feature splicing method to splice the shallow features output by the encoder at the corresponding level with the upsampled features of the decoder along the channel dimension to form fused features; One possible implementation is a feature extraction module, which may include a first convolutional layer, a first activation layer, a first batch normalization layer, a second convolutional layer, a second activation layer, and a second batch normalization layer connected in sequence, for performing convolution processing on the fused features to obtain the extracted decoded features.

[0040] Optionally, the first and second convolutional layers can be 3×3 convolutional layers, and the first and second activation layers can use the ReLU activation function; the first batch normalization layer and the second batch normalization layer preferably use batch normalization processing to perform local feature reconstruction and distribution adjustment on the fused features, thereby improving the stability and effectiveness of feature recovery in the decoding stage.

[0041] As the decoding process proceeds step by step, the spatial resolution of the feature maps output by each upsampling unit gradually increases, while the number of channels decreases step by step, so that deep semantic information is gradually mapped back to high-resolution feature maps in the process of restoring spatial details.

[0042] The aforementioned decoder receives the enhanced features output from the bottleneck layer and gradually restores the spatial size of the feature map through multi-level upsampling units. Specifically, the enhanced features output from the bottleneck layer are first input into the first-level upsampling unit, where their spatial size is enlarged by the upsampling layer to obtain upsampled features. Then, the upsampled features are input into the skip connection fusion unit and fused with the features output from the corresponding level of the encoder to introduce the shallow edge information, texture information, and local structural information retained in the encoding stage into the current decoding stage. Afterward, the fused features are input into the feature extraction module, where convolution, activation, and batch normalization processes are used to integrate the local detail information and deep semantic information in the fused features to obtain the current-level decoded features. The current-level decoded features are then passed to the next-level upsampling unit for further processing. This process is repeated level by level in each decoding unit, gradually restoring the spatial resolution of the feature map and ultimately outputting restored features that combine global semantic information and local boundary detail information.

[0043] In this embodiment, by setting up an upsampling layer, a skip connection fusion unit, and a feature extraction module in the decoder, the deep enhancement features from the bottleneck layer can be fused with the shallow detail features output by the corresponding layer of the encoder while restoring spatial resolution. This improves the problem of blurred boundaries and missing fine structures that are easily caused when decoding relies solely on deep semantic features. Furthermore, by performing convolutional extraction on the fused features, the ability to restore the edge, texture, and local structural information of the target region can be further enhanced, thereby improving the boundary clarity and structural integrity of the image processing results.

[0044] A further technical solution involves connecting an output layer after the decoder, which includes a lightweight MBConv module, a channel mapping layer, and a probability output layer. The lightweight MBConv module is used to further extract features from the recovered features output by the decoder; The channel mapping layer is used to map the number of feature channels extracted by the lightweight MBConv module to the number of target categories; the channel mapping layer can be a 1×1 convolutional layer.

[0045] The probability output layer is used to normalize the number of target categories after mapping in order to generate a pixel-level probability map; optionally, the probability output layer can use the Softmax function to normalize the number of target categories.

[0046] In this embodiment, by setting a lightweight MBConv module in the output layer, the features recovered by the decoder can be further integrated and enhanced before the final output, thereby improving the problems of unclear boundary transitions and insufficient class distinction that may exist when directly performing class mapping; by setting a 1×1 convolutional layer for channel mapping, class dimension transformation can be achieved while maintaining the spatial resolution; by generating a pixel-level probability map through the Softmax function, the probability of each pixel belonging to different classes can be output, thereby improving the discriminability of the segmentation results and the standardization of the output format.

[0047] To illustrate the image processing performance of the network model in this embodiment, a comparative experiment was conducted. Figures 3 to 5 The figures shown are the performance metrics of the Unet, Unet++, and models from this embodiment compared in a comparative experiment. The comparative experiment used a 3T MRI dataset and employed quantum feature compression and spatial resolution downsampling dimensionality reduction methods. During testing, a composite loss function, BCE_Dice, was used, comprising binary cross-entropy (BCE) and Dice loss.

[0048] The cross-entropy loss can be expressed as: ; in, This represents the true label (0 or 1) of the i-th pixel. To predict the probability that the i-th pixel is the target for the model. This represents the total number of pixels in a single image.

[0049] The Dice loss can be expressed as: ; in, It is a smoothing factor to prevent the denominator from being zero.

[0050] Dice loss is used to evaluate the spatial overlap between the predicted mask P and the real mask Y. It is essentially a continuously differentiable approximation of the F1 score and is sensitive to the overall consistency of the foreground (tumor region) pixels.

[0051] The final composite loss function is expressed as: ; in, =0.01 is the weighting coefficient.

[0052] The optimization method uses the Adam optimizer with default parameters and dynamic learning rate adjustment. When the validation loss does not decrease for two consecutive epochs, the learning rate is multiplied by 0.2 and set to 0.0001.

[0053] The model loss value, Iou, Dice, and accuracy are used as evaluation criteria for model performance, and the 95% Hausdorff coefficient of the model in this embodiment is displayed.

[0054] like Figures 3 to 5 As shown, Figure 3 The experiment showcases various experimental metrics of the UNet++ comparison model, demonstrating the training dynamics and performance of the traditional UNet++ model, serving as a benchmark for medium-to-high complexity comparisons. While UNet++ improves multi-scale feature fusion through dense skip connections and nested structures, it still suffers from insufficient global dependency modeling, leading to blurred boundaries and missed subtle lesions on high-resolution 3T MRI. Experiments Figure 3 The solid blue line represents the training loss (BCE_Dice composite loss), reflecting the model's fit on the training set; a lower value is better. The dashed orange line represents the validation loss (Val Loss), reflecting the model's generalization ability on the validation set and used to monitor overfitting. The solid green line represents the Dice score, evaluating the spatial overlap between the predicted segmentation mask and the ground truth mask; the value is 0-1, with higher values ​​being better, especially sensitive to foreground tumor regions. The solid purple line represents the IoU (Intersection over Union, Jaccard index), evaluating the accuracy of the overlap of segmented regions; the value is 0-1, with higher values ​​being better. The solid red line represents the precision, evaluating the proportion of truly positive pixels among those predicted as positive, avoiding false positives.

[0055] Figure 4 The graphs show the training curves of the model in this embodiment on the same 3T MRI dataset, highlighting the superior convergence, higher accuracy, and stronger boundary recovery ability of the model in this embodiment. Through sequential coupling of MBConv-SSM and dual-path collaboration between bottleneck layer quantum multi-head attention and SSM, global context modeling is more efficient, significantly outperforming traditional models on high-resolution images; the blue solid line represents the training loss, which decreases faster and more smoothly. Figure 4The orange dashed line represents the validation loss, which is the lowest and most stable, reflecting the ability to resist overfitting; the green solid line represents the Dice coefficient, which reaches the highest value among all comparison models and has the fastest curve rise; the purple solid line represents the IoU, the highest overlap accuracy; the red solid line represents the accuracy, which is the highest among all comparison models and has the smallest fluctuation; the dark blue thick solid line represents the 95% Hausdorff distance coefficient (HD95), which measures the maximum error of the segmentation boundary, in pixels or mm. The lower the value, the better, reflecting the ability to recover boundary details. Other models do not show this indicator to highlight the advantages of the model in this embodiment.

[0056] Figure 5 The model used as a benchmark demonstrates the limitations of the traditional U-Net, whose structure consists of an encoder-decoder architecture with simple skip connections. On high-resolution, multimodal 3T MRI, local convolutions struggle to capture long-range dependencies, resulting in slow loss reduction, low Dice / IoU, and blurred boundaries. Figure 5 The solid blue line represents the training loss, which decreases slowly and may plateau later; the dashed orange line represents the validation loss, which is higher and fluctuates more; the solid green line represents the Dice coefficient, which is lower than UNet++ and this invention; the solid purple line represents the IoU, which is lower; and the solid red line represents the accuracy, which is the worst performance among all models.

[0057] pass Figure 3 , Figure 4 , Figure 5 A side-by-side comparison clearly demonstrates that the model in this embodiment converges faster and is more stable during training: Figure 4 The loss curve (Train / Val) has the steepest descent slope and reaches its lowest point earliest, while Figure 5 (UNet) slowest decline Figure 3 While UNet++ offers improvements, it still experiences a significant plateau. This invention significantly reduces training costs through efficient local extraction using MBConv, long-range dependency modeling using SSM, and global association using quantum attention.

[0058] Leading in segmentation accuracy across the board: Dice and IoU metrics Figure 4 The highest accuracy, with an improvement of 5% to 10% or more, indicates optimal overlap of the foreground area, i.e., the tumor or lesion. Figure 4 Highest, fewest false positives. Figure 3 Superior Figure 5 But still lagging behind Figure 4 This indicates that dense hop connections cannot completely replace the global modeling of the SSM+ quantum mechanism.

[0059] It has the strongest ability to restore boundary details: Figure 4The displayed HD95 curve, with the lowest value, demonstrates that the present invention significantly outperforms the U-Net series in characterizing subtle lesions, complex boundaries, and long-distance tissue associations. Traditional models are prone to "unclear boundary segmentation and omission of fine structures," while the present invention effectively solves this problem through quantum complex mapping + unitary matrix transformation + SSM dual-path synergy.

[0060] Overall efficiency and robustness, under the same 3T MRI dataset and composite loss, Figure 4 The curve exhibits minimal fluctuations and the strongest generalization ability, demonstrating that this embodiment balances image segmentation accuracy, boundary detail restoration capability, and computational efficiency.

[0061] Figure 3 and Figure 5 As a benchmark, it highlights the limitations of traditional models: good local details but weak global dependencies and high computational cost. Figure 4 The intuitive demonstration shows that the MBConv-SSM cooperative quantum model in this embodiment achieves the best performance on all metrics.

[0062] Example 2 Based on the MBConv-SSM cooperative quantum image processing system described in Example 1, this example provides an MBConv-SSM cooperative quantum image processing method, including the following steps: Step S1: Preprocess the acquired input image; Step S2: Input the preprocessed image into the encoder and extract features step by step through multi-level downsampling units. Each downsampling unit performs the following steps in sequence: extracting local features from the input features using the MBConv module; and modeling the global dependency of the features after local feature extraction using the Selective State Space Model (SSM) layer. Step S3: Input the deep features output by the encoder into the bottleneck layer, and perform global modeling in parallel through the quantum multi-head attention path and the SSM enhancement path; Specifically, the deep features output by the encoder are input into the bottleneck layer and executed in parallel within the bottleneck layer: attention modeling of the deep features is performed through a quantum multi-head attention path; state space modeling of the deep features is performed through an SSM enhancement path. Step S4: Fuse the output of the quantum multi-head attention path and the output of the SSM enhancement path to obtain the enhanced features; Step S5: Upsample and restore the enhanced features input to the decoder, and combine them with the skip connection features of the corresponding level of the encoder to perform feature fusion, obtain the restored features, and perform category mapping to obtain the image processing result.

[0063] This embodiment preprocesses the input image, followed by stepwise local feature extraction and global dependency modeling by the encoder, dual-path parallel global modeling of the quantum multi-head attention path and SSM enhancement path in the bottleneck layer, dual-path output fusion, and upsampling restoration and cross-layer feature fusion by the decoder. This enables the restoration of image spatial resolution while taking into account local details and global contextual information, thereby improving the representation ability of complex structural regions, long-distance dependent regions, and blurred boundary regions. It also addresses the problems of easy omission of fine structures, unclear boundary segmentation, and insufficient overall segmentation accuracy in existing image segmentation methods.

[0064] In step S1, the input image is preprocessed, including: Step S11, Normalization Processing: The pixel values ​​of the input image are scaled to normalize the pixel values ​​to the [0,1] range; Step S12, Size Adjustment Process: Adjust the resolution of the input image to uniformly adjust the input image to a preset size, preferably 256×256 pixels. When the input image is a multimodal 3T MRI image, the images of different modalities are stitched together along the channel dimension to form multi-channel input data. The number of channels in the multi-channel input data is preferably 3 or 4.

[0065] Step S2, which uses the MBConv module to extract local features from the input features, includes the following steps: Step S21: Perform dimensionality upscaling, convolutional feature extraction, and dimensionality reduction compression on the input image features, and perform residual connection to obtain inverse residual output features; Step S22: Perform spatial convolution on each channel of the inverted residual output feature to extract local spatial features, and then perform pointwise convolution to achieve feature fusion between channels and adjustment of the number of channels to obtain the feature after depth separable convolution processing. Step S23: Perform nonlinear mapping on the features after depthwise separable convolution processing; Step S24: Normalize the data after nonlinear mapping to obtain the extracted local features; The above steps involve sequentially performing dimensionality upscaling, convolutional feature extraction, dimensionality reduction compression, and residual connection processing on the input features. Further, depthwise separable convolution, nonlinear mapping, and normalization processing are performed. This allows for the full extraction of edge, texture, and local structural information from the input features while controlling the number of parameters and computational complexity. This improves the expressive power and stability of local features, thereby providing a more effective feature foundation for subsequent global dependency modeling and addressing the problems of insufficient local detail representation and easy omission of fine structures in existing image processing methods.

[0066] Step S2, which utilizes the Selective State-Space Model (SSM) layer to perform global dependency modeling on the features extracted from local features, includes the following steps: Step S201: The feature map after local feature extraction is serialized and expanded according to the spatial dimension to obtain the input feature sequence; Step S202: Use the Selective State Space Model (SSM) layer to perform state space modeling on the input feature sequence; during the state space modeling process, update the hidden state based on the current input feature sequence to obtain the output sequence features containing long-distance dependency information; Step S203: Reconstruct the output sequence features into a feature map and pass it to the next level downsampling unit.

[0067] In the above steps, by expanding the feature map after local feature extraction into a feature sequence according to the spatial dimension, and using the Selective State Space Model (SSM) layer to perform state space modeling on the feature sequence, the hidden state can be updated based on the current input feature sequence during the feature sequence transmission process, thereby extracting long-distance dependency information between different spatial locations. After being reconstructed into a feature map, the long-distance dependency information is introduced into the subsequent downsampling process, thereby improving the representation ability of complex structural regions, long-distance associated regions, and overall semantic information, and improving the problems of insufficient global dependency modeling and insufficient understanding of complex regions in existing image processing methods.

[0068] In step S3, global correlation modeling of the input deep features is performed through the quantum multi-head attention path, including the following steps: Step S311: Perform a linear transformation on the input deep features through a learnable linear layer to generate query features Q, key features K, and value features V; Step S312: Decompose the query feature Q and key feature K into real and imaginary parts, or perform a preset transformation on the query feature Q and key feature K to map the query feature Q and key feature K into a complex vector space to obtain a complex feature representation; Step S313: Based on complex feature representation, calculate the correlation between query feature Q and key feature K using complex inner product, and apply unitary matrix transformation to key feature K and / or query feature Q to enhance the high-dimensional association expression between features and obtain attention association results; Step S314: Normalize the attention association results to generate attention weights, and perform weighted summation on the value features V according to the attention weights to obtain the weighted attention features; Step S315: Map the weighted attention features back to the real-valued feature space to obtain the quantum multi-head attention path output features.

[0069] This embodiment enhances the ability to represent complex nonlinear relationships between distant locations in the input deep features by performing linear projection, complex space mapping, correlation calculation based on complex inner product and unitary matrix transformation, attention weighting, and real-valued space output mapping on the input deep features. It highlights important feature information related to the target region and forms output features containing global correlation information, thereby improving the recognition ability of complex structural regions, distant correlation regions, and fuzzy boundary regions, and addressing the problem of insufficient global correlation modeling ability in existing image processing methods.

[0070] Step S3, the method for global modeling through SSM-enhanced paths, includes the following steps: Step S321: Input the input deep features into the SSM enhancement path, and use the selective state space model to perform state space modeling processing on the input deep features; Step S322: During the state space modeling process, the hidden state is updated based on the current input features to extract long-distance dependency information from the deep input features, thereby obtaining enhanced features containing long-distance dependency information.

[0071] In step S4, the outputs of the quantum multi-head attention path and the SSM enhancement path are fused. The fusion of the output features of the quantum multi-head attention path and the output features of the SSM enhancement path can be achieved by element-wise addition to obtain the enhancement features. In step S5, the enhanced features are input to the decoder for upsampling and recovery, and then combined with the skip connection features of the corresponding level of the encoder for feature fusion to obtain the recovered features and perform class mapping to obtain the image processing result, including the following steps: Step S51: Input the enhanced features into the decoder, and gradually restore the spatial resolution of the feature map through multi-level upsampling, skip connection feature fusion and convolutional refinement to obtain the restored features; Step S511: Input the enhanced features into the upsampling layer of the decoder, and perform upsampling processing on the enhanced features to restore the spatial size of the feature map and obtain the upsampled features; Step S512: The upsampled features are fused with the features output by the downsampled unit of the corresponding level of the encoder by the jump connection fusion unit to obtain the fused features; Step S513: Input the fused features into the feature extraction module. Through convolution, activation and batch normalization processing, the fused features are integrated with local detail information and deep semantic information to obtain the current level decoding features. Step S514: Pass the current level decoding features to the next level upsampling unit for further processing until the recovered features output by the decoder are obtained; Step S52: Input the recovered features into the lightweight MBConv module for feature extraction to obtain the extracted output features; Step S53: Map the extracted output features to the number of feature channels to obtain a category response feature map corresponding to the number of target categories; Step S54: Normalize the category response feature map to generate a pixel-level probability map and obtain the image processing result.

[0072] This embodiment performs multi-level upsampling recovery, skip connection feature fusion, convolutional refinement, lightweight feature extraction, category mapping, and normalization on the enhanced features. This enables the fusion of shallow detail information and deep semantic information while restoring the spatial resolution of the image, enhancing the class discrimination ability of the output features, and generating pixel-level probability distributions. This improves the boundary clarity, structural integrity, and classification accuracy of the image processing results.

[0073] The method in this embodiment organically combines the long-distance dependency modeling capability of the Mamba structure, the multi-scale feature fusion capability of the U-Net architecture, and the efficient global correlation modeling capability of the quantum multi-head attention mechanism. This enables the image segmentation system to simultaneously extract local detail features and represent global contextual information when processing high-resolution medical images, thereby improving the segmentation accuracy of complex structural regions, blurred boundary regions, and subtle target regions. Simultaneously, the lightweight structural design reduces the number of model parameters and computational complexity, which is beneficial for improving training and inference efficiency and enhancing the system's deployment feasibility in resource-constrained environments. Therefore, the method in this embodiment not only meets the comprehensive requirements of accuracy, efficiency, and robustness for medical image segmentation scenarios such as 3T MRI, but can also be extended to fields such as natural image segmentation, remote sensing image analysis, and autonomous driving environmental perception, demonstrating good versatility and application value.

[0074] Example 3 Based on Embodiment 2, this embodiment provides an electronic device, including a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When the computer instructions are executed by the processor, they complete the steps in the MBConv-SSM-based collaborative quantum image processing method described in Embodiment 2.

[0075] Example 4 Based on Embodiment 2, this embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, they complete the steps in the MBConv-SSM-based collaborative quantum image processing method described in Embodiment 1.

[0076] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0077] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A quantum image processing system based on MBConv-SSM, characterized in that, include: The preprocessing module is used to preprocess the input image and then input the preprocessed image into the encoder. The encoder includes a multi-level cascaded downsampling unit. Each downsampling unit includes a sequentially connected MBConv module and a selective state space model layer. The MBConv module is used to extract local features from the input features, and the SSM layer is used to perform global dependency modeling on the features processed by the MBConv module and output deep features. The bottleneck layer includes a parallel quantum multi-head attention path and an SSM enhancement path. The outputs of the quantum multi-head attention path and the SSM enhancement path are fused to obtain the enhanced features. The decoder is used to upsample the enhanced features and fuse the skip connection features of the corresponding level of the encoder for feature recovery. The recovered features are then classified to obtain the image processing result.

2. The MBConv-SSM-based collaborative quantum image processing system as described in claim 1, characterized in that, The encoder includes multiple cascaded downsampling units, each of which includes sequentially connected MBConv modules and a selective state-space model layer; Each MBConv module includes a sequentially connected inverted residual structure, a depthwise separable convolutional layer, a Swish activation function layer, and a batch normalization layer, which are used to implement depthwise separable convolution, function activation, and batch normalization operations on the image to obtain a local feature map. The selective state space model layer receives the feature sequence formed by expanding the local feature map along the spatial dimension. Based on the current input feature sequence, it updates the hidden state to enhance important information in the feature sequence and suppress redundant information, thereby outputting sequence features containing long-distance dependency information.

3. The MBConv-SSM-based collaborative quantum image processing system as described in claim 1, characterized in that, In the MBConv module: The inverse residual structure is used to perform dimensionality upscaling, convolutional feature extraction, and dimensionality reduction compression on the input image features, and retains the original feature information through residual connections to obtain the inverse residual output features. Depth-separable convolutional layers are used to perform spatial convolution on each channel of the inverse residual output feature to extract local spatial features. Then, pointwise convolution is used to achieve feature fusion between channels and adjustment of the number of channels to obtain the features after depth-separable convolution. The Swish activation function layer performs a non-linear mapping on the features processed by depthwise separable convolutions. The batch normalization layer is used to normalize the data after nonlinear mapping to obtain the extracted local features.

4. The MBConv-SSM-based collaborative quantum image processing system as described in claim 1, characterized in that, The bottleneck layer, located between the encoder and decoder, is used to perform global context enhancement on the deep features output by the encoder. It includes parallel quantum multi-head attention paths and SSM enhancement paths, as well as feature fusion units located after the quantum multi-head attention paths and SSM enhancement paths. Quantum multi-head attention path is used to perform global correlation modeling on deep features of the input; deep features are the output features of the last stage of the encoder; SSM enhancement path, used for state-space modeling of deep features of the same input; The feature fusion unit is used to fuse the features output by the quantum multi-head attention path and the SSM enhancement path to obtain the enhanced bottleneck layer features, which are then output to the decoder.

5. The MBConv-SSM-based collaborative quantum image processing system as described in claim 1, characterized in that, The quantum multi-head attention path includes a feature projection unit, a complex mapping unit, a quantum correlation computation unit, a weighting unit, and an output projection unit connected in sequence. The feature projection unit is used to perform a linear transformation on the input deep features through a learnable linear layer to generate query features Q, key features K, and value features V. The complex mapping unit is used to split the query feature Q and the key feature K into real and imaginary parts, or to perform a preset transformation on the query feature Q and the key feature K to map the query feature Q and the key feature K to the complex vector space to obtain the complex feature representation; The quantum correlation computing unit is used to calculate the correlation between query feature Q and key feature K based on complex inner product, and to apply unitary matrix transformation to key feature K and / or query feature Q to enhance the high-dimensional correlation expression between features and obtain attention correlation results. The weighting unit is used to perform normalization processing on the feature representation in the form of constructed complex numbers output by the quantum correlation computing unit, generate attention weights, and perform weighted summation on the value feature V according to the attention weights to obtain the weighted attention feature; The output projection unit is used to map the weighted attention features from complex or high-dimensional representation back to the real-valued feature space through a linear projection layer, thereby obtaining the quantum multi-head attention path output features.

6. The MBConv-SSM-based collaborative quantum image processing system as described in claim 1, characterized in that, The decoder, located after the bottleneck layer, includes a multi-level cascaded upsampling unit, each of which includes an upsampling layer, a skip connection fusion unit, and a feature extraction module. The upsampling layer is used to restore the spatial size of the input feature map to obtain upsampled features. The skip connection fusion unit is used to fuse upsampled features with features output by the downsampled unit of the corresponding level of the encoder. The feature extraction module is used to perform convolution processing on the features fused by the skip connection fusion unit to obtain the current level decoding features, and output them to the next level upsampling unit or output layer.

7. The MBConv-SSM-based collaborative quantum image processing system as described in claim 1, characterized in that, The decoder is connected to an output layer, which includes a lightweight MBConv module, a channel mapping layer, and a probability output layer. The lightweight MBConv module is used to extract features from the recovered features output by the decoder; The channel mapping layer is used to map the number of feature channels extracted by the lightweight MBConv module to the number of target categories; The probability output layer is used to normalize the number of target categories after mapping in order to generate a pixel-level probability map.

8. A method for quantum image processing based on MBConv-SSM, characterized in that, Includes the following steps: Preprocess the acquired input image; The preprocessed image is input into the encoder, and features are extracted step by step through a multi-level downsampling unit. Each downsampling unit performs the following steps in sequence: local feature extraction is performed on the input features using the MBConv module; global dependency modeling is performed on the features after local feature extraction using the selective state space model (SSM) layer. The deep features output by the encoder are input into the bottleneck layer, and global modeling is performed in parallel through the quantum multi-head attention path and the SSM enhancement path. The outputs of the quantum multi-head attention path and the SSM enhancement path are fused to obtain the enhanced features; The enhanced features are input to the decoder for upsampling and recovery, and then combined with the skip connection features of the corresponding level of the encoder for feature fusion to obtain the recovered features. These features are then mapped to categories to obtain the image processing result.

9. The MBConv-SSM-based collaborative quantum image processing method as described in claim 8, characterized in that, Global correlation modeling of deep input features using quantum multi-head attention path includes the following steps: The deep features of the input are linearly transformed by a learnable linear layer to generate query features Q, key features K, and value features V. The query feature Q and the key feature K are split into real and imaginary parts, or a preset transformation is performed on the query feature Q and the key feature K to map the query feature Q and the key feature K to the complex vector space to obtain the complex feature representation; Based on complex feature representation, the correlation between query feature Q and key feature K is calculated using complex inner product, and unitary matrix transformation is applied to key feature K and / or query feature Q to enhance the high-dimensional association expression between features, thus obtaining attention association results; The attention association results are normalized to generate attention weights, and the value features V are weighted and summed according to the attention weights to obtain the weighted attention features. The weighted attention features are mapped back to the real-valued feature space to obtain the quantum multi-head attention path output features.

10. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the steps in the MBConv-SSM-based collaborative quantum image processing method as described in any one of claims 8-9.