Image segmentation method and system based on double-branch network architecture, medium and equipment

Through the image segmentation method of the dual-branch network architecture, the diffusion model and mask autocode model are used to extract features and perform cross-modal attention fusion, which solves the segmentation problem of traditional medical image segmentation models under different scales and complex backgrounds, and achieves higher accuracy lesion boundary extraction and robustness.

CN120388037AActive Publication Date: 2025-07-29SHANDONG UNIV

Patent Information

Application Number
CN202510884178.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-29
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Traditional medical image segmentation models are difficult to take into account semantic information at different scales and levels at the same time, and noise and fuzzy problems increase the difficulty of accurate segmentation.

Method used

The image segmentation method based on the dual-branch network architecture is adopted, and the semantic and local features are extracted respectively using the diffusion model and the masked self-coding model, and feature alignment and fusion are performed through the cross-modal attention mechanism, combining multi-scale feature fusion and adaptive denoising step adjustment.

Benefits of technology

It improves the segmentation accuracy of medical image segmentation and the detailed extraction ability of lesion boundaries, adapts to the challenges of complex backgrounds and multi-scale, and enhances the robustness and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388037A_ABST
    Figure CN120388037A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and provides an image segmentation method and system based on a double-branch network architecture, a medium and equipment, and the method comprises the steps: for a to-be-segmented image, measuring the complexity of the to-be-segmented image through edge detection and texture analysis, distributing the denoising step length of a diffusion model, and extracting semantic features; after random shielding and fuzzy processing are carried out on a to-be-segmented image, local features and global features are extracted through a mask self-encoding model, and fusion features are obtained through multi-scale feature fusion; and the semantic features and the fusion features are paired, the similarity between the two paired features is calculated, an attention weight is generated, and after the paired features are fused based on the attention weight, the to-be-segmented image is segmented. The segmentation precision, the detail extraction of the focus boundary and the capturing capability of the multi-scale focus are remarkably improved, and the method adapts to challenges of different scales and complex backgrounds in medical images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and in particular relates to an image segmentation method, system, medium and device based on a dual-branch network architecture. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] In recent years, with the progress of deep learning in the field of image segmentation, various convolutional neural networks (CNNs) and models based on self-attention mechanisms have gradually been applied to medical image segmentation tasks.

[0004] However, medical images usually have complex background interference, subtle lesion boundaries and multi-scale structural features. Traditional medical image segmentation usually relies on feature extraction and segmentation prediction of a single model. A single model cannot simultaneously take into account semantic information of different scales and levels, resulting in limited segmentation effect; in addition, noise and blur problems in medical images further increase the difficulty of accurate segmentation. Summary of the Invention

[0005] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides an image segmentation method, system, medium and equipment based on a dual-branch network architecture, which uses a dual-branch architecture for image segmentation. One branch uses a diffusion model to process the image, and the other branch uses a masked autoencoder model to process the image. Through the cross-modal attention mechanism, the features extracted by the two models are aligned, the semantic representation is enhanced, and there are significant improvements in segmentation accuracy, detail extraction of lesion boundaries, and the ability to capture multi-scale lesions, adapting to the challenges of different scales and complex backgrounds in medical images.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions: A first aspect of the present invention provides an image segmentation method based on a dual-branch network architecture, comprising: Obtain the image to be segmented; For the image to be segmented, the complexity of the image to be segmented is measured through edge detection and texture analysis. After allocating the denoising step size of the diffusion model according to the complexity, the semantic features are extracted through the diffusion model. After random occlusion and blurring of the segmented image, local and global features are extracted through the mask autoencoder model, and fused features are obtained through multi-scale feature fusion. The semantic features and fusion features are paired, the similarity between the two paired features is calculated, and the attention weight is generated. After the paired features are fused based on the attention weight, the segmentation representation is obtained through the activation function, and the image to be segmented is segmented based on the segmentation representation.

[0007] Furthermore, the denoising step size is inversely proportional to the complexity of the image to be segmented.

[0008] Furthermore, the diffusion model and the mask autoencoder model are trained using pixel-level loss, feature alignment loss, and hierarchical reconstruction loss.

[0009] Furthermore, the feature alignment loss is: ,in, L is the total number of layers, The diffusion model and the masked autoencoder model are respectively l Layer characteristics, is a hyperparameter.

[0010] Furthermore, the masked autoencoder model simulates the missing image blocks by performing random occlusion and blurring on the image to be segmented, and then reconstructs the missing parts through the encoder. The encoder extracts local features through the multi-scale convolutional layer and captures global features through the Transformer layer; the local features and global features are fused at a multi-scale level to obtain a fused feature map.

[0011] Furthermore, the complexity is: , texture complexity , average edge strength complexity ,in, is the edge strength, and are the gradients in the horizontal and vertical directions respectively, N represents the total number of edge pixels involved in the calculation, α and β is the weight coefficient.

[0012] A second aspect of the present invention provides an image segmentation system based on a dual-branch network architecture, comprising: An image acquisition module is configured to: acquire an image to be segmented; A first feature extraction module is configured to: measure the complexity of the image to be segmented through edge detection and texture analysis, assign a denoising step size of a diffusion model according to the complexity, and then extract semantic features through the diffusion model; The second feature extraction module is configured to: after randomly occluding and blurring the image to be segmented, extract local features and global features through a mask autoencoder model, and obtain fused features through multi-scale feature fusion; The image segmentation module is configured to: pair semantic features and fusion features, calculate the similarity between the two paired features, generate attention weights, fuse the paired features based on the attention weights, obtain segmentation representations through activation functions, and segment the image to be segmented based on the segmentation representations.

[0013] Further, the denoising step size is inversely proportional to the complexity of the image to be segmented.

[0014] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the image segmentation method based on a dual-branch network architecture as described above.

[0015] The fourth aspect of the present invention provides a computer device, including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor. When the processor executes the program, it implements the steps in the image segmentation method based on a dual-branch network architecture as described above.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention uses a dual-branch structure for image segmentation. One branch uses a diffusion model to process the image, and the other branch uses a masked autoencoder model to process the image. Through the cross-modal attention mechanism, the features extracted by the two models are aligned, enhancing the semantic representation. There is a significant improvement in segmentation accuracy, the extraction of details of lesion boundaries, and the ability to capture multi-scale lesions, adapting to the challenges of different scales and complex backgrounds in medical images.

[0017] The present invention aligns the features of the diffusion model and the MIM model through the cross-modal attention mechanism, fusing semantic features and fused features at each layer, and being more coordinated in capturing details and the overall structure. This mechanism enables the image segmentation model to simultaneously focus on fine boundary features and overall semantic consistency at the pixel level, thereby improving the segmentation effect.

[0018] By introducing an adaptive reverse diffusion mechanism, the diffusion model can dynamically adjust the denoising step size according to the image feature distribution and the current denoising state to more finely retain and extract semantic features. The denoising step size of the diffusion model is allocated according to the complexity. If the edges and textures are complex and the detail information is rich, the step size is reduced to more finely denoise and retain the detail features, ensuring that the details in the complex area can be gradually restored and avoiding information loss. If the overall complexity is low, the step size is increased to reduce the fineness of denoising to denoise more quickly in the simple area.

[0019] The training of the diffusion model and the masked autoencoder model of the present invention adopts pixel-level loss, feature alignment loss, and hierarchical reconstruction loss. Among them, the feature alignment loss can make the features of the diffusion model and the masked autoencoder model more similar at different layers, ensuring the consistency of the features extracted by the two models in terms of structure and semantics, thereby improving the overall performance of the model. The use of other loss functions such as hierarchical reconstruction loss also helps the model to accurately learn and optimize at different levels, further improving the segmentation ability of the model.

[0020] The masked autoencoder model of the present invention simulates the situation of missing image patches by performing random occlusion and blurring on the image to be segmented, and then learns to reconstruct the missing parts. This method enhances the model's ability to process incomplete information in the image, enabling it to maintain good segmentation performance even when facing the situation of missing or unclear image information that may occur in the actual scene, and improving the robustness and generalization ability of the model.

[0021] In the masked autoencoder model of the present invention, the encoder extracts local features through multi-scale convolutional layers and captures global features through Transformer layers. The multi-scale feature fusion can fully combine local detail information and global structure information, enabling the model to understand the image more comprehensively and helping to more accurately identify and segment different objects and regions in the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0023] Figure 1 is a flowchart of the image segmentation method based on a dual-branch network architecture according to Embodiment 1 of the present invention; Figure 2 is a structural diagram of the diffusion model according to Embodiment 1 of the present invention; Figure 3 is a graph of experimental results according to Embodiment 1 of the present invention; Figure 4 is a schematic structural diagram of a computer device according to Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0025] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0026] Embodiment 1 This embodiment provides an image segmentation method based on a dual-branch network architecture.

[0027] The image segmentation method based on a dual-branch network architecture provided in this embodiment proposes a dual-branch network architecture that fuses multiple features to enhance the feature expression ability of the image segmentation model and improve the segmentation accuracy. One branch uses a diffusion model to extract semantic features at different levels, and the other branch uses a pre-trained network of MIM to extract features. Through the cross-modal attention mechanism, the image segmentation model can align the two types of features, enhancing the semantic representation, and thus making predictions through a pixel classifier.

[0028] Diffusion models perform excellently in generative tasks, capable of gradually removing noise and extracting multi-level semantic features, and are suitable for recovering the detailed information of images from noisy data; while the Masked Image Modeling (MIM) network captures the global structural features of images by reconstructing occluded images. The two models are complementary in feature extraction at different levels, and combining them can obtain an omni-directional semantic representation of the target area.

[0029] The image segmentation method based on a dual-branch network architecture provided in this embodiment is applicable to medical image segmentation.

[0030] The image segmentation method based on a dual-branch network architecture provided in this embodiment, as Figure 1 shown, includes the following steps: Step 1: Obtain the image to be segmented. For the image to be segmented, measure the complexity of the image to be segmented through edge detection and texture analysis. After allocating the denoising step size of the diffusion model according to the current complexity, extract semantic features at different scales through the diffusion model.

[0031] The principle of the diffusion model is to convert a clean image into a noisy image by gradually adding noise, and then recover the original image by denoising in the reverse process. The diffusion model generates image features at multiple time steps T and uses a noise predictor (i.e., the U-Net network model) to generate activation states at different stages. The U-Net (U-shaped network) model is the core module in the diffusion model, which mainly iteratively denoises the Gaussian noise matrix in the diffusion cycle. Using the reverse diffusion process of the diffusion model, the intermediate activation states of the labeled image are used to extract semantic features. In the reverse diffusion process, different levels of semantic features can be extracted at different levels of the U-Net architecture for subsequent analysis and processing.

[0032] Specifically, this embodiment proposes an adaptive reverse diffusion mechanism. In the reverse denoising process, a dynamic step size adjustment based on image complexity is introduced. The specific implementation method is to calculate the complexity of the current state of the image before each reverse diffusion step and allocate a more detailed step size according to the complexity.

[0033] Step 101. Initial settings for reverse diffusion: First, a diffusion model is pre-trained based on a large-scale image dataset (ImageNet); subsequently, the pre-trained model is fine-tuned using a domain-specific ultrasound dataset to adapt to the feature distribution of medical images; after fine-tuning, the diffusion model starts to execute the reverse diffusion process: gradually denoising the noisy image using a noise predictor with a U-Net architecture. This process restores the image from a random noise state to a clear image by gradually removing the noise.

[0034] Before entering each reverse diffusion step, the diffusion model dynamically calculates the complexity of the current image to determine the step size of the current denoising.

[0035] Step 102. Calculate the complexity of the current image: Before each reverse diffusion step, the current image state is extracted, and the complexity of the image is measured through edge detection and texture analysis.

[0036] Based on the complexity determination of edge detection, edges can represent the contours of lesion areas. The edge density and edge strength of the image can be used to judge the image complexity. The Sobel operator or Canny edge detection operator is used to obtain the edge gradient of the image.

[0037] Specifically, for a given image , the gradients and [[ID=S17]] are obtained using the Sobel operator in the horizontal and vertical directions respectively, where and are the horizontal and vertical Sobel kernels respectively.

[0038] The edge strength can be expressed as: .

[0039] The formula for the average edge strength complexity is as follows: , where N represents the total number of edge pixels participating in the calculation, which can be expressed as: , where W and H are the width and height of the image respectively; 1(·) is an indicator function that takes the value of 1 when the condition is true and 0 otherwise; indicates that the current pixel is considered an edge.

[0040] Based on the complexity determination of texture features, on the basis of the Sobel operator, the texture complexity is represented by calculating the average value of the local gradient intensity: .

[0041] The complexities of the two can be combined into a comprehensive complexity metric through a linear combination: , where α and β are weight coefficients.

[0042] Step 103: Adjust the denoising step size according to the overall complexity: During the back diffusion process, the diffusion model allocates a dynamic step size according to the current complexity: if the overall complexity is high (for example, complex edges and textures, rich detail information), the step size is reduced to achieve more detailed denoising and retain detail features; if the overall complexity is low, the step size is increased and the denoising fineness is reduced to improve efficiency.

[0043] Adaptive denoising with dynamic step size enhances the retention of local image details, making the extracted multi-scale semantic features richer and more accurate.

[0044] The traditional diffusion model's reverse denoising process gradually restores the image using a fixed noise step size. This strategy can be inadequate for complex medical images or those with subtle features. By introducing an adaptive reverse diffusion mechanism, the diffusion model can dynamically adjust the denoising step size and intensity based on the image feature distribution and the current denoising state to more precisely preserve and extract semantic features. For example, if the diffusion model detects that the current image feature is very important, the step size can be reduced for more detailed denoising. Conversely, if the noise has little impact on the image semantics, the step size can be increased, reducing the number of denoising steps to improve efficiency.

[0045] Step 104: Perform dynamic step-size backdiffusion denoising: Using the adjusted step size, denoising is performed in each backdiffusion step. A small step size ensures that details in complex areas are gradually restored, avoiding information loss; a large step size allows for faster denoising in simple areas.

[0046] In each denoising step, different layers of U-Net generate multi-scale activation states, through which semantic features at different levels are extracted.

[0047] Step 105: Multi-scale feature fusion and semantic feature extraction.

[0048] In the middle layer of each back-diffusion step, semantic features are extracted from the activation states of different scales of U-Net, and a multi-scale feature pyramid structure is used to ensure the comprehensiveness of feature information.

[0049] like Figure 2As shown in the figure, in order to extract higher-quality semantic features with multi-scale fine-grainedness, this embodiment introduces a multi-scale dynamic weight feature pyramid mechanism to enhance the expression ability of features at different scales in medical images. A multi-scale U-Net architecture is adopted, and feature extraction and noise prediction are carried out simultaneously at different scales during the diffusion process to capture rich information from the global structure to local details in the image. Specifically, in the U-Net encoder stage, a multi-scale feature pyramid is constructed. The features at each scale are dynamically weighted by a scale evaluation module, and the weight ratio is adjusted by soft normalization (Softmax), and then weighted fusion is further performed. To ensure that the features at each scale can effectively represent information of different granularities, different-sized convolutional kernels are used in the generation process of the feature pyramid, so as to promote the diffusion model to take into account both the macroscopic structure and the fine-grained features of tiny lesions when restoring features. In addition, to further improve feature selectivity, in the U-Net encoder stage, a hybrid weighted attention mechanism is introduced, that is, a dual-branch structure of channel attention and spatial attention, which can dynamically focus on more diagnostically significant key regions and channels during the multi-scale feature fusion process, and improve the model's response ability to key information.

[0050] Step 2: For the image to be segmented, through the MIM model and multi-scale feature fusion, a fused feature map is obtained.

[0051] The MIM model simulates the situation of missing some image patches by randomly occluding and blurring the image. The MIM model learns the ability to reconstruct these missing parts through the encoder, so as to extract the detailed representation of the image. This process uses convolutional operations for local feature extraction and captures global information through Vision Transformer (Visual Transformation Network). Transformer (Transformation Network) is a deep learning model architecture that introduces a self-attention mechanism. Vision Transformer applies Transformer to visual tasks and has strong capabilities for global information modeling.

[0052] The MIM model combines multi-scale convolution and Vision Transformer structure. Local convolution operations help capture details and local information, while global Transformer operations effectively obtain long-range dependencies and global information. This way ensures that the model can take into account the coherence of global semantics while processing details.

[0053] Step 201: Input module: Random occlusion and blurring processing.

[0054] The input image is preprocessed by randomly blocking certain image blocks, causing information loss in these areas. The blocked areas are then blurred to simulate the blurring or loss of image details. This allows the MIM model to learn how to complete and restore the lost areas.

[0055] Step 202: MIM encoder structure.

[0056] Multi-scale convolutional layers: Multi-scale convolution kernels are first applied in the encoder to capture details of the image at different scales. These convolution operations capture local features at different scales, allowing the model to have a deeper understanding of the details of the occluded area and its neighboring parts.

[0057] Use convolution kernels of different sizes (such as 3×3, 5×5, etc.) to extract multi-scale features to ensure that the detailed changes in the image are captured.

[0058] Transformer layer: After the multi-scale convolutional layer, the Transformer module is introduced to extract global semantic information. The Transformer structure uses a self-attention mechanism to model the long-range dependencies between regions in the image, thereby acquiring global information.

[0059] In the attention mechanism, the features of multi-scale convolution output are considered so that the information of the occluded area can be better integrated with the surrounding areas.

[0060] Through the self-attention mechanism, the MIM model can complete and restore the semantic content of the occluded area at the global level.

[0061] Step 203: Multi-scale feature fusion module: First, the multi-scale feature fusion module receives the local features extracted by the multi-scale convolution layer (such as 3×3, 5×5, and 7×7 convolution outputs) and the global features of the Transformer layer. The dimensions of each feature may be different, so alignment and dimensionality reduction processing are required.

[0062] Among them, feature alignment and dimensionality reduction processing include: (1) Channel dimension reduction: Since the features from convolutions and Transformers of different scales may have different channel numbers, the channel numbers of these features are first aligned through 1×1 convolution to ensure that all features have consistent channel dimensions.

[0063] (2) Spatial alignment: If the spatial resolutions of multi-scale feature maps are different, bilinear interpolation or transposed convolution can be used to upsample or downsample the feature maps to align all feature maps to the same spatial size for subsequent fusion.

[0064] Next, the multi-scale feature weight generator includes: (1) Feature importance weight: To dynamically adjust the importance of features at different scales, a weight generator is introduced. Usually, a lightweight fully connected layer or an attention mechanism is used to assign weights to features at each scale. By learning these weights, the MIM model can adaptively adjust the contributions of features at each scale according to the content of the input image.

[0065] (2) Calculation method: The features at each scale are input into the weight generator to generate corresponding weight values, which are then normalized by softmax (normalized exponential function) so that the sum of the weights of all scale features is 1.

[0066] Next, feature weighted fusion includes: (1) Scale-by-scale weighting: Multiply the features at different scales by their corresponding weights to ensure that the contribution ratio of features at each scale conforms to the weight distribution during fusion.

[0067] (2) Feature summation: Element-wise sum all the weighted features at all scales to obtain the fused feature map. At this time, the detailed information and global semantic information are integrated to form a unified multi-scale feature representation.

[0068] Step 204, Output layer (enhancement of the fused features): The fused feature map is further processed by a 3×3 convolutional kernel to enhance local feature details and reduce information loss during the fusion process.

[0069] Step 205, Perform a non-linear transformation through an activation function (such as ReLU or Leaky ReLU) to ensure that the representational ability of the feature map is further enhanced and the segmentation accuracy of the MIM model is improved. The full name of ReLU (Rectified Linear Unit) is Rectified Linear Unit, which is a commonly used activation function in neural networks. Generally speaking, it refers to the ramp function in mathematics, and the formula is as follows: . The leaky rectified linear unit (Leaky Rectified Linear Unit, LeakyReLU) is a variant of the ReLU activation function.

[0070] Step 3, Pair the semantic features and the fused feature map, calculate the similarity between the two paired features to generate attention weights, fuse the paired features based on the attention weights, obtain the segmentation representation through the activation function, and segment the image to be segmented based on the segmentation representation.

[0071] A cross-attention module is proposed to align features extracted from the diffusion model and the MIM model. Features extracted from different diffusion decoder layers are aligned with corresponding MIM features. The diffusion model and the MIM model each process different aspects of the image. The diffusion model focuses on capturing semantic features through progressive denoising, while the MIM model learns global and local information of the image by reconstructing occluded areas. The cross-attention module aligns the features extracted from different diffusion decoder layers with the corresponding MIM model features.

[0072] The local features extracted by each diffusion decoder layer are aligned with the corresponding hierarchical features of the MIM model to ensure that the fusion of local and global information can be fully utilized. This alignment operation not only enhances the mutual understanding of features from different models, but also improves segmentation accuracy.

[0073] To address the problem of insufficient context caused by attention calculation limited to the current level, the cross-modal attention module introduces a cross-level attention transfer mechanism. Through this mechanism, low-level detail features can effectively flow to high-level feature layers, helping to generate a more complete semantic representation.

[0074] The cross-level attention transfer mechanism specifically includes: Information transfer between layers: Setting up cross-layer connections to combine low-level features with high-level features so that low-level details can provide support for high-level features; By introducing hierarchical information into cross-modal data through layer-by-layer fusion, the alignment accuracy is improved.

[0075] In this embodiment, the specific processing steps of the cross-modal attention module are as follows: Step 301: Obtain features of the diffusion model decoder layer and the MIM model encoder layer, and pair them at the same level to form feature pairs.

[0076] Step 302, cross-model attention calculation: For each pair of features, first generate attention weights by calculating the similarity between the features. These weights are used to measure the similarity between the current diffusion feature (semantic feature) and the MIM feature (fusion feature). The specific implementation method is: Calculate the feature similarity matrix through dot product: ;in: and Represents semantic features and fusion features The feature vectors of the i-th and j-th positions in ; represents the dot product, denotes the norm of the feature vector; S(i, j) represents the similarity between position i and position j, ranging from [-1, 1]; Based on the similarity matrix S, calculate the attention weights through the softmax function: ; where: S(i, j) is the and similarity, and k represents the feature index; measures the impact on , and through softmax normalization, ensure that the sum of the weights at each position is 1; , measures the impact on , and through softmax normalization, ensure that the sum of the weights at each position is 1.

[0077] Step 303, weighted fusion feature output.

[0078] Weight and sum each pair of features according to the attention weights to obtain the finally fused features. Process the fused features through an activation function to form the final segmentation representation, ensuring that global and local information is integrated and optimized.

[0079] Weightedly fuse the two types of features according to the attention weights to obtain the fused feature representation . The bidirectional attention mechanism ensures that the fusion process considers the interaction between the two types of features and avoids bias towards a single source.

[0080] ; where α and β are hyperparameters used to adjust the weights of the diffusion features and MIM features.

[0081] Step 304, output result. The fused features output by the cross-modal attention module enter the downstream pixel classifier to further improve the segmentation accuracy.

[0082] In this embodiment, the diffusion model, the masked autoencoder model, the cross-modal attention module, and the pixel classifier constitute an image segmentation model.

[0083] In this embodiment, two loss functions are used to train the image segmentation model: Pixel-level loss: Measure the quality of the segmentation result through the Dice coefficient, and optimize the features extracted by the diffusion model and the MIM model; The Dice coefficient is a set similarity metric function, usually used to calculate the similarity between two samples, and is a common evaluation metric for semantic segmentation, with a value range of [0, 1]; Cross-modal alignment loss: used to ensure that the features extracted by the diffusion model and the MIM model are well aligned and fused in the cross-modal attention mechanism.

[0084] Specifically, the feature fusion between the diffusion model and the MIM model is optimized through the cross-modal alignment loss. This loss function measures the similarity between the features of the two modalities, ensuring that they are fully aligned in the high-dimensional space. The cross-modal alignment loss can be divided into two main components: feature alignment loss and hierarchical reconstruction loss, each of which contributes to effective alignment between cross-modal features.

[0085] Feature alignment loss: measures the similarity between features from different modalities: , where L is the total number of layers, The diffusion model and MIM model are l Layer characteristics, is a hyperparameter, cos Represents cosine similarity.

[0086] Hierarchical reconstruction loss can be used to ensure that the cross-modal attention module maintains the integrity of information when fusing features. The mean squared error (MSE) can be used to measure the difference between the reconstructed features and the true labels: ,in, It is l The output of layer reconstruction, is the true label, is a hyperparameter used to control the importance of reconstruction loss at different levels, and N is the number of samples.

[0087] Image segmentation model training is performed through a joint optimization approach, gradually adjusting the feature extraction processes of the diffusion model and the MIM model, and ensuring feature alignment between the two through a cross-modal attention mechanism. The entire training process includes forward propagation, loss calculation, backpropagation, and gradient updates, and segmentation performance is gradually optimized through multiple rounds of iteration.

[0088] The image segmentation method based on the dual-branch network architecture provided in this embodiment has the following experimental results: Figure 3 shown.

[0089] The image segmentation method based on a dual-branch network architecture provided in this embodiment uses a dual-branch architecture for image segmentation. One branch uses a diffusion model to process the image, while the other branch uses a pre-trained MIM network to capture details in the medical image. The cross-modal attention mechanism aligns the two features, enhancing the semantic representation, thereby enabling prediction through a pixel classifier. This has the following technical effects: Rich multi-level semantic feature extraction: In the process of gradually restoring the image using the diffusion model branch, the diffusion model can generate multi-level semantic features, including both the detailed features of the lesion area and the boundary features between the background and the lesion, providing strong support for fine segmentation; Enhanced global structure representation: Through the MIM network branch, the MIM model can learn the large-scale structure information in the image from a global perspective, and has stronger semantic understanding ability in the discrimination between the lesion area and the background area, which helps to improve the robustness of the image segmentation model in complex backgrounds; Cross-modal alignment enhanced feature fusion: Align the features of the diffusion model and the MIM model through the cross-modal attention mechanism, fuse local and global features at each layer, and are more coordinated in capturing details and overall structures. This mechanism enables the image segmentation model to simultaneously focus on fine boundary features and overall semantic consistency at the pixel level, thereby improving the segmentation effect; Accurate segmentation prediction: After combining the features of the two branches, use a pixel classifier to perform pixel-level classification on the enhanced semantic features. Through this fusion strategy, the image segmentation model has significant improvements in segmentation accuracy, detailed extraction of lesion boundaries, and the ability to capture multi-scale lesions, adapting to the challenges of different scales and complex backgrounds in medical images.

[0090] Embodiment 2 This embodiment provides an image segmentation system based on a dual-branch network architecture, including: An image acquisition module, which is configured to: acquire the image to be segmented; A first feature extraction module, which is configured to: for the image to be segmented, measure the complexity of the image to be segmented through edge detection and texture analysis, allocate the denoising step size of the diffusion model according to the complexity, and then extract semantic features through the diffusion model; A second feature extraction module, which is configured to: after performing random occlusion and blurring processing on the image to be segmented, extract local and global features through the masked autoencoder model, and obtain fused features through multi-scale feature fusion; An image segmentation module, which is configured to: pair the semantic features and the fused features, calculate the similarity between the two paired features, generate attention weights, fuse the paired features based on the attention weights, obtain a segmentation representation through an activation function, and segment the image to be segmented based on the segmentation representation.

[0091] Furthermore, the denoising step size is inversely proportional to the complexity of the image to be segmented.

[0092] It should be noted here that each module in this embodiment corresponds to each step in Embodiment 1, and its specific implementation process is the same, so it will not be repeated here.

[0093] Embodiment III This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the image segmentation method based on the dual-branch network architecture as described in Embodiment I above.

[0094] Embodiment IV This embodiment provides a computer device, as Figure 4 shown, including a computer-readable storage medium 1003, a processor 1001, a communication interface 1002, and a computer program stored on the computer-readable storage medium 1003 and executable on the processor 1001. Among them, the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected through a bus or other means. Among them, the communication interface 1002 is used to receive and send data, and when the processor 1001 executes the program, it implements the steps in the image segmentation method based on the dual-branch network architecture as described in Embodiment I above.

[0095] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An image segmentation method based on a dual-branch network architecture, characterized in that, including: obtain the image to be segmented; For the image to be segmented, measure the complexity of the image to be segmented through edge detection and texture analysis, and after allocating the denoising step size of the diffusion model according to the complexity, extract semantic features through the diffusion model; After performing random occlusion and blurring processing on the image to be segmented, use a masked autoencoder model to extract local features and global features, and obtain fused features through multi-scale feature fusion; Pair the semantic features and the fused features, calculate the similarity between the two paired features, generate attention weights, fuse the paired features based on the attention weights, obtain a segmentation representation through an activation function, and segment the image to be segmented based on the segmentation representation.

2. The image segmentation method based on a dual-branch network architecture according to claim 1, wherein The denoising step size is inversely proportional to the complexity of the image to be segmented.

3. The image segmentation method based on a dual-branch network architecture according to claim 1, wherein The training of the diffusion model and the masked autoencoder model adopts pixel-level loss, feature alignment loss, and hierarchical reconstruction loss.

4. The image segmentation method based on a dual-branch network architecture according to claim 3, wherein, The feature alignment loss is as follows: , where L is the total number of layers, are the features of the diffusion model and the masked autoencoder model at the l -th layer respectively, is a hyperparameter.

5. The image segmentation method based on a dual-branch network architecture according to claim 1, wherein The masked autoencoder model simulates the situation of missing image patches by performing random occlusion and blurring processing on the image to be segmented, and then the encoder learns to reconstruct the missing part. The encoder extracts local features through multi-scale convolutional layers and captures global features through Transformer layers; the local features and global features are fused through multi-scale feature fusion to obtain a fused feature map.

6. The image segmentation method based on a dual-branch network architecture according to claim 1, characterized in that The complexity is as follows: , the texture complexity , the average edge strength complexity , where is the edge strength, and are the gradients in the horizontal and vertical directions respectively, N represents the total number of edge pixels participating in the calculation, α and β are the weight coefficients.

7. An image segmentation system based on a dual-branch network architecture, characterized in that, including: an image acquisition module configured to: obtain the image to be segmented; a first feature extraction module configured to: for the image to be segmented, measure the complexity of the image to be segmented through edge detection and texture analysis, and after allocating the denoising step size of the diffusion model according to the complexity, extract semantic features through the diffusion model; a second feature extraction module configured to: after performing random occlusion and blurring processing on the image to be segmented, use a masked autoencoder model to extract local features and global features, and obtain fused features through multi-scale feature fusion; an image segmentation module configured to: pair the semantic features and the fused features, calculate the similarity between the two paired features, generate attention weights, fuse the paired features based on the attention weights, obtain a segmentation representation through an activation function, and segment the image to be segmented based on the segmentation representation.

8. The image segmentation system based on the dual-branch network architecture according to claim 7, wherein, The denoising step size is inversely proportional to the complexity of the image to be segmented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the image segmentation method based on a dual-branch network architecture according to any one of claims 1-6.

10. A computer device, comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the image segmentation method based on a dual-branch network architecture according to any one of claims 1-6.

Citation Information

Patent Citations

  • Medical image segmentation method based on cross attention and cross-scale fusion

    CN116385724A

  • Lane line detection method based on mask image modeling

    CN117523518A

  • Medical image segmentation method based on feature interaction

    CN118134952A

  • Cervical vertebra segmentation and key point detection method based on diffusion model

    CN119180954A

Cited By

  • PVNet and double-branch feature fusion object pose detection method and device

    CN120747226A

  • Multi-scale background segmentation and retrieval method and device, computer equipment and storage medium

    CN121210701A