Medical image segmentation method and system based on multi-scale feature fusion and correction

By adopting multi-scale feature fusion and correction methods in medical image segmentation, the problems of insufficient feature fusion, insufficient boundary supervision and poor noise processing in the prior art are solved, and more efficient segmentation effect and robustness are achieved.

CN119941747APending Publication Date: 2025-05-06SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510086421.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing medical image segmentation methods lack bidirectional interaction when fusing local and global features, lack of boundary supervision, difficult to effectively deal with noise and non-robust features, and a single-scale model cannot adapt to lesion areas of different sizes.

Method used

Using a method based on multi-scale feature fusion and correction, features of different scales are extracted through the PVT network, the two-way interactive fusion module realizes two-way interaction of local and global features, the boundary supervision module generates boundary embedding, the feature separation correction module corrects non-robust features, and generates feature maps that enhance boundary attention through the cross-scale boundary guidance module.

Benefits of technology

It improves the segmentation effect of medical image segmentation, enhances the robustness of the model, can better deal with boundary blur and noise interference, and adapts to lesion areas of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941747A_ABST
    Figure CN119941747A_ABST
Patent Text Reader

Abstract

The invention provides a medical image segmentation method and system based on multi-scale feature fusion and correction, and belongs to the technical field of medical image segmentation. The method comprises the following steps: acquiring and preprocessing a medical image to be segmented; in the PVT network, obtaining a first medical image feature map and inputting the first medical image feature map into the bidirectional interaction fusion module and the boundary supervision module, and obtaining a second medical image feature map and boundary embedding; inputting a high-level feature map in the second medical image feature map into a feature separation and correction module to obtain a third medical image feature map; and embedding and inputting the third medical image feature map, the uncorrected second medical image feature map and the boundary into a cross-scale boundary guide module to generate a fourth medical image feature map, and obtaining a final medical image segmentation result by using a predictor. The complex relation between the local information and the global information of the medical image is disclosed, boundary embedding is utilized, the problem that the boundary of the lesion area is fuzzy is better solved, and robustness is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and in particular relates to a medical image segmentation method and system based on multi-scale feature fusion and correction. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Medical image segmentation is a complex and critical step in the field of medical image processing and analysis. Its purpose is to segment parts of medical images with certain special meanings and extract relevant features to provide a reliable basis for clinical diagnosis and pathological research, and to assist doctors in making more accurate diagnoses. At present, there are two main types of segmentation methods: convolutional neural network-based methods and Transformer-based methods.

[0004] In current research, Vision Transformers (ViTs) have excellent long-range modeling capabilities and context-aware features, but lack local inductive bias and have certain limitations in capturing local information; Convolutional neural networks have stronger local modeling capabilities and can capture local spatial information well, providing image-related prior knowledge, but because the convolution operation obtains local receptive fields, it limits the model's ability to obtain global information and cannot capture long-distance dependencies. Therefore, in order to take into account the advantages of both, many current methods use local-global sequential structures or local-global parallel structures to achieve the fusion of local and global features, forming a more powerful hybrid model to improve performance in medical image segmentation.

[0005] However, the above technologies still have the following problems when used for medical image segmentation: (1) The fusion of local features and global features only adopts a simple sequential structure or parallel structure, lacking the two-way interaction of the two types of information and failing to fully and effectively reveal the complex relationship between local information and global information.

[0006] (2) In medical images, the accuracy of boundaries directly affects the quality of segmentation results and the reliability of clinical diagnosis. However, existing models lack supervision of boundaries.

[0007] (3) Medical images containing noise, incorrect annotations, or low quality may cause the model to learn some non-robust features. Current methods simply separate and discard these non-robust features, lacking correction for non-robust features, which may lead to the loss of useful information.

[0008] (4) The size of lesions in medical images may vary greatly, from very small lesions to lesions covering larger areas. However, current single-scale models often cannot effectively segment all lesion regions at the same time. Summary of the invention

[0009] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a medical image segmentation method and system based on multi-scale feature fusion and correction, which extracts features of the input medical image at different scales, captures feature information from large to small and from coarse to fine, and better understands and processes objects and details at various scales; generates boundary embedding, realizes boundary supervision, and better deals with the problem of blurred boundaries in lesion areas; corrects non-robust features, makes full use of potential useful clues for model prediction, and enhances the robustness of the model.

[0010] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of the present invention provides a medical image segmentation method based on multi-scale feature fusion and correction; Medical image segmentation methods based on multi-scale feature fusion and correction, including: Obtaining the medical image to be segmented and performing preprocessing; Inputting the preprocessed medical image into the PVT network to obtain the first medical image feature maps of different scales; Inputting the first medical image feature map into the bidirectional interactive fusion module and the boundary supervision module respectively, obtaining the second medical image feature map after bidirectional interactive fusion of local information and global information and the boundary embedding corresponding to the second medical image feature map; Inputting a high-level feature map in the second medical image feature map into a feature separation and correction module to obtain a corrected third medical image feature map; Inputting the third medical image feature map, the uncorrected second medical image feature map, and the boundary embedding into a cross-scale boundary guidance module to generate a fourth medical image feature map with enhanced boundary attention; The acquired fourth medical image feature map is processed using the predictor to obtain a final medical image segmentation result.

[0011] As a further technical solution, the preprocessing process includes: using data enhancement to expand the medical image to be segmented; the data enhancement includes: horizontal flipping, vertical flipping, translation, scaling, rotation, and adding Gaussian noise to the data set.

[0012] As a further technical solution, the process of using the bidirectional interactive fusion module to obtain the second medical image feature map after bidirectional interactive fusion of local information and global information is as follows: The bidirectional interactive fusion module comprises a global feature extraction layer, a local feature extraction layer and a bidirectional interactive fusion layer; wherein the global feature extraction layer adopts a fine-grained downsampling strategy to downsample the key vector and the value vector of the first medical image feature map, and obtains the global features of the first medical image feature map through a multi-head self-attention module; The local feature extraction layer uses deep convolution to extract local information of the first medical image feature map, generates context-aware weights through a Sigmoid function and combines the local information to generate local features; The bidirectional interactive fusion layer multiplies the global features by the local features point by point after passing through the Sigmoid function; multiplies the local features by the global features point by point after passing through the Sigmoid function; combines the results of the two point-by-point multiplications, and obtains a second medical image feature map through linear projection.

[0013] As a further technical solution, the process of obtaining boundary embedding using the boundary supervision module is: Normalizing and residually connecting the first medical image feature map using the multi-head self-attention layer and the multi-layer perceptron layer of the boundary supervision module to obtain the first medical image feature map containing semantic information; Classifying the first medical image feature map containing semantic information through a linear predictor to obtain a boundary key point map; Multiplying the boundary key point map and the first medical image feature map containing semantic information point by point and summing them up to obtain an enhanced first medical image feature map; The enhanced first medical image feature map and the randomly initialized boundary embedding are input into the Transformer decoder together to obtain the boundary embedding containing boundary information.

[0014] As a further technical solution, the high-level feature map in the second medical image feature map is input into the feature separation and correction module to obtain the corrected third medical image feature map, specifically: The feature separation and correction module includes a feature separation layer and a feature correction layer; the feature separation layer obtains a robust feature map and a non-robust feature map of a high-level feature map in a feature map of a second medical image through a separation network; Use the feature correction layer to correct the non-robust feature map; The robust feature map and the corrected non-robust feature map are added element by element to obtain a corrected third medical image feature map.

[0015] As a further technical solution, a loss function is also provided in the feature separation correction module; Among them, the loss function is:

[0016]

[0017]

[0018]

[0019] In the formula, is the total loss of the feature separation correction module; To supervise the robust feature map in the feature separation layer; Supervision of non-robust feature maps in the feature separation layer; Supervision for feature correction layer; and Represent the Dice loss function and the cross entropy function respectively, and Represents the true labels and predicted segmentation result maps of different scales respectively; , , They respectively represent the robust feature map, non-robust feature map and third medical image feature map generated by the feature separation and correction module.

[0020] As a further technical solution, the third medical image feature map, the uncorrected second medical image feature map and the boundary embedding are input into a cross-scale boundary guidance module to generate a fourth medical image feature map with enhanced boundary attention, specifically: Inputting the uncorrected second medical image feature map and the corresponding boundary embedding as well as the third medical image feature map and the corresponding boundary embedding into the multi-head attention layer of the cross-scale boundary guidance module and performing normalization and residual connection to output the feature map; The feature map and boundary embedding of the lower scale and the feature map and boundary embedding of the higher scale are input into the multi-head attention and normalized and residually connected to output the feature map; the boundary embedding at one scale is used to gradually refine the features at another scale, provide complementary boundary knowledge, and generate a fourth medical image feature map with enhanced boundary attention.

[0021] A second aspect of the present invention provides a medical image segmentation system based on multi-scale feature fusion and correction.

[0022] Medical image segmentation system based on multi-scale feature fusion and correction, including: The medical image acquisition module is configured to: acquire the medical image to be segmented and perform preprocessing; A first medical image feature map acquisition module is configured to: input the preprocessed medical image into the PVT network to acquire first medical image feature maps of different scales; The second medical image feature map and boundary embedding acquisition module is configured to: input the first medical image feature map into the bidirectional interactive fusion module and the boundary supervision module respectively, and obtain the second medical image feature map after the bidirectional interactive fusion of local information and global information and the boundary embedding corresponding to the second medical image feature map; The third medical image feature map acquisition module is configured to: input the high-level feature map in the second medical image feature map into the feature separation and correction module to obtain a corrected third medical image feature map; a fourth medical image feature map acquisition module, configured to: input the third medical image feature map, the uncorrected second medical image feature map and the boundary embedding into the cross-scale boundary guidance module to generate a fourth medical image feature map with enhanced boundary attention; The medical image segmentation result acquisition module is configured to: use the predictor to process the acquired fourth medical image feature map to obtain a final medical image segmentation result.

[0023] The third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the medical image segmentation method based on multi-scale feature fusion and correction as described in the first aspect of the present invention.

[0024] The fourth aspect of the present invention provides an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the medical image segmentation method based on multi-scale feature fusion and correction as described in the first aspect of the present invention are implemented.

[0025] One or more of the above technical solutions have the following beneficial effects: (1) In terms of segmentation effect, the present invention proposes a medical image segmentation model based on multi-scale feature fusion and feature correction. Features are extracted from the input image at different scales to guide the model to capture feature information from large to small and from coarse to fine, so that the model can better understand and process objects and details at various scales; generate boundary embedding, implement boundary supervision, and better deal with the problem of blurred boundaries in lesion areas; correct non-robust features, make full use of potential useful clues predicted by the model, and enhance the robustness of the model.

[0026] (2) In terms of practicality and scalability, the present invention uses a boundary supervision module and a feature separation correction module. This enables it to better capture boundary detail information and high-level semantic information in images, better cope with the interference of noise in medical images, and improve the robustness of the model. In the future, it can be better applied in clinical practice to assist doctors in diagnosis.

[0027] (3) In terms of computational efficiency, the present invention only applies the feature separation correction module at the highest two scales; it uses the pyramid structured PVT as the backbone network, which reduces computational complexity and memory consumption and improves the speed of model operation. The advantages of the additional aspects of the present invention will be partially given in the following description, and partially will become apparent from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0029] Figure 1 This is a flow chart of the method of the first embodiment.

[0030] Figure 2 This is the overall architecture diagram of the first embodiment.

[0031] Figure 3 It is a schematic diagram of data enhancement in the preprocessing process of the first embodiment.

[0032] Figure 4 It is a schematic diagram of the structure of the PVT network in the first embodiment.

[0033] Figure 5 It is a schematic diagram of the bidirectional interactive fusion module in the first embodiment.

[0034] Figure 6 Schematic diagram of the boundary supervision module in the first embodiment.

[0035] Figure 7 It is a schematic diagram of the feature separation and correction module in the first embodiment.

[0036] Figure 8 Schematic diagram of the cross-scale boundary guidance module in the first embodiment.

[0037] Fig. 9 It is a system structure diagram of the second embodiment. DETAILED DESCRIPTION

[0038] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0039] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0040] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0041] Embodiment 1 This embodiment discloses a medical image segmentation method based on multi-scale feature fusion and correction; like Figure 1 and Figure 2 As shown, the medical image segmentation method based on multi-scale feature fusion and correction includes: Step S1, obtaining a medical image to be segmented and performing preprocessing; The preprocessing process includes: using data enhancement to expand the medical image to be segmented; the data enhancement includes: horizontal flipping, vertical flipping, translation, scaling, rotation, and adding Gaussian noise to the data set. Figure 3 ,In this embodiment, the results of expanding medical images through data ,augmentation are demonstrated.

[0042] Step S2, inputting the preprocessed medical image into the PVT network to obtain first medical image feature maps of different scales; Among them, combined Figure 4 , the PVT network is a pyramid structure, which can reduce computational complexity and memory consumption and increase the speed of model operation. PVT has four stages, each of which is used to generate feature maps of different scales. The structure of each stage is similar, the difference is that for stage i, it has i Transformer encoder layers. At the beginning of the i-th stage, the feature map is decomposed into multiple blocks and passes through linear projection and normalization layers. Then, it passes through i Transformer encoder layers. Each Transformer encoder layer passes through the attention layer and the feedforward layer in turn to extract image features.

[0043] The PVT network is used to extract the input medical image features and obtain the first medical image feature maps of four different scales, wherein the first medical image feature map includes four scales, which are respectively: , , , , among which, from arrive The scale gradually decreases and the level gradually increases.

[0044] Step S3, inputting the first medical image feature map into the bidirectional interactive fusion module and the boundary supervision module respectively, obtaining the second medical image feature map after bidirectional interactive fusion of local information and global information and the boundary embedding corresponding to the second medical image feature map; Combination Figure 5, the bidirectional interactive fusion module includes a global feature extraction layer, a local feature extraction layer, and a bidirectional interactive fusion layer; the global feature extraction layer adopts a fine-grained downsampling strategy to downsample the keys and values ​​of the first medical image feature map to minimize the loss of global information. The fine-grained downsampling strategy consists of several basic units. Each unit uses a deep convolution with a kernel size of 5×5 and a stride of 2, followed by a 1×1 convolution; after that, the query vector, value vector, and key vector generated on the first medical image feature map are processed by a multi-head self-attention module to obtain the global features of the first medical image feature map.

[0045] The local feature extraction layer uses a deep convolution to embed context awareness into the convolution. Context awareness refers to the relationship between different elements in the input data learned by the model and other information useful for model prediction, enabling it to extract local representations. Context-aware weights are generated through Sigmoid and combined with deep convolution to aggregate local information and generate local features; The bidirectional interactive fusion layer multiplies the global features with the local features point by point after passing through the Sigmoid function; the local features are multiplied with the global features point by point after passing through the Sigmoid function; finally, these representations are merged again by point-by-point multiplication, and the mixing between channels is achieved through linear projection to obtain the second medical image feature map, where the second medical image feature maps are , , , After completing this two-way interactive fusion process, both the local representation and the global representation contain information about each other, which can better help the model understand the relationship between local information and global information and improve the model's ability to capture information.

[0046] In addition, the first medical image feature map is input into the boundary supervision module, combined with Figure 6 The boundary supervision module consists of a multi-head self-attention layer, a multi-layer perceptron, and a Transformer decoder. The first medical feature map input first enters the multi-head self-attention layer and is normalized and residually connected; then, it enters the multi-layer perceptron and is normalized and residually connected to obtain the first medical image feature map containing semantic information. A linear predictor with Sigmoid activation is used to classify each patch of the first medical image feature map containing semantic information to obtain a predicted boundary key point map. A patch refers to a small area on the feature map, which contains the visual feature representation of the area and is supervised by the boundary supervision map pre-generated by the Canny edge detection algorithm.

[0047] Among them, the process of generating the boundary supervision map is: First, the image is smoothed using a Gaussian filter to reduce the effect of noise on edge detection. The gradient magnitude and direction of the image are obtained by calculating the first-order derivative of each pixel in the image. The gradient magnitude indicates the strength of the edge, while the gradient direction indicates the direction of the edge. The Canny edge detection algorithm checks each pixel along the gradient direction and only retains those pixels with the largest gradient magnitude, i.e., the edges. This means that if a pixel is not a local maximum in the gradient direction, it will be suppressed (i.e., set to zero). In order to determine which edges are real, the Canny algorithm sets a high threshold and a low threshold, the high threshold is used to detect strong edges, and the low threshold is used to detect weak edges. Strong edges are those edges with a gradient magnitude exceeding the high threshold, while weak edges are those edges with a gradient magnitude between the two thresholds. For each strong edge point, the algorithm searches for weak edge points in its neighborhood. If a weak edge point is close enough to a strong edge point (this distance is defined by the neighborhood range of the pixel point, which is a 3x3 area), the weak edge point is considered to be connected to the strong edge point, and therefore the weak edge point is considered to be part of the edge. If the gradient magnitude of a pixel is above the high threshold, it is determined to be a strong edge. If a pixel's gradient magnitude is below the low threshold, it is determined to be a non-edge. For pixels with gradient magnitudes between the two (weak edges), they are determined to be edges only if they are connected to a strong edge.

[0048] After the above process, a boundary supervision map of the lesion area in each medical image is obtained during the model training process. The boundary supervision map can be used to supervise the randomly initialized boundary embedding, allowing it to better learn the boundary information, so that the boundary embedding can guide the model to focus on the correct boundary area of ​​the lesion in the subsequent process.

[0049] Furthermore, the predicted boundary key point map is point-by-point multiplied and summed with the first medical image feature map containing semantic information to obtain an enhanced first medical image feature map. Finally, the enhanced first medical image feature map and the randomly initialized boundary embedding are sent to the Transformer decoder together, and the optimization training is continuously performed to obtain a boundary embedding with rich boundary information. In this embodiment, the randomly initialized boundary embedding refers to a vector randomly generated by the code in advance before the model training. During the model training process, as the model continues to back propagate, these boundary embeddings will truly learn the representation of the boundary of the lesion area, thereby continuously optimizing the feature map. During random initialization, the nn.Embedding() function is used to randomly obtain the value of the vector that constitutes the boundary embedding. The boundary embedding obtained is , , , , respectively corresponding to the generated second medical image feature map.

[0050] Step S4: input the high-level feature map in the second medical image feature map into the feature separation and correction module to obtain a corrected third medical image feature map. and ; Combination Figure 7 The feature separation and correction module includes a feature separation layer and a feature correction layer. A separation network is introduced in the feature separation layer, and the robustness of each feature unit is learned using the separation network. The separation network contains three basic units, each of which consists of a convolutional layer, a batch normalization layer, and a Relu activation layer. The last basic unit only contains a convolutional layer. The high-level feature map in the second medical image feature map is converted into 、 Separate and obtain the robust feature map p, which represents the robustness score of the unit corresponding to f, where a higher score means a stronger robustness of feature activation. Based on the robustness score, the feature map 、 Decompose into robust features and non-robust features. Use Gumbelsoftmax to obtain differentiable soft masks , so that:

[0051] in, is used to normalize the robustness graph function. , Indicates from The sample drawn from the distribution is of the form ,in ,and It is used to control , The influence of temperature coefficient. It is used to indicate the robustness of the feature. t The closer it is to 1, the more robust the feature at that position is. t The closer it is to 0, the lower the robustness of the feature at that location. Get the robustness feature map, Obtain a non-robust feature map. Among them, = , =1- , using two loss functions at the same time and The separated robust features and non-robust features are supervised.

[0052] In the feature correction layer, non-robust features are adjusted to capture additional useful clues. The correction network is introduced. The correction network contains three basic units, each of which consists of a convolutional layer, a batch normalization layer, and a Relu activation layer. The last basic unit only contains a convolutional layer. It uses non-robust feature maps As input, the negative mask Applied to recalibrate the unit. And add the result to To calculate the recalibrated feature map Finally, the robust feature map And the corrected non-robust feature map Perform element-by-element addition to obtain the third medical image feature map Add loss , which is used to supervise the entire feature separation and correction module.

[0053] It should be noted that medical images may contain some noise, incorrect annotations, and some low-quality images, which will cause the model to learn some non-robust features and disrupt the model's learning process. The use of the feature separation and correction module can separate and correct non-robust features, further capture useful discriminative clues, and enhance the robustness of the model.

[0054] Step S5: The third medical image feature map , uncorrected second medical image feature map , and boundary embedding , , , The input is input into a cross-scale boundary guidance module to generate a fourth medical image feature map that strengthens boundary attention; Combination Figure 8 In order to achieve the gradual refinement of the boundary, the neighboring layers are used to provide complementary boundary knowledge. The four feature maps and boundary embeddings are compared in order, and the low-scale feature in each group is defined as , the boundary embedding corresponding to the low-scale features is defined as ; The high-scale features are defined as , the boundary embedding corresponding to the high-scale features is defined as . The low-scale features and boundary embedding And high-scale features and boundary embedding As input, As a query, As key and value input into the multi-head attention and normalized and residual connected to get the output features , you can The boundary knowledge in is transferred to For each point in As a query, As key and value input into the multi-head attention and normalized and residual connected to get the output features , which can effectively integrate low-scale boundary detail information into high-scale feature maps. After upsampling and conduct Then it is input into the linear layer to generate enhanced low-scale features Utilizing boundary embedding at one scale to gradually refine features at another scale can provide complementary boundary knowledge and ultimately generate a fourth medical image feature map with enhanced boundary attention.

[0055] Step S6, using the predictor to process the acquired fourth medical image feature map to obtain a final medical image segmentation result. In this embodiment, the predictor uses a linear classifier, which is composed of a 1x1 convolution layer, a batch normalization layer, a ReLU activation layer and another 1x1 convolution layer, which is used to classify the pixels in the image to obtain the final segmentation result map.

[0056] Furthermore, in this embodiment, a loss function is also used for correction, wherein the total loss function is:

[0057] is the lesion segmentation loss, is the boundary supervision loss, is the total loss of the feature separation correction module.

[0058] in, , , Defined as:

[0059]

[0060]

[0061]

[0062]

[0063]

[0064] and is defined as follows:

[0065]

[0066] in, and Represent the Dice loss function and the cross entropy function respectively, and Represents the true labels and predicted segmentation result maps of different scales respectively. and They represent the boundary point map generated by the boundary supervision module and the boundary point map generated by the real label respectively. , , They respectively represent the robust feature map, non-robust feature map and third medical image feature map generated by the feature separation and correction module.

[0067] Embodiment 2 This embodiment discloses a medical image segmentation system based on multi-scale feature fusion and correction; like Fig. 9 As shown, the medical image segmentation system based on multi-scale feature fusion and correction includes: The medical image acquisition module is configured to: acquire the medical image to be segmented and perform preprocessing; A first medical image feature map acquisition module is configured to: input the preprocessed medical image into the PVT network to acquire first medical image feature maps of different scales; The second medical image feature map and boundary embedding acquisition module is configured to: input the first medical image feature map into the bidirectional interactive fusion module and the boundary supervision module respectively, and obtain the second medical image feature map after the bidirectional interactive fusion of local information and global information and the boundary embedding corresponding to the second medical image feature map; The third medical image feature map acquisition module is configured to: input the high-level feature map in the second medical image feature map into the feature separation and correction module to obtain a corrected third medical image feature map; a fourth medical image feature map acquisition module, configured to: input the third medical image feature map, the uncorrected second medical image feature map and the boundary embedding into the cross-scale boundary guidance module to generate a fourth medical image feature map with enhanced boundary attention; The medical image segmentation result acquisition module is configured to: use the predictor to process the acquired fourth medical image feature map to obtain a final medical image segmentation result.

[0068] Embodiment 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0069] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the medical image segmentation method based on multi-scale feature fusion and correction as described in Example 1.

[0070] Embodiment 4 The purpose of this embodiment is to provide an electronic device.

[0071] An electronic device comprises a memory, a processor and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the medical image segmentation method based on multi-scale feature fusion and correction as described in Example 1 are implemented.

[0072] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0073] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0074] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A medical image segmentation method based on multi-scale feature fusion and correction, characterized in that: include: Obtaining the medical image to be segmented and performing preprocessing; Inputting the preprocessed medical image into the PVT network to obtain the first medical image feature maps of different scales; Inputting the first medical image feature map into the bidirectional interactive fusion module and the boundary supervision module respectively, obtaining the second medical image feature map after bidirectional interactive fusion of local information and global information and the boundary embedding corresponding to the second medical image feature map; Inputting a high-level feature map in the second medical image feature map into a feature separation and correction module to obtain a corrected third medical image feature map; Inputting the third medical image feature map, the uncorrected second medical image feature map, and the boundary embedding into a cross-scale boundary guidance module to generate a fourth medical image feature map with enhanced boundary attention; The acquired fourth medical image feature map is processed using the predictor to obtain a final medical image segmentation result.

2. The medical image segmentation method based on multi-scale feature fusion and correction as claimed in claim 1, characterized in that: The preprocessing process includes: expanding the medical image to be segmented by using data enhancement; the data enhancement includes: horizontal flipping, vertical flipping, translation, scaling, rotation, and adding Gaussian noise to the data set.

3. The medical image segmentation method based on multi-scale feature fusion and correction as claimed in claim 1, characterized in that: The process of using the bidirectional interactive fusion module to obtain the second medical image feature map after bidirectional interactive fusion of local information and global information is as follows: The bidirectional interactive fusion module includes a global feature extraction layer, a local feature extraction layer and a bidirectional interactive fusion layer; The global feature extraction layer adopts a fine-grained downsampling strategy to downsample the key vector and value vector of the first medical image feature map, and obtains the global features of the first medical image feature map through a multi-head self-attention module; The local feature extraction layer uses deep convolution to extract local information of the first medical image feature map, generates context-aware weights through a Sigmoid function and combines the local information to generate local features; The bidirectional interactive fusion layer multiplies the global features by the local features point by point after passing through the Sigmoid function; multiplies the local features by the global features point by point after passing through the Sigmoid function; combines the results of the two point-by-point multiplications, and obtains a second medical image feature map through linear projection.

4. The medical image segmentation method based on multi-scale feature fusion and correction as claimed in claim 1, characterized in that: The process of obtaining boundary embedding using the boundary supervision module is: Normalizing and residually connecting the first medical image feature map using the multi-head self-attention layer and the multi-layer perceptron layer of the boundary supervision module to obtain the first medical image feature map containing semantic information; Classifying the first medical image feature map containing semantic information through a linear predictor to obtain a boundary key point map; Multiplying the boundary key point map and the first medical image feature map containing semantic information point by point and summing them up to obtain an enhanced first medical image feature map; The enhanced first medical image feature map and the randomly initialized boundary embedding are input into the Transformer decoder together to obtain the boundary embedding containing boundary information.

5. The medical image segmentation method based on multi-scale feature fusion and correction as claimed in claim 1, characterized in that: The high-level feature map in the second medical image feature map is input into the feature separation and correction module to obtain the corrected third medical image feature map, specifically: The feature separation and correction module includes a feature separation layer and a feature correction layer; the feature separation layer obtains a robust feature map and a non-robust feature map of a high-level feature map in a feature map of a second medical image through a separation network; Use the feature correction layer to correct the non-robust feature map; The robust feature map and the corrected non-robust feature map are added element by element to obtain a corrected third medical image feature map.

6. The medical image segmentation method based on multi-scale feature fusion and correction as claimed in claim 1, characterized in that: The feature separation and correction module is also provided with a loss function; Among them, the loss function is: In the formula, is the total loss of the feature separation correction module; To supervise the robust feature map in the feature separation layer; Supervision of non-robust feature maps in the feature separation layer; Supervision for feature correction layer; and Represent the Dice loss function and the cross entropy function respectively, and Represents the true labels and predicted segmentation result maps of different scales respectively; , , They respectively represent the robust feature map, non-robust feature map and third medical image feature map generated by the feature separation and correction module.

7. The medical image segmentation method based on multi-scale feature fusion and correction as claimed in claim 1, characterized in that: The third medical image feature map, the uncorrected second medical image feature map and the boundary embedding are input into the cross-scale boundary guidance module to generate a fourth medical image feature map with enhanced boundary attention, specifically: Inputting the uncorrected second medical image feature map and the corresponding boundary embedding as well as the third medical image feature map and the corresponding boundary embedding into the multi-head attention layer of the cross-scale boundary guidance module and performing normalization and residual connection to output the feature map; Input the feature map and boundary embedding of the lower scale and the feature map and boundary embedding of the higher scale into the multi-head attention and perform normalization and residual connection to output the feature map; Boundary embedding at one scale is used to gradually refine features at another scale, providing complementary boundary knowledge and generating a fourth medical image feature map with enhanced boundary attention.

8. A medical image segmentation system based on multi-scale feature fusion and correction, characterized by: include: The medical image acquisition module is configured to: acquire the medical image to be segmented and perform preprocessing; A first medical image feature map acquisition module is configured to: input the preprocessed medical image into the PVT network to acquire first medical image feature maps of different scales; The second medical image feature map and boundary embedding acquisition module is configured to: input the first medical image feature map into the bidirectional interactive fusion module and the boundary supervision module respectively, and obtain the second medical image feature map after the bidirectional interactive fusion of local information and global information and the boundary embedding corresponding to the second medical image feature map; The third medical image feature map acquisition module is configured to: input the high-level feature map in the second medical image feature map into the feature separation and correction module to obtain a corrected third medical image feature map; a fourth medical image feature map acquisition module, configured to: input the third medical image feature map, the uncorrected second medical image feature map and the boundary embedding into the cross-scale boundary guidance module to generate a fourth medical image feature map with enhanced boundary attention; The medical image segmentation result acquisition module is configured to: use the predictor to process the acquired fourth medical image feature map to obtain a final medical image segmentation result.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the medical image segmentation method based on multi-scale feature fusion and correction as described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the medical image segmentation method based on multi-scale feature fusion and correction as described in any one of claims 1 to 7 are implemented.