Medical image segmentation method and device, electronic equipment and storage medium

By combining CNN and Transformer networks in medical image segmentation, and using the Transformer network to extract global features and the CNN network layer to extract local features, the problems of insufficient global semantic modeling capability and high computational cost in the hybrid architecture are solved, and more efficient medical image segmentation is achieved.

CN120833487APending Publication Date: 2025-10-24NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510995457.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing hybrid CNN-Transformer architectures in medical image segmentation suffer from insufficient global semantic modeling capabilities and high computational overhead due to over-reliance on CNN encoders.

Method used

A method combining a CNN network with multiple independent Transformer networks is adopted. Global features are extracted by the Transformer network and input into the corresponding CNN network layer for local feature extraction. This ensures that the feature map scale of the Transformer network is the same as that of the CNN network, avoiding over-reliance on local features and reducing computational overhead without adding a parameter fusion module.

Benefits of technology

It improves the accuracy of medical image segmentation and reduces the consumption of computing resources, achieving more efficient feature extraction and segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833487A_ABST
    Figure CN120833487A_ABST
Patent Text Reader

Abstract

The invention provides a medical image segmentation method and device, electronic equipment and a storage medium, a used target image segmentation model comprises a plurality of CNN network layers and a plurality of independent Transform networks, the Transform networks are in one-to-one correspondence with the CNN network layers, the feature map scale of the Transform networks is the same as that of the corresponding CNN network layers, and in the training process, the feature map scale of the Transform networks is smaller than that of the CNN network layers. According to the method, all Transform networks are used for extracting global features of different scales of training images, the global features are input to corresponding CNN network layers, local features are extracted under the guidance of the global features, the CNN is prevented from only paying attention to the local features, and the accuracy and integrity of CNN network feature extraction and the accuracy of target image segmentation are improved. The feature map scale of the Transform network is the same as that of the corresponding CNN network layer, so that an additional parameter fusion module is not needed, and the resource overhead is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a medical image segmentation method and device, an electronic device and a storage medium. BACKGROUND

[0002] Medical image segmentation is a key step in computer-aided diagnosis and surgery and radiotherapy planning and implementation. Medical image segmentation refers to dividing specific regions or structures in a medical image, such as tumors, organs, blood vessels, etc., which helps doctors diagnose and treat more accurately. In the related art, medical image segmentation is usually performed by an AI model, such as a U-Net model, a Transformer model, etc.

[0003] Although the U-shaped structure (such as UNet and its variants) realizes multi-scale feature aggregation through hierarchical design and performs well in image segmentation tasks, it relies on local convolution operations, which fundamentally limits the extraction of long-distance dependencies.

[0004] Vision Transformer (ViTs) has become an important alternative solution for modeling long-range dependencies in visual tasks. However, it often faces challenges in local feature extraction, which can be alleviated by introducing a convolution Stem structure in the initial visual processing stage. This complementary relationship has promoted the development of hybrid CNN-Transformer architectures.

[0005] Although hybrid CNN-Transformer architectures have shown good performance, there are still two core limitations in current implementations: the strategy of injecting local features extracted by CNN into the Transformer encoder to compensate for the lack of local modeling ability of Transformer improves the global representation ability of ViT, but also leads to excessive dependence on the inductive bias of the CNN encoder, which affects the ability of the Transformer to independently model global semantics, thereby affecting image segmentation accuracy. In addition, although CNN-Transformer hybrid architectures can achieve the performance of pure Transformer methods, the complex structure design and fusion bottleneck, especially the complex fusion modules introduced to improve performance, make the computational overhead of this hybrid architecture significantly higher than that of the Transformer-based U-shaped architecture. SUMMARY

[0006] Therefore, the embodiments of the present application provide a medical image segmentation method and device, an electronic device and a storage medium to improve the accuracy of medical image segmentation and reduce the computational overhead.

[0007] According to an aspect of the present application, a medical image segmentation method is provided, the method comprising: obtaining a target medical image; inputting the target medical image into a pre-trained target image segmentation model, and obtaining a target segmentation result output by the target image segmentation model for the target medical image; wherein the target image segmentation model is pre-trained by the following steps: obtaining a training data set, the training data set containing a plurality of training images, each training image corresponding to a segmentation label, the segmentation label being used to identify the position of a target region in the training image; inputting the training data set into an initial image segmentation model, the image segmentation model comprising a CNN network and a plurality of independent Transformer networks, the CNN network comprising a plurality of CNN network layers; the Transformer network corresponds to the CNN network layer one by one, and the feature map scale of the Transformer network is the same as that of the corresponding CNN network; for each training image, extracting a plurality of global features of the training image through each Transformer network, and inputting the plurality of global features into the CNN network layer corresponding to the Transformer network; using each CNN network layer to extract features based on the training image and the global features output by the corresponding Transformer network of the CNN network layer, to obtain a plurality of local features; stacking a plurality of local features to obtain a target feature of the training image, and outputting a predicted segmentation result of the training image based on the target feature; training the image segmentation model based on the difference between the predicted segmentation result of each training image and the segmentation label of each training image; in the case where the difference converges, determining the current image segmentation model as the target image segmentation model.

[0008] In a possible embodiment, the image segmentation model further comprises a pyramid input network, the pyramid input network comprising a plurality of convolution layers; the method further comprises: using the pyramid input network to downsample each training image to obtain a plurality of downsampled images of each training image; inputting the downsampled images of corresponding size into each Transformer network based on the feature map scale of the CNN network layer corresponding to each Transformer network.

[0009] In a possible embodiment, the CNN network comprises a CNN encoder and a CNN decoder, the CNN encoder comprises an encoding convolutional block and a plurality of encoding residual blocks connected in series, the CNN decoder comprises a decoding convolutional block and a plurality of decoding residual blocks, one encoding residual block and its corresponding decoding residual block form a CNN network layer, and the decoding convolutional block and each decoding residual block are connected to a prediction head; The feature extraction is performed based on the training image and global features output by the corresponding Transformer network of the CNN network layer, and a plurality of local features are obtained, including: The training image is input into the CNN encoder, the first local feature of the training image is extracted by using the encoding convolutional block, and the first local feature is input into the encoding residual block connected to the encoding convolutional block; The first local feature and the first global feature output by the Transformer network corresponding to the encoding residual block connected to the encoding convolutional block are spliced to obtain a first fusion feature; The first fusion feature is subjected to local feature extraction by using the encoding residual block connected to the encoding convolutional block to obtain a second local feature; The second local feature is input into a next-level encoding residual block, so that the next-level encoding residual block splices the second local feature and global features input by the Transformer network corresponding to the next-level encoding residual block, and performs local feature extraction on the spliced features to obtain a new second local feature, and the step of inputting the second local feature into the next-level encoding residual block is returned until all the encoding residual blocks perform local feature extraction; The first local feature and each second local feature obtained by performing feature extraction on each encoding residual block are input into the encoding convolutional block, the decoding convolutional block corresponding to each encoding residual block, and each decoding residual block to obtain a plurality of local features of the training image output by the CNN decoder; The plurality of local features are superimposed to obtain the target feature of the training image, including: The plurality of local features of the training image output by the CNN decoder are input into the corresponding prediction head, and the features output by each prediction head are superimposed to obtain the target feature of the training image.

[0010] In a possible embodiment, patch parameters of the Transformer network satisfy the following condition: RCi / 2 < Patch size < RCi, where RCi is a receptive field of a CNN network layer i corresponding to the Transformer network, the receptive field is related to a convolution kernel size of the CNN network layer, and the patch size is the patch parameter of the Transformer network.

[0011] In a possible embodiment, the image segmentation model further comprises a Stem network corresponding to each Transformer network; and the method further comprises: inputting the plurality of down-sampled images into the Stem network corresponding to the Transformer network, and extracting reference local features of the down-sampled images through the Stem network; inputting the reference local features into the Transformer network, so that the Transformer network extracts a plurality of global features of the training image based on the reference local features and the down-sampled images.

[0012] In a possible embodiment, the image segmentation model further comprises a strip attention module corresponding to each Transformer network and receiving an input of the corresponding Transformer network; and the method further comprises: the Transformer network inputs the global features into the strip attention module corresponding to the Transformer network; the strip attention module performs strip attention calculation based on the global features of the training image, a previous frame image and a next frame image of the training image in a time sequence, to obtain a strip feature map; the strip feature map is input into the CNN network layer corresponding to the Transformer network.

[0013] According to another aspect of the present application, a medical image segmentation device is provided, which comprises: an acquisition module configured to acquire a target medical image; an input module configured to input the target medical image into a pre-trained target image segmentation model, and acquire a target segmentation result output by the target image segmentation model for the target medical image; a training module configured to pre-train the target image segmentation model by the following steps: obtaining a training data set, the training data set comprising a plurality of training images, each of the training images corresponding to a segmentation label used to identify a location of a target region in the training image; inputting the training data set into an initial image segmentation model, the image segmentation model comprising a CNN network and a plurality of independent Transformer networks, the CNN network comprising a plurality of CNN network layers, each of the Transformer networks corresponding to one of the CNN network layers, and a feature map scale of each of the Transformer networks being the same as a feature map scale of the corresponding CNN network layer; for each of the training images, extracting a plurality of global features of the training image by each of the Transformer networks, and inputting the plurality of global features into the corresponding CNN network layer of the Transformer network; performing feature extraction on the training image and the global features output by the corresponding Transformer network of the CNN network layer based on each of the CNN network layers to obtain a plurality of local features; stacking the plurality of local features to obtain a target feature of the training image, and outputting a predicted segmentation result of the training image based on the target feature; training the image segmentation model based on a difference between the predicted segmentation result of each of the training images and the segmentation label of each of the training images; in a case where the difference converges, determining a current image segmentation model as a target image segmentation model.

[0014] In a possible embodiment, the image segmentation model further comprises a pyramid input network, the pyramid input network comprising a plurality of convolution layers; the training module is further configured to perform down-sampling on each of the training images by using the pyramid input network to obtain a plurality of down-sampled images of each of the training images; inputting a down-sampled image of a corresponding size into each of the Transformer networks based on a feature map scale of the corresponding CNN network layer of each of the Transformer networks; the CNN network comprises a CNN encoder and a CNN decoder, the CNN encoder comprising an encoding convolution block and a plurality of encoding residual blocks connected in series, the CNN decoder comprising a decoding convolution block and a plurality of decoding residual blocks, one encoding residual block and its corresponding decoding residual block forming one CNN network layer, and an output end of the decoding convolution block and each of the decoding residual blocks being connected to a prediction head; The CNN network layer based on the training image and the global feature output by the corresponding Transformer network of the CNN network layer is used for feature extraction to obtain a plurality of local features, including: The training image is input into the CNN encoder, the first local feature of the training image is extracted by using the encoding convolution block, and the first local feature is input into the encoding residual block connected with the encoding convolution block; The first local feature and the first global feature output by the Transformer network corresponding to the encoding residual block connected with the encoding convolution block are spliced to obtain a first fusion feature; The first fusion feature is locally feature-extracted by using the encoding residual block connected with the encoding convolution block to obtain a second local feature; The second local feature is input into a next-level encoding residual block, so that the next-level encoding residual block splices the second local feature and the global feature input by the Transformer network corresponding to the next-level encoding residual block, and locally feature-extracts the spliced feature to obtain a new second local feature, and returns to the step of inputting the second local feature into the next-level encoding residual block until all encoding residual blocks are locally feature-extracted; The first local feature and each second local feature obtained by feature extraction of each encoding residual block are input into the encoding convolution block and the decoding convolution block and each decoding residual block corresponding to each encoding residual block to obtain a plurality of local features of the training image output by the CNN decoder; The plurality of local features are superimposed to obtain the target feature of the training image, including: The plurality of local features of the training image output by the CNN decoder are input into the corresponding prediction head, and the features output by each prediction head are superimposed to obtain the target feature of the training image; The patch parameter of the Transformer network satisfies the following condition: RCi / 2 The image segmentation model further includes a Stem network, and the Stem network corresponds to the Transformer network one by one; the training module is further configured to input a plurality of the down-sampled images into the Stem network corresponding to the corresponding Transformer network, and extract reference local features of the down-sampled images through the Stem network. Inputting the reference local features into a corresponding Transformer network, so that the Transformer network extracts a plurality of global features of the training image based on the reference local features and the downsampled image; The image segmentation model also includes a strip attention module, which corresponds one-to-one to the Transformer network and receives input from the corresponding Transformer network; the training module is also used for the Transformer network to input global features into the corresponding strip attention module; The stripe attention module performs stripe attention calculation based on the global features of the training image, the previous frame image and the next frame image of the training image in the time series to obtain a stripe feature map; The strip feature map is input into the CNN network layer corresponding to the corresponding Transformer network.

[0015] According to another aspect of the present invention, there is provided an electronic device, comprising: processor; and Memory for storing programs, The program includes instructions, which, when executed by the processor, enable the processor to perform any of the above-mentioned medical image segmentation methods.

[0016] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute any of the above-mentioned medical image segmentation methods.

[0017] One or more technical solutions provided in the embodiments of the present application input the target medical image to be segmented into a pre-trained target image segmentation model, and obtain the target image segmentation result output by the target image segmentation model. The target image segmentation model includes a CNN network and a plurality of independent Transformer networks. The CNN network includes a plurality of CNN network layers, and the plurality of independent Transformer networks correspond one-to-one to the CNN network layers. The feature map scale of the Transformer network is the same as the feature map scale of the corresponding CNN network. During the training process of the target image segmentation model, the global features of different scales of a training image are first extracted by using each independent Transformer network, and the global features extracted by each independent Transformer network are input into the corresponding CNN network layer, so that the CNN network layer extracts local features under the guidance of the global features, avoids the CNN from only focusing on the local features itself during the process of extracting local features, and thus improves the accuracy and completeness of the feature extraction of the CNN network, and also improves the accuracy of the subsequent target image segmentation model. In addition, since the feature map scale of the Transformer network is the same as the feature map scale of the corresponding CNN network, the size of the Transformer is aligned with the receptive field of the CNN, and thus the global features extracted by the Transformer can guide the CNN without increasing the parameter fusion module. Compared with the existing model which needs to increase the parameter fusion module, the resource consumption is greatly reduced. BRIEF DESCRIPTION OF DRAWINGS

[0018] In the following description of exemplary embodiments in conjunction with the accompanying drawings, more details, features and advantages of the present application are disclosed in the drawings, in which: Figure 1 A flowchart of a medical image segmentation method provided by the embodiments of the present application is shown. Figure 2 A flowchart of a target image segmentation model training process provided in the embodiments of the present application is shown. Figure 3 A structure diagram of a CNN network in a target image segmentation model provided by the embodiments of the present application is shown. Figure 4 A flowchart of a medical image segmentation method provided by the embodiments of the present application is shown. Figure 5 A flowchart of a medical image segmentation device provided by the embodiments of the present application is shown. Figure 6 A structural block diagram of an exemplary electronic device capable of implementing the embodiments of the present application is shown. DETAILED DESCRIPTION

[0019] Embodiments of the present application will be described in more detail with reference to the drawings. While several embodiments of the application are shown in the drawings, it is understood that the application can be practiced by many other forms that will be readily apparent to those skilled in the art and that the scope of the application is defined not by the embodiments depicted but by the claims that follow. It is also understood that the drawings are not necessarily to scale and that exaggerated details of an embodiment can be shown for illustrative purposes.

[0020] It should be understood that each of the steps in the method embodiments of the present application can be performed in a different order and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present application is not limited in this respect.

[0021] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising, but not limited to." The term "based on" is "based at least in part on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments." Related terms are defined in the following description. It should be noted that reference to a "first," "second," etc. concept does not limit the function of these devices, modules, or units but is used to differentiate between different devices, modules, or units.

[0022] It should be noted that the terms "a" and "an" and "the" and similar referents used in the context of the present application are to be construed to cover both singular as well as plural and are used interchangeably with "one or more."

[0023] The names of the messages or information exchanged between the devices in the embodiments of the present application are used only for illustrative purposes and are not intended to limit the scope of the messages or information.

[0024] In order to improve the accuracy of medical image segmentation, the present application provides a medical image segmentation method, device, electronic equipment and storage medium. The medical image segmentation method provided by the present application can be applied to any electronic equipment with medical image segmentation function. The electronic equipment can be a computer or a mobile terminal, etc. The scheme of the present application will be described below with reference to the drawings: Figure 1 A flowchart of a medical image segmentation method provided by an embodiment of the present application can include the following steps: S101, obtaining a target medical image; S102, input the target medical image into a pre-trained target image segmentation model, and obtain a target segmentation result output by the target image segmentation model for the target medical image.

[0025] The target medical image segmentation model is pre-trained, as shown in the following S201-S207. Figure 2 The target medical image segmentation model can be pre-trained by the following S201-S207: S201, obtain a training data set, the training data set containing a plurality of training images, each training image corresponding to a segmentation label, the segmentation label being used to identify the position of a target region in the training image; S202, input the training data set into an initial image segmentation model, the image segmentation model including a CNN network and a plurality of independent Transformer networks, the CNN network including a plurality of CNN network layers; the Transformer network corresponds to the CNN network layer one by one, and the feature map scale of the Transformer network is the same as that of the corresponding CNN network; S203, for each training image, extracting a plurality of global features of the training image by each Transformer network, and inputting the plurality of global features into the CNN network layer corresponding to the Transformer network; S204, using each CNN network layer to extract features based on the training image and the global features output by the corresponding Transformer network of the CNN network layer, to obtain a plurality of local features; S205, superimposing a plurality of local features to obtain a target feature of the training image, and outputting a predicted segmentation result of the training image based on the target feature; S206, training the image segmentation model based on the difference between the predicted segmentation result of each training image and the segmentation label of each training image; S207, in the case where the difference converges, determining the current image segmentation model as a target image segmentation model.

[0026] In the embodiment of the present application, the target medical image to be segmented is input into the pre-trained target image segmentation model, and the target image segmentation result output by the target image segmentation model is obtained. The target image segmentation model includes a CNN network and a plurality of independent Transformer networks. The CNN network includes a plurality of CNN network layers. The plurality of independent Transformer networks correspond to the CNN network layers one by one, and the feature map scale of the Transformer network is the same as that of the corresponding CNN network. During the training of the target image segmentation model, the global features of different scales of the training image are first extracted by using each independent Transformer network, and the global features extracted by each independent Transformer network are input into the corresponding CNN network layer, so that the CNN network layer extracts local features under the guidance of the global features, avoids the CNN from only focusing on the local features itself during the process of extracting local features, and thus improves the accuracy and completeness of the feature extraction of the CNN network, and also improves the accuracy of the subsequent target image segmentation model. In addition, since the feature map scale of the Transformer network is the same as that of the corresponding CNN network, the size of the Transformer is aligned with the receptive field of the CNN, and thus the global features extracted by the Transformer can guide the CNN without increasing the parameter fusion module. Compared with the existing model that needs to increase the parameter fusion module, the resource consumption is greatly reduced.

[0027] The S101-S102 and S201-S207 are exemplarily described as follows: In the present application, the target medical image to be segmented can be obtained by any feasible way, such as a medical image derived from a medical device, or a CT (Computed Tomography), MRI (Magnetic Resonance Imaging) image provided by a user, etc. The target medical image can include organs, lesion regions, blood vessels, etc. After obtaining the target medical image, it can be input into the target image segmentation model to obtain the target image segmentation result output by the target image segmentation model. The target image segmentation result can be defined according to the actual application scenario. For example, the target image segmentation result can include the position information of the target anatomical structure or the lesion region, the size information of the lesion region, etc.

[0028] The target image segmentation model is a pre-trained model, which can be trained by S201-S207. The training data set used in the training process of the target image segmentation model can be obtained by any feasible method, for example, the training data set can be obtained from open source data sets, internal historical data, etc. The training data set contains multiple training images, each of which can include organs, lesion regions, or blood vessels, etc. and each training image corresponds to a segmentation label, which can be the position information of the target region in the training image, i.e. the region including the above organs, lesion regions or blood vessels in the training image.

[0029] After obtaining the training data set, the training data set can be input into an initial image segmentation model, which can include a CNN network (Convolutional Neural Network, CNN) and multiple independent Transformer networks (Isolated Hierarchical Transformer, IHT). The CNN network of the image segmentation model contains multiple CNN network layers, the number of which can be set according to the actual application scenario, such as 4. As a possible implementation, the CNN network can be a U-shaped architecture of U-Net, i.e. it can include multiple CNN encoders and CNN decoders, each pair of CNN encoder and CNN decoder forming a CNN network layer.

[0030] In one possible embodiment, the CNN network includes a CNN encoder and a CNN decoder, the CNN encoder includes an encoding convolution block and multiple encoding residual blocks connected in series, and the CNN decoder includes a decoding convolution block and multiple decoding residual blocks. One encoding residual block and its corresponding decoding residual block form a CNN network layer, and the output ends of the decoding convolution block and each decoding residual block are connected to a prediction head. The feature extraction based on the training image and the global features output by the corresponding Transformer network of each CNN network layer includes: S31, input the training image into the CNN encoder, extract the first local feature of the training image by the encoding convolution block, and input the first local feature into the encoding residual block connected with the encoding convolution block; S32, splice the first local feature and the first global feature output by the Transformer network corresponding to the encoding residual block connected with the encoding convolution block to obtain the first fusion feature; S33, locally extracting features from the first fusion feature by using an encoding residual block connected to the encoding convolution block to obtain a second local feature; S33, inputting the second local feature into a next-level encoding residual block, so that the next-level encoding residual block concatenates the second local feature and a global feature input into a Transformer network corresponding to the next-level encoding residual block, and locally extracts features from the concatenated features to obtain a new second local feature, and returning to the step of inputting the second local feature into the next-level encoding residual block until all encoding residual blocks locally extract features; S34, inputting the first local feature and each second local feature obtained by extracting features from each encoding residual block into an encoding convolution block and a decoding convolution block corresponding to each encoding residual block and each decoding residual block to obtain a plurality of local features of the training image output by the CNN decoder; The superimposing of the plurality of local features to obtain the target feature of the training image comprises: Inputting the plurality of local features of the training image output by the CNN decoder into corresponding prediction heads, and superimposing features output by each prediction head to obtain the target feature of the training image.

[0031] As a possible implementation, the CNN encoder specifically is a ResNet-based path, which can include one convolution block (CB) and three residual blocks (RBs). The convolution block can include two convolution layers with a kernel size of 3x3 for extracting initial low-level features. The residual blocks can include a ResNet18 base block for extracting adjacent features of different scale images. The spatial resolution of the feature map output by the residual blocks decreases step by step, and the number of channels increases step by step. Specifically, the output of each residual block can be input into a pooling layer (maxpooling). The pooling layer is used for downsampling, which can sample the feature map output by the residual block to half of the feature map. Specifically, the length and width of the feature map can be reduced to half of the length and width of the feature map input into the pooling layer.

[0032] In a possible embodiment, the CNN encoder can further include a separate ResNet residual block for further compressing spatial information and enhancing the expression of high-order semantic features. For example, the separate ResNet residual block can output a feature map with a spatial resolution of H / 16xW / 16 and a channel number of 512, where H and W are the height and width of the input training image. That is, in the CNN encoder, the input training image will go through the CB module convolution and the feature extraction of the three residual modules, and the feature extraction of the separate residual block. The channel number of the output feature map is 64-128-256-512-512, respectively.

[0033] The features extracted by each module in the CNN encoder will be input into the corresponding module of the CNN decoder. Based on the above example, the CNN decoder can also include three residual blocks and one convolution block. For the sake of convenience, the residual blocks contained in the same CNN network layer are marked with the same serial number in this article. For example, for the third CNN network layer, the residual blocks contained therein can be named the third encoding residual block and the third decoding residual block.

[0034] The above multiple independent Transformer networks can be specifically independent ViT (Vision Transformer) modules, each of which is a single-head self-attention module. Specifically, the single-head self-attention module can include a patch embedding layer and a self-attention layer, wherein the patch embedding layer can process the input image according to the preset patch size. Specifically, the patch embedding layer can divide the input features into (H×W) / 4 / p s=2 2 patches, where s is the layer identifier of the CNN network layer corresponding to the Transformer network, and p represents the patch size. For example, if a CNN encoder contains one convolutional block and three residual blocks, then the CNN encoder contains four layers. Each patch size is p2 × p2, which is then flattened into a one-dimensional vector v pos ∈ , C2 represents the number of channels of the input feature. The first-dimensional vector is then input to the self-attention layer, which can perform self-attention calculation on the input features according to preset parameters. Self-attention calculation refers to dynamically calculating the attention weight of each position to other positions to generate a feature representation that integrates global information. The above self-attention calculation process can be implemented through the Q (query), K (key), and V (value) matrices. The Q, K, and V matrices are trainable parameters, and their initial values ​​can be flexibly set according to the actual application scenario, but it is necessary to ensure that Q, K, V∈ , where np represents the number of patches.

[0035] For example, the output SA of the self-attention layer can be calculated by the following formula: (1) Among them, d kis the dimension of the key matrix. The features output by the Transformer network can be input into the CNN network, specifically, can be input into the corresponding CNN network layer in the CNN network. The correspondence between the Transformer network and the CNN network layer can be pre-set, and the patch parameter in the Transformer network is set based on the field of view of the CNN network layer. In the present application, the receptive field of the CNN network layer is defined by the following formula: (2) wherein the receptive field scale of the l-th layer is determined by the step size of the l-th layer , the size of the convolution kernel , and the receptive field scale of the l-th layer . For a multi-path network with the same step length, the overall receptive field is determined by the path with the largest receptive field. For example, for a convolution layer with a 2*2 convolution kernel and a step size of 2, the receptive field is 2*2+2=6. The patch parameter of the Transformer network can specifically be smaller than the receptive field of the CNN network layer but larger than half of the receptive field of the CNN network layer, i.e., RC i / 2 < patch size < RC i . For example, the receptive field of the first residual block is 6, and the patch size of the Transformer network corresponding to the first residual block can be set to 4.

[0036] As a possible implementation, the CNN network can include 5 CNN network layers, which respectively include an encoding convolution block, a decoding convolution block, a first encoding residual block, a first decoding residual block, a second encoding residual block, a second decoding residual block, a third encoding residual block, a third decoding residual block, a fourth encoding residual block, and a fourth decoding residual block. There can be four independent Transformer networks, which respectively correspond to the first CNN network layer composed of the first encoding residual block and the first decoding residual block, the second CNN network layer composed of the second encoding residual block and the second decoding residual block, the third CNN network layer composed of the third encoding residual block and the third decoding residual block, and the fourth CNN network layer composed of the fourth encoding residual block and the fourth decoding residual block. For ease of description, the Transformer networks are referred to as the first Transformer network, the second Transformer network, the third Transformer network, and the fourth Transformer network hereinafter.

[0037] As a possible implementation, the CNN network can include 5 CNN network layers, which respectively include an encoding convolution block, a decoding convolution block, a first encoding residual block, a first decoding residual block, a second encoding residual block, a second decoding residual block, a third encoding residual block, a third decoding residual block, a fourth encoding residual block, and a fourth decoding residual block. There can be four independent Transformer networks, which respectively correspond to the first CNN network layer composed of the first encoding residual block and the first decoding residual block, the second CNN network layer composed of the second encoding residual block and the second decoding residual block, the third CNN network layer composed of the third encoding residual block and the third decoding residual block, and the fourth CNN network layer composed of the fourth encoding residual block and the fourth decoding residual block. For ease of description, the Transformer networks are referred to as the first Transformer network, the second Transformer network, the third Transformer network, and the fourth Transformer network hereinafter.

[0038] ​​The above encoding convolution block and decoding convolution block can both contain two convolution layers with a convolution kernel size of 3×3. Each residual block can be constructed through a double 3x3 convolution layer and a jump connection. The first block of each stage (i.e., the above layer) is downsampled, the feature map size is halved at each stage, and the number of channels is doubled. Each maximum pooling operation doubles the receptive field of the input. For example, if there is a pixel in the output feature map of the maximum pooling layer at the bottom of the fourth stage, the receptive field scale of the pixel in the feature map before the pooling operation is 2, that is, . Calculated from the bottom up, the receptive field scales of each layer in the fourth stage are 、 In the embodiment of the present invention, IHT extracts features starting from the second stage, so there is no need to configure the image block size in the convolution block of the first stage. However, the long-distance dependent features extracted in the second stage need to "match" the features extracted by CNN in the first stage. Therefore, the features input by IHT need to match the feature size of the corresponding residual block. Each maximum pooling of the residual block will halve the receptive field, and this effect is cumulative. Therefore, the image block size of the second and third stages should be 4 or larger ( / ), while the image block size for the fourth and fifth stages should be 8 or larger ( / , / ). In this paper, PI-IHT achieves the best segmentation accuracy using a block size of 4−4−8−8.

[0039] like Figure 3 As shown, the image input to the CNN network passes through the encoding convolution block ( Figure 3 After convolution on the left side (CB), the first convolution feature F is obtained. e 1 Input to the first encoding residual block RB, the feature block output by the self-attention module Features extracted by CNN encoder cascaded and input to the RB module of the second stage. In particular, the feature block From the first stage characteristics After downsampling, the number of channels is 64. The generation method is the same.

[0040] The first global feature extracted by the first Transformer network is also input into the first encoding residual block. The first encoding residual block can perform local feature extraction and maxpooling on the first convolutional feature under the guidance of the first global feature, to obtain a first local feature input into a second encoding residual block. For example, the residual block can concatenate the first global feature and the first convolutional feature, and extract the local feature of the concatenated feature to obtain the first local feature. The second global feature extracted by the second Transformer network is also input into the second encoding residual block. In this way, the fourth local feature output by the fourth encoding residual block is input into an independent REsNet block, and is output by the feature extraction and maxpooling of the REsNet block to a fourth decoding residual block. The fourth decoding residual block can perform local feature extraction based on the feature map and the fourth local feature output by the fourth encoding residual block, and restore the resolution of the feature map by bilinear interpolation to expand the size of the feature map by one time, to obtain a fourth decoding local feature input into a third decoding residual block. The third decoding residual block can restore the resolution of the feature map based on the fourth decoding local feature and the third local feature output by the third encoding residual block by bilinear interpolation, to obtain a third decoding local feature. In this way, the second decoding local feature and the first decoding local feature are obtained.

[0041] To effectively capture the relevant context information of the multi-organ region, a prediction head is arranged at each stage of the decoder. The prediction head uses a 1x1 convolution kernel and performs non-linear mapping through a ReLU activation function. The output results of all stages are upsampled to match the resolution of the input image, and then added to obtain the target feature, and the Softmax activation function is used to generate the final prediction result. All convolution layers use ReLU as the activation function, and the features output by each decoder and the independent ResNet block can be added after adjusting the channel number and size by 1x1 convolution to obtain the output result.

[0042] Then, the difference between the output result and the segmentation label of each training image can be calculated. The difference can be calculated based on a cross entropy (CE) function. Then, the parameters in the image segmentation model can be adjusted by a feasible parameter adjustment method such as gradient descent method or gradient ascent method, until the difference converges, and the target image segmentation model is determined.

[0043] In a possible embodiment, in order to enhance the multi-scale modeling capability of the independent structure in the hierarchical ViT path, a pyramid input strategy can be used to input data to multiple independent Transformers, that is, a multi-scale input image is provided to the Transformer pyramid parsing module to make up for the deficiency of the Transformer pyramid parsing module in multi-scale information. That is, the image segmentation model can further include a pyramid input network, the pyramid input network including multiple convolutional layers; correspondingly, the method further includes: S21, down-sampling each training image by using the pyramid input network to obtain multiple down-sampled images of each training image; S22, inputting a down-sampled image of a corresponding size into each Transformer network based on the feature map size of the CNN network layer corresponding to each Transformer network.

[0044] Specifically, before the training image is input into the multiple independent Transformer networks, the training image can be input into an input network composed of multiple convolutional layers. The input network can be composed of multiple convolutional layers with a kernel size of 2*2, a stride of 2, and an output channel of 64, for example, the input network can include four convolutional layers. The feature map size of the pyramid input network is consistent with the size of each CNN network layer in the CNN network, that is, when the CNN network layer includes four layers, the convolutional layer in the pyramid input network also includes four layers, and the corresponding convolutional layer has the same feature map size as the CNN network layer.

[0045] Based on the above example, let a two-dimensional input image be X t ∈ , where t is the time sequence index of the image, the resolution is HxW, and the number of channels is C1=1. At the s=2 to s=5 stage of the network (corresponding to the 2-5 layer CNN network layer), the spatial size of the pyramid input image is H / 2 s−1 xW / 2 s−1 , and the corresponding number of channels is C 2−5 =64-128-256-512. Taking the second stage of the image segmentation model as an example, the pyramid input image ∈ is obtained by applying a convolution operation to the original input image , the kernel size is 2x2, the stride is 2, and the number of output channels is C2=64. The image size obtained after such processing is consistent with the dimension of the feature block extracted by the CNN at the corresponding stage, thereby ensuring the dimension alignment in the feature fusion process. As can be seen, the input image is down-sampled into a pyramid structure step by step from the second stage.

[0046] By the technical solution, each stage of the Transformer encoder independently processes the input of a specific scale, and breaks the sequence information flow in the traditional pyramid structure through a multi-scale input module. This design effectively preserves low-level features, i.e., pixel-level details (such as edge bumps), and avoids the smoothing effect caused by the down-sampling process of the Transformer. At the same time, the structure can also maintain the feature expression of the high-level stage, such as pure long-distance dependence (such as organ topological information), and avoid early feature pollution.

[0047] In addition, by spatially aligning the patch size of the Transformer and the receptive field of the CNN at each stage, the spatial-semantic consistency between the two encoders is strengthened. Global features are only injected into the CNN stage consistent with their semantic level, so there is no need for a complex fusion mechanism, and the overall GPU memory usage is reduced by about 21% compared to TransUNet.

[0048] As described above, the Transformer network usually focuses on the correlation between each position in the feature map to extract global features, which leads to less attention of the Transformer network itself to local features. In order to enhance the perception of local features by the Transformer network and improve the integrity of global features, in one possible embodiment, a Stem network can be added before the Transformer encoder, which is set for each independent Transformer network, that is, in the case of 4 Transformer networks, the Stem network also has 4.

[0049] Correspondingly, the above method can also include: inputting a plurality of the down-sampled images into the corresponding Stem network of the corresponding Transformer network, and extracting reference local features of the down-sampled images through the Stem network; inputting the reference local features into the corresponding Transformer network, so that the Transformer network extracts a plurality of global features of the training image based on the reference local features and the down-sampled images.

[0050] The parameters of the Stem network corresponding to the Transformer network can be determined according to the patch parameters of the Transformer, that is, the receptive field of the Stem network needs to be the same as the receptive field of the Transformer network. For example, the convolution kernel of the Stem network can be greater than half of the patch size of the corresponding Transformer network, but less than the patch size of the corresponding Transformer network, that is, the convolution kernel coresize of the Stem network satisfies: patch size / 2 < coresize < patch size.

[0051] Through the above technical solutions, a convolution Stem extension module (ECS, Extension of Convolutional Stems) is integrated in the Transformer encoder, that is, an additional convolution Stem is added before each IHT module to extract feature maps F sub , and feature fusion is performed in the second to fifth stages of the Transformer encoder. These features extracted by the Stem are injected into the IHT module, effectively alleviating the problem of over-smoothing of global representation.

[0052] In the medical image segmentation scene, there is a cross-scale similarity of organ morphology in medical images, that is, although the size of the organ in different slices is different (different scales), they belong to the same anatomical structure, so there is similarity in shape, texture and other features, so that the same organ can be tracked in consecutive slices. For example, abdominal CT scans are usually collected in the form of a transverse sequence. When observing the slices from top to bottom, the morphological continuity of the organ appears as gradual expansion or contraction across slices. This cross-scale similarity makes the same organ appear in different scales in consecutive slices.

[0053] In order to identify the cross-scale similarity of the organ, an adaptive stripe patch strategy can be set. Here, the stripe refers to a strip attention mechanism (Strip Attention), which is a kind of attention module designed for long-distance, directional dependency relationship, mainly used to enhance the model's perception ability of strip-shaped structures (such as roads, text lines, edges, etc.) in images, videos or sequence data. The logic of the strip attention mechanism is to truncate the feature map through the strip-shaped patch. The strips in the strip attention usually appear in pairs.

[0054] The above method can further include that the Transformer network inputs the global features into the corresponding strip attention module; The stripe attention module performs stripe attention calculation based on global features of the training image, a previous frame image and a next frame image of the training image in a time sequence, to obtain a stripe feature map; The stripe feature map is input into a corresponding CNN network layer of a corresponding Transformer network As a possible implementation, a cross-scale stripe patch attention (CSPA) module of PTC can be arranged at an output end of the Transformer to perform an adaptive stripe patch strategy and capture hierarchical correlation representation of content between adjacent slices.

[0055] By setting a smaller stripe width s w W and height s h H, it is ensured that each stripe patch only covers a local homogeneous region. Horizontal and vertical stripe attention is calculated in parallel. In order to effectively extract such features and reduce the computational overhead, the channels of the feature block can be divided into two groups, which are used to calculate the horizontal and vertical stripe attention respectively.

[0056] The number of stripes in each channel group is The size of the stripe patch is fixed as HorW=1. Therefore, the computational complexity is , which is equivalent to the traditional stripe attention method.

[0057] The stripe size of the CSPA module can be set according to the patch size of the corresponding Transformer network. For example, if the patch size of the Transformer network is 4, the stripe size can be set as (4, 1) and (1, 4). Based on the above example, in the case of four independent Transformer networks with patch sizes of [4-4-8-8], the stripe patch size in the CSPA module corresponding to each Transformer network can be set as [2 × (( sw = 4, sh = 1) − (1, 4)) − 2× (( sw = 8, sh = 1) − (1, 8)).

[0058] In single-head attention, attention calculation can be performed based on the previous and next frame images of the training image in the time sequence and the global features of the training image. Specifically, the Transformer network extracts feature blocks (as Key) from the three consecutive images (as Query) with The CSPA between (as Value) is calculated as shown in the following formula: (3) in, is the CSPA calculation result, represents the i-th horizontal strip of the Query feature, and The input feature blocks are and The Key and Value obtained after applying patch embedding.

[0059] Finally, the output of CSPA is composed of the concatenation of horizontal and vertical attention results, and the output is adjusted to the original feature map through the reshape function. The same size is then input into the corresponding CNN network layer to generate the final feature map .

[0060] like Figure 4 As shown, Figure 4 Another flow chart of a medical image segmentation method provided in an embodiment of the present invention is a flow chart of applying a target image segmentation model to segment a target medical image. The target image segmentation model may include a pyramid input network, a CNN network, multiple independent Transformer networks (IHTs), an ECS module, and a CSPA module.

[0061] ImageX t As the input to the CNN network and multiple independent Transformer networks IHT, the CNN encoder contains a convolutional block and four residual blocks. The feature extraction process includes the following steps: First layer: image X t After inputting into the CNN network, the CNN convolution block extracts X t The local feature C1(X t ), and reduce the feature to half its size through maximum pooling to obtain feature C1'(X t ).

[0062] And we can use the convolution layer to reduce the scale of A by half with a 2*2 convolution step of 2 to obtain the feature T1(X t ), this convolution layer is a separately set convolution layer, and together with the convolution layer with a convolution kernel of 2*2 and a step size of 2 mentioned later, it constitutes a pyramid input network for pre-processing the input of the Transformer network.

[0063] The second layer, the patch of 4 of the Transformer network extracts the global feature in T1(X t ) to obtain T1'(X t ) (corresponding to F t2 in Figure 4 ), and then spliced with C1'(X t ) to obtain T1'C1'(X t ), and the first residual block in the CNN extracts the local feature in T1'C1'(X t ) to obtain C2(X t ), and then C2(X t ) is reduced by half scale by using maximum pooling to obtain the feature C2'(X t ), and T1(X t ) is reduced by half scale by using 2*2 convolution with a step of 2 to obtain the feature T2(X t ).

[0064] The third layer: the patch of 4 of the Transformer network extracts the global feature in T2(X t ) to obtain T2'(X t ) (corresponding to F t3 in Figure 4 ), and then spliced with C2'(X t ) to obtain T2'C2'(X t ), and the second residual block in the CNN extracts the local feature in T2'C2'(X t ) to obtain the feature C3(X t ), and then C3(X t ) is reduced by half scale by using maximum pooling to obtain the feature C3'(X t ), and T2(X t ) is reduced by half scale by using 2*2 convolution with a step of 2 to obtain the feature T3(X t ).

[0065] The fourth layer, the patch of 8 of the Transformer network extracts the global feature in T3(X t ) to obtain T3'(X t ) (corresponding to F t4 in Figure 4 ), and then spliced with C3'(X t ) to obtain T3'C3'(X t ), and the third residual block in the CNN network extracts the local feature in T3'C3'(X t ) to obtain the feature C4(X t ), and then C4(X t ) is reduced by half scale by using maximum pooling to obtain the feature C4'(X t ), and T3(X​​​t ) Use 2*2 convolution step 2 to reduce the scale by half and get feature T4(X t ).

[0066] The fifth layer, the Transformer network with patch 8 extracts the global features in T4(A) to obtain T4'(X t ) (corresponding to Figure 4 F in t5 ), then with C4'(X t ) is spliced ​​into T4'C4'(X t ), the fourth residual block in the CNN network extracts T4'C4'(X t ) in the local features to obtain C5(X t ), then C5(X t ) Use upsampling to expand half the scale to C5'(X t ), for C5(X t ) Use 1*1 convolution and upsample to the input image X t Same size, get C5''(X t ).

[0067] The sixth layer uses jump connections to splice C4(X t ) and C5'(A) to obtain C4C5'(X t ), use the residual block to extract C4C5'(X t ) in the local features to obtain C6(X t ), then C6(X t ) Use upsampling to expand half the scale to C6'(X t ), for C6(X t ) Use 1*1 convolution and upsample to the input image X t The same size gets C6''(X t ).

[0068] The seventh layer uses jump connections to splice C3(X t ) and C6'(X t ) to obtain C3C6'(X t ), use the residual block to extract C3C6'(X t ) in the local features to obtain C7(X t ), then C7(X t ) Use upsampling to expand half the scale to C7'(X t ), for C7(X t ) Use 1*1 convolution and upsample to the input image X t The same size gets C7''(X t ).

[0069] The eighth layer uses jump connections to splice C2(Xt ) and C7'(X t ) to obtain C2C7'(X t ), and the local features in C2C7'(X t ) are extracted by a residual block to obtain C8(X t ), and then C8(X t ) is expanded by half scale by upsampling to obtain C8'(X t ), and C8'(X t ) is input into a 1*1 convolution and upsampling to the input image X t to obtain C8''(X t ).

[0070] The ninth layer, C1(X t ) and C8'(X t ) are spliced by a skip connection to obtain C1C8'(X t ), and the local features in C1C8'(X t ) are extracted by a CNN convolution block to obtain C9(X t ), and then C9(X t ) is output by a 1*1 convolution to obtain C9''(X t ).

[0071] Output, the sum of C5''(X t )+C6''(X t )+C7''(X t )+C8''(X t )+C9''(X t ) is obtained by an activation function softmax to obtain the final output, and each layer above refers to different stages in the network.

[0072] The above features T1(X t ), T2(X t ), T3(X t ) and T4(X t ) can be obtained by pyramid input network downsampling, and the obtained down-sampling features can be input into ECS1. The ECS1 is a Stem structure, and the convolution kernel size of each layer in the ECS1 is set according to the patch size of the corresponding Transformer network. The ECS1 can further extract features from the down-sampling features to generate reference local features and input them into the corresponding Transformer network together with the down-sampling features.

[0073] The four Transformer networks in this embodiment have patch sizes of 4-4-8-8 respectively. Specifically, each convolutional layer in the pyramid input network can input down-sampled features of different scales to different Transformer networks. For example, the down-sampled features obtained by the first and second convolutional layers can be input to two Transformer networks with a patch size of 4, and the down-sampled features obtained by the third and fourth convolutional layers can be input to two Transformer networks with a patch size of 8.

[0074] Each Transformer network performs global feature extraction on the input features to obtain global features, which are input to the corresponding CNN encoder in the CNN network. Meanwhile, the target global features output by each Transformer network are input to the ECS2 module. The convolution kernel size of the ECS2 module is also set according to the patch size of the Transformer, which can be the same as that of the ECS1. The ECS2 can down-sample the target global features output by the Transformer network and input the down-sampled results to the CSPA module. The previous and next frames of the target medical image can also be input to the CSPA module.

[0075] The CSPA module corresponds to the Transformer network. Specifically, the strip patch size in the CSPA module is determined according to the patch size of the Transformer network. The CSPA module can perform strip self-attention calculation on the down-sampled results of the target global features of the target image input thereto. The calculation process can refer to the previous and next frames. The strip self-attention calculation specifically includes horizontal attention calculation and vertical attention calculation. The output results of the CSPA are spliced from the horizontal and vertical attention results and reshaped to the same size as the target global features, and then input to the corresponding CNN encoder in the corresponding Transformer network layer.

[0076] The CNN encoder can splice the features input to the corresponding Transformer network layer and the CSPA module and the local features extracted by the convolutional layer in the CNN network from the medical image, and perform feature extraction on the spliced features. The features extracted by each CNN encoder are input into the corresponding CNN decoder, and the result output by the last CNN encoder is input into a separate ResNet module. The outputs of each CNN decoder and the separate ResNet module are all subjected to a 1*1 convolutional layer, and the convolutional results are added to obtain the target feature of the target image, and based on the target feature, the target segmentation result of the target image can be determined.

[0077] By using the embodiment of the application, the global semantic prior extracted by the Transformer is used to guide the local feature extraction of the CNN encoder, so that the overall semantic expression capability is improved. In addition, an isolated hierarchical Transformer (IHT) structure with a pyramid input is proposed. The structure realizes the abstract independence between stages through strict multi-scale input processing, and avoids cross-stage interference. The Transformer sub-module of each stage uses a single-head self-attention mechanism to extract the global features of the corresponding pyramid input. Furthermore, according to the alignment relationship between the patch size and the receptive field of the two encoders in the same stage, feature alignment and efficient fusion without complex fusion modules are realized.

[0078] The PTC architecture provided by the application realizes the balanced fusion of local (CNN) and global (Transformer) feature semantics while optimizing the computing efficiency (compared with TransUNet, the GPU memory occupancy is reduced by about 21%). The decoder adopts a direct feature summation (DFS) method to directly weight the outputs of multiple stages to generate a unified prediction result.

[0079] In addition, the PTC architecture has good scalability and supports 2D and 3D dual modes. For example, by adding an extension of convolutional stems (ECS) module before the Transformer encoder, the stability of feature extraction in 2D object segmentation can be enhanced, and the boundary segmentation accuracy can be improved. For 3D scenes, 3D convolution kernels can be directly used for feature encoding, or a cross-scale stripe patch attention (CSPA) mechanism can be used to improve the processing efficiency of 3D sequence images while maintaining the accuracy.

[0080] Based on the same inventive concept, the embodiment of the present invention also provides a medical image segmentation device, such as Figure 5 As shown, the apparatus 500 may include: An acquisition module 501 is configured to acquire a target medical image; An input module 502 is configured to input the target medical image into a pre-trained target image segmentation model and obtain a target segmentation result output by the target image segmentation model for the target medical image; The training module 503 is configured to pre-train the target image segmentation model through the following steps: Acquire a training data set, wherein the training data set includes a plurality of training images, each of the training images corresponds to a segmentation label, and the segmentation label is used to identify a location of a target area in the training image; Inputting the training data set into an initial image segmentation model, wherein the image segmentation model includes a CNN network and multiple independent Transformer networks, wherein the CNN network includes multiple CNN network layers; the Transformer networks correspond one-to-one to the CNN network layers, and the feature map scale of the Transformer network is the same as the feature map scale of the corresponding CNN network; For each of the training images, extracting multiple global features of the training image through each of the Transformer networks, and inputting the multiple global features into the CNN network layer corresponding to the Transformer network; Using each of the CNN network layers to extract features based on the training image and the global features output by the Transformer network corresponding to the CNN network layer, to obtain multiple local features; Superimposing a plurality of the local features to obtain a target feature of the training image, and outputting a predicted segmentation result of the training image based on the target feature; Training the image segmentation model based on the difference between the predicted segmentation result of each training image and the segmentation label of each training image; When the difference converges, the current image segmentation model is determined to be the target image segmentation model.

[0081] In a possible embodiment, the image segmentation model further includes a pyramid input network, wherein the pyramid input network includes multiple convolutional layers; the training module is further configured to downsample each of the training images using the pyramid input network to obtain multiple downsampled images of each of the training images; inputting a down-sampled image of a corresponding size into each of the Transformer networks based on a feature map scale of a corresponding CNN network layer of each of the Transformer networks; The CNN network comprises a CNN encoder and a CNN decoder. The CNN encoder comprises an encoding convolution block and a plurality of encoding residual blocks connected in series. The CNN decoder comprises a decoding convolution block and a plurality of decoding residual blocks. One encoding residual block and its corresponding decoding residual block form one CNN network layer. The output ends of the decoding convolution block and each decoding residual block are connected to a prediction head. The CNN network layer uses the training image and the global feature output by the corresponding Transformer network to perform feature extraction, obtaining a plurality of local features, including: The training image is input into the CNN encoder, and the first local feature of the training image is extracted by the encoding convolution block, and the first local feature is input into the encoding residual block connected to the encoding convolution block. The first local feature and the first global feature output by the Transformer network corresponding to the encoding residual block connected to the encoding convolution block are spliced to obtain a first fusion feature. The first fusion feature is locally feature-extracted by the encoding residual block connected to the encoding convolution block to obtain a second local feature. The second local feature is input into a next-level encoding residual block, so that the next-level encoding residual block splices the second local feature and the global feature input by the Transformer network corresponding to the next-level encoding residual block, and locally feature-extracts the spliced feature to obtain a new second local feature, returning to the step of inputting the second local feature into the next-level encoding residual block until all encoding residual blocks perform local feature extraction. The first local feature and each second local feature obtained by feature extraction of each encoding residual block are input into the encoding convolution block, the decoding convolution block corresponding to each encoding residual block, and each decoding residual block to obtain a plurality of local features of the training image output by the CNN decoder. The plurality of local features are superimposed to obtain the target feature of the training image, including: The plurality of local features of the training image output by the CNN decoder are input into the corresponding prediction head, and the features output by each prediction head are superimposed to obtain the target feature of the training image. The patch parameter of the Transformer network satisfies the following condition: RCi / 2 < Patch size < RCi, wherein RCi is a receptive field of a CNN network layer i corresponding to the Transformer network, the receptive field is related to a convolution kernel size of the CNN network layer, and the patch size is the patch parameter of the Transformer network. The image segmentation model further comprises a Stem network corresponding to the Transformer network; and the training module is further configured to input the plurality of down-sampled images into the Stem network corresponding to the Transformer network, and extract reference local features of the down-sampled images through the Stem network. The reference local features are input into the Transformer network, so that the Transformer network extracts a plurality of global features of the training image based on the reference local features and the down-sampled images. The image segmentation model further comprises a strip attention module corresponding to the Transformer network and receiving an input of the corresponding Transformer network; and the training module is further configured to input the global features into the strip attention module corresponding to the Transformer network. The strip attention module performs strip attention calculation based on the global features of the training image, a previous frame image and a next frame image of the training image in a time sequence, to obtain a strip feature map. The strip feature map is input into the CNN network layer corresponding to the Transformer network.

[0082] In the present application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0083] The exemplary embodiments of the present application also provide an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication. The memory stores a computer program capable of being executed by the at least one processor, and the computer program is used to make the electronic device execute the method according to the embodiments of the present application when executed by the at least one processor.

[0084] The exemplary embodiments of the present application also provide a non-transitory computer readable storage medium storing a computer program, wherein the computer program is used to make the computer execute the method according to the embodiments of the present application when executed by a processor of the computer.

[0085] An exemplary embodiment of the present invention further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor of a computer, the computer is configured to cause the computer to perform a method according to an embodiment of the present invention.

[0086] refer to Figure 6 , a block diagram of an electronic device 600 that can serve as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0087] like Figure 6 As shown, electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of electronic device 600. Computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0088] Multiple components within electronic device 600 are connected to I / O interface 605, including an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. Input unit 606 can be any type of device capable of inputting information into electronic device 600. Input unit 606 can receive input numeric or character information and generate key input signals related to user settings and / or function control of the electronic device. Output unit 607 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 608 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 609 allows electronic device 600 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0089] The computing unit 601 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above. For example, in some embodiments, any of the medical image segmentation methods described above can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to perform any of the medical image segmentation methods described above by other any suitable means, such as by means of firmware.

[0090] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be retrieved from a machine-readable medium or device, stored media, and / or making a transmission, such as over the Internet, over a wireless network, and / or over a wired network. The machine-readable medium or device can include any suitable medium or device that is non-transitory and includes a machine-readable storage medium and / or a machine-readable transmission medium. The machine-readable storage medium can include a tangible device that can retain, store, or maintain the program code for a period of time. The machine-readable transmission medium can include a machine-readable storage medium or a machine-readable signal medium that can include any suitable medium or device that can transmit or carry the program code for a period of time. The program code can cause a machine, such as a processor or controller, to implement the functions / acts specified in the flowcharts and / or block diagrams.

[0091] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store program code for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of a program of instructions in a searchable database, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0092] As used in this description, the terms "machine-readable medium," "computer-readable medium," and "computer-readable media" refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that can be used to provide machine instructions and / or data to a programmable processor.

[0093] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0094] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0095] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

Claims

1. A medical image segmentation method characterized by, The method comprises: obtaining a target medical image; inputting the target medical image into a pre-trained target image segmentation model, and obtaining a target segmentation result output by the target image segmentation model for the target medical image; wherein the target image segmentation model is pre-trained through the following steps: obtaining a training data set, the training data set comprising a plurality of training images, each training image corresponding to a segmentation label, the segmentation label being used to identify the position of a target region in the training image; inputting the training data set into an initial image segmentation model, the image segmentation model comprising a CNN network and a plurality of independent Transformer networks, the CNN network comprising a plurality of CNN network layers; the Transformer network corresponds to the CNN network layer one by one, and the feature map scale of the Transformer network is the same as that of the corresponding CNN network; for each training image, extracting a plurality of global features of the training image through each Transformer network, and inputting the plurality of global features into the CNN network layer corresponding to the Transformer network; using each CNN network layer to extract features based on the training image and the global features output by the corresponding Transformer network, to obtain a plurality of local features; stacking a plurality of the local features to obtain a target feature of the training image, and outputting a predicted segmentation result of the training image based on the target feature; training the image segmentation model based on the difference between the predicted segmentation result of each training image and the segmentation label of each training image; in the case where the difference converges, determining the current image segmentation model as the target image segmentation model.

2. The method of claim 1, wherein, The image segmentation model further comprises a pyramid input network, the pyramid input network comprising a plurality of convolution layers; the method further comprises: down-sampling each training image using the pyramid input network to obtain a plurality of down-sampled images of each training image; inputting a down-sampled image of corresponding size into each Transformer network based on the feature map scale of the CNN network layer corresponding to each Transformer network.

3. The method of claim 1, wherein, The CNN network comprises a CNN encoder and a CNN decoder, the CNN encoder comprising an encoding convolution block and a plurality of encoding residual blocks, the encoding convolution block and the plurality of encoding residual blocks being connected in series, the CNN decoder comprising a decoding convolution block and a plurality of decoding residual blocks, one encoding residual block and its corresponding decoding residual block forming one CNN network layer, and the output end of the decoding convolution block and each decoding residual block being connected to a prediction head; the use of each CNN network layer to extract features based on the training image and the global features output by the corresponding Transformer network to obtain a plurality of local features comprises: inputting the training image into the CNN encoder, extracting first local features of the training image by using the encoding convolutional block, and inputting the first local features into an encoding residual block connected to the encoding convolutional block; concatenating the first local features and first global features output by a Transformer network corresponding to the encoding residual block connected to the encoding convolutional block to obtain first fusion features; extracting local features of the first fusion features by using the encoding residual block connected to the encoding convolutional block to obtain second local features; inputting the second local features into a next-level encoding residual block, so that the next-level encoding residual block concatenates the second local features and global features input by a Transformer network corresponding to the next-level encoding residual block, and extracts local features of the concatenated features to obtain new second local features, and returning to the step of inputting the second local features into the next-level encoding residual block until all the encoding residual blocks extract local features; inputting the first local features and second local features extracted by each of the encoding residual blocks into an encoding convolutional block, a decoding convolutional block corresponding to each of the encoding residual blocks, and each of the decoding residual blocks to obtain a plurality of local features of the training image output by the CNN decoder; the step of superimposing the plurality of local features to obtain target features of the training image comprises: inputting the plurality of local features of the training image output by the CNN decoder into corresponding prediction heads, and superimposing features output by each of the prediction heads to obtain the target features of the training image.

4. The method of claim 1, wherein, The patch parameter of the Transformer network satisfies the following condition: RCi / 2 < Patch size < RCi, where RCi is a receptive field of a CNN network layer i corresponding to the Transformer network, the receptive field is related to a convolution kernel size of the CNN network layer, and patch size is the patch parameter of the Transformer network.

5. The method of claim 2, wherein, The image segmentation model further comprises a Stem network, and the Stem network corresponds to the Transformer network in a one-to-one manner; the method further comprises: inputting a plurality of the down-sampled images into a Stem network corresponding to a corresponding Transformer network, and extracting reference local features of the down-sampled images by using the Stem network; inputting the reference local features into the corresponding Transformer network, so that the Transformer network extracts a plurality of global features of the training image based on the reference local features and the down-sampled images.

6. The method of claim 1, wherein, The image segmentation model further comprises a strip attention module, the strip attention module corresponds to the Transformer network in a one-to-one manner, and receives input of the corresponding Transformer network; the method further comprises: the Transformer network inputs the global features into the corresponding strip attention module; The stripe attention module performs stripe attention calculation based on the global features of the training image, the previous frame image and the next frame image of the training image in the time series to obtain a stripe feature map; The strip feature map is input into the CNN network layer corresponding to the corresponding Transformer network.

7. A medical image segmentation apparatus characterized by comprising: The device comprises: An acquisition module, used for acquiring a target medical image; An input module, configured to input the target medical image into a pre-trained target image segmentation model, and obtain a target segmentation result output by the target image segmentation model for the target medical image; The training module is used to pre-train the target image segmentation model through the following steps: Acquire a training data set, wherein the training data set includes a plurality of training images, each of the training images corresponds to a segmentation label, and the segmentation label is used to identify a location of a target area in the training image; Inputting the training data set into an initial image segmentation model, wherein the image segmentation model includes a CNN network and multiple independent Transformer networks, wherein the CNN network includes multiple CNN network layers; the Transformer networks correspond one-to-one to the CNN network layers, and the feature map scale of the Transformer network is the same as the feature map scale of the corresponding CNN network; For each of the training images, extracting multiple global features of the training image through each of the Transformer networks, and inputting the multiple global features into the CNN network layer corresponding to the Transformer network; Using each of the CNN network layers to extract features based on the training image and the global features output by the Transformer network corresponding to the CNN network layer, to obtain multiple local features; Superimposing a plurality of the local features to obtain a target feature of the training image, and outputting a predicted segmentation result of the training image based on the target feature; Training the image segmentation model based on the difference between the predicted segmentation result of each training image and the segmentation label of each training image; When the difference converges, the current image segmentation model is determined to be the target image segmentation model.

8. The apparatus of claim 7, wherein, The image segmentation model further includes a pyramid input network, which includes multiple convolutional layers; the training module is further configured to downsample each of the training images using the pyramid input network to obtain multiple downsampled images of each of the training images; Inputting downsampled images of corresponding sizes into each of the Transformer networks based on the feature map scale of the CNN network layer corresponding to each of the Transformer networks; The CNN network comprises a CNN encoder and a CNN decoder, the CNN encoder comprises an encoding convolution block and a plurality of encoding residual blocks connected in series, the CNN decoder comprises a decoding convolution block and a plurality of decoding residual blocks, one encoding residual block and its corresponding decoding residual block constitute a CNN network layer, and the output ends of the decoding convolution block and each decoding residual block are connected to a prediction head; The CNN network layer based on the training image and the global feature output by the corresponding Transformer network of the CNN network layer is used for feature extraction to obtain a plurality of local features, including: The training image is input into the CNN encoder, the first local feature of the training image is extracted by the encoding convolution block, and the first local feature is input into the encoding residual block connected with the encoding convolution block; The first local feature and the first global feature output by the Transformer network corresponding to the encoding residual block connected with the encoding convolution block are spliced to obtain a first fusion feature; The first fusion feature is subjected to local feature extraction by the encoding residual block connected with the encoding convolution block to obtain a second local feature; The second local feature is input into a next-level encoding residual block, so that the next-level encoding residual block splices the second local feature and the global feature input by the Transformer network corresponding to the next-level encoding residual block, and performs local feature extraction on the spliced feature to obtain a new second local feature, and returns to the step of inputting the second local feature into the next-level encoding residual block until all encoding residual blocks perform local feature extraction; The first local feature and each second local feature obtained by feature extraction of each encoding residual block are input into the encoding convolution block, the decoding convolution block corresponding to each encoding residual block and each decoding residual block to obtain a plurality of local features of the training image output by the CNN decoder; The plurality of local features are superimposed to obtain the target feature of the training image, including: The plurality of local features of the training image output by the CNN decoder are input into the corresponding prediction head, and the features output by each prediction head are superimposed to obtain the target feature of the training image; The patch parameter of the Transformer network satisfies the following condition: RCi / 2 < Patch size < RCi, wherein RCi is the receptive field of the CNN network layer i corresponding to the Transformer network, the receptive field is related to the convolution kernel size of the CNN network layer, and patch size is the patch parameter of the Transformer network. The image segmentation model further comprises a Stem network corresponding to each Transformer network; the training module is further configured to input the plurality of down-sampled images into the Stem network corresponding to the Transformer network, and extract reference local features of the down-sampled images through the Stem network; input the reference local features into the corresponding Transformer network, so that the Transformer network extracts a plurality of global features of the training image based on the reference local features and the down-sampled images; The image segmentation model further comprises a strip attention module corresponding to each Transformer network, which receives the input of the corresponding Transformer network; the training module is further configured to input the global features into the corresponding strip attention module by the Transformer network; The strip attention module performs strip attention calculation based on the global features of the training image, the previous frame image and the next frame image of the training image in the time sequence, to obtain a strip feature map; input the strip feature map into the CNN network layer corresponding to the corresponding Transformer network.

9. An electronic device comprising: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the method of any one of claims 1-6.

10. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to make the computer perform the method of any one of claims 1-6.