Image segmentation method, device, computer device, and storage medium

Through multi-angle feature extraction and adaptive learning to generate fusion feature maps, the problem of single features of existing medical image segmentation methods is solved, and the accuracy and universality of image segmentation are improved.

CN118887405BActive Publication Date: 2025-06-24SHANGHAI LIANYING ZHIYUAN MEDICAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411172208.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-06-24
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

The image features obtained by existing medical image segmentation methods are single, resulting in poor image segmentation effect and lack of spatial information and depth information.

Method used

By extracting multi-angle feature of the image to be segmented, the first target feature map and the second target feature map are obtained, and adaptive learning is combined with the space adapter and the depth adapter to generate a fusion feature map for segmentation processing.

Benefits of technology

It improves the accuracy and universality of image segmentation, and can more effectively extract the regions of interest in medical images, which are suitable for two-dimensional and three-dimensional image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887405B_ABST
    Figure CN118887405B_ABST
Patent Text Reader

Abstract

The present application relates to an image segmentation method, apparatus, computer device, and storage medium. Feature extraction is performed on an image to be segmented to obtain a target feature map of the image to be segmented, and based on the target feature map, segmentation processing is performed on the image to be segmented to obtain a region-of-interest image in the image to be segmented; the target feature map includes a first target feature map and a second target feature map. Since the image features obtained by existing image segmentation methods are single, resulting in poor image segmentation effects, in the embodiments of the present application, the first target feature map and the second target feature map are obtained by performing feature extraction on the image to be segmented from different angles, and the first target feature map and the second target feature map are used to perform segmentation processing on the image to be segmented, which can improve the segmentation accuracy of the image to be segmented. Moreover, the method proposed in the present application can not only segment two-dimensional images, but also be applicable to the segmentation of three-dimensional images, improving the universality of image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical technologies, and particularly to an image segmentation method, apparatus, computer device, and storage medium. Background Art

[0002] Medical image segmentation can extract key information from specific tissue images and is a key component in clinical practice. The segmented medical images are provided to doctors for different tasks such as quantitative analysis of tissue volume, diagnosis, localization of pathologically changed tissues, depiction of anatomical structures, and treatment planning. Therefore, how to effectively segment medical images and improve the segmentation effect of medical images is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide an image segmentation method, apparatus, computer device, and storage medium that can improve the image segmentation effect.

[0004] In a first aspect, the present application provides an image segmentation method, including:

[0005] Performing feature extraction on the image to be segmented to obtain a target feature map of the image to be segmented; the target feature map includes a first target feature map and a second target feature map; the first target feature map and the second target feature map are feature maps obtained by performing feature extraction on the image to be segmented from different angles;

[0006] Performing segmentation processing on the image to be segmented based on the target feature map to obtain a region of interest image in the image to be segmented.

[0007] In one embodiment, the performing feature extraction on the image to be segmented to obtain a target feature map of the image to be segmented includes:

[0008] Performing embedding processing on the image to be segmented to obtain an image embedding vector of the image to be segmented;

[0009] Performing feature extraction on the image embedding vector to obtain the first target feature map of the image to be segmented;

[0010] Transposing the image embedding vector to obtain a transposed image embedding vector;

[0011] Performing feature extraction on the transposed image embedding vector to obtain the second target feature map of the image to be segmented.

[0012] In one embodiment, the image to be segmented includes a three-dimensional image to be segmented, and the first target feature map includes a target spatial feature map; the extracting the first target feature map of the image to be segmented from the image embedding vector includes:

[0013] Performing spatial feature extraction on the image embedding vector to obtain an initial spatial feature map of the image to be segmented;

[0014] Inputting the initial spatial feature map into a spatial adaptor for adaptive learning to obtain the target spatial feature map.

[0015] In one embodiment, the image to be segmented includes a three-dimensional image to be segmented, and the second target feature map includes a target depth feature map; the extracting the second target feature map of the image to be segmented from the transposed image embedding vector includes:

[0016] Performing depth feature extraction on the transposed image embedding vector to obtain an initial depth feature map of the image to be segmented;

[0017] Inputting the initial depth feature map into a depth adaptor for adaptive learning to obtain the target depth feature map.

[0018] In one embodiment, the segmenting the image to be segmented based on the target feature map to obtain a region of interest image in the image to be segmented includes:

[0019] Transposing the second target feature map to obtain a transposed second target feature map;

[0020] Fusing the first target feature map and the transposed second target feature map to obtain a fused feature map;

[0021] Segmenting the image to be segmented based on the fused feature map to obtain the region of interest image.

[0022] In one embodiment, the segmenting the image to be segmented based on the fused feature map to obtain the region of interest image includes:

[0023] Obtaining interactive prompt information;

[0024] Segmenting the image to be segmented based on the interactive prompt information and the fused feature map to obtain the region of interest image.

[0025] In one embodiment, the segmenting the image to be segmented based on the interactive prompt information and the fused feature map to obtain the region of interest image includes:

[0026] Associate the image embedding vector with the interactive prompt information to obtain associated information;

[0027] Based on the associated information and the fused feature map, perform segmentation processing on the image to be segmented to obtain the region of interest image.

[0028] In a second aspect, the present application also provides an image segmentation device, including:

[0029] An extraction module, configured to perform feature extraction on the image to be segmented to obtain a target feature map of the image to be segmented; the target feature map includes a first target feature map and a second target feature map; the first target feature map and the second target feature map are feature maps obtained by performing feature extraction on the image to be segmented from different angles;

[0030] A segmentation module, configured to perform segmentation processing on the image to be segmented based on the target feature map to obtain a region of interest image in the image to be segmented.

[0031] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0032] Perform feature extraction on the image to be segmented to obtain a target feature map of the image to be segmented; the target feature map includes a first target feature map and a second target feature map; the first target feature map and the second target feature map are feature maps obtained by performing feature extraction on the image to be segmented from different angles;

[0033] Based on the target feature map, perform segmentation processing on the image to be segmented to obtain a region of interest image in the image to be segmented.

[0034] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0035] Perform feature extraction on the image to be segmented to obtain a target feature map of the image to be segmented; the target feature map includes a first target feature map and a second target feature map; the first target feature map and the second target feature map are feature maps obtained by performing feature extraction on the image to be segmented from different angles;

[0036] Based on the target feature map, perform segmentation processing on the image to be segmented to obtain a region of interest image in the image to be segmented.

[0037] In a fifth aspect, the present application further provides a computer program product, including a computer program which, when executed by a processor, implements the following steps:

[0038] Extract features from the image to be segmented to obtain the target feature map of the image to be segmented; the target feature map includes a first target feature map and a second target feature map; the first target feature map and the second target feature map are feature maps obtained by extracting features from the image to be segmented from different angles;

[0039] Perform segmentation processing on the image to be segmented based on the target feature map to obtain the region-of-interest image in the image to be segmented.

[0040] The above image segmentation method, device, computer device and storage medium extract features from the image to be segmented to obtain the target feature map of the image to be segmented, and perform segmentation processing on the image to be segmented based on the target feature map to obtain the region-of-interest image in the image to be segmented; the target feature map includes a first target feature map and a second target feature map; the first target feature map and the second target feature map are feature maps obtained by extracting features from the image to be segmented from different angles. Since the image features obtained by existing image segmentation methods are single, resulting in poor image segmentation effects, in the embodiments of the present application, the first target feature map and the second target feature map are obtained by extracting features from the image to be segmented from different angles, and the first target feature map and the second target feature map are used to perform segmentation processing on the image to be segmented, which can improve the segmentation accuracy of the image to be segmented. Moreover, the method proposed in the present application can not only segment two-dimensional images, but also be applicable to the segmentation of three-dimensional images, improving the universality of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0042] Figure 1 It is an application environment diagram of the image segmentation method in an embodiment;

[0043] Figure 2 It is a flowchart of the image segmentation method in an embodiment;

[0044] Figure 3 It is a flowchart of the image segmentation method in another embodiment;

[0045] Figure 4Schematic flowchart of a method for determining a target spatial feature map in an embodiment;

[0046] Figure 5 Schematic flowchart of a method for determining a target spatial feature map and a target depth feature map in an embodiment;

[0047] Figure 6 Schematic diagram of a spatial adapter in an embodiment;

[0048] Figure 7 Schematic flowchart of a method for determining a target depth feature map in an embodiment;

[0049] Figure 8 Schematic flowchart of a method for determining an image of a region of interest in an embodiment;

[0050] Figure 9 Schematic flowchart of a method for determining an image of a region of interest in another embodiment;

[0051] Figure 10 Block diagram of the structure of an image segmentation device in an embodiment;

[0052] Figure 11 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0053] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] Medical image segmentation is a key component in clinical practice, which helps with accurate diagnosis, treatment planning, and disease monitoring. However, most current medical image segmentation methods mainly rely on customized models, which show limited generality in different tasks. For example, the Segment Anything Model (SAM) and the Segment Anything in Medical Images (MedSAM). The MedSAM model can perform one-click segmentation on any object from a photo or video image and can be zero-shot migrated to other tasks. However, the SAM and MedSAM models have poor segmentation effects on medical images, lacking spatial information and depth information, resulting in unsatisfactory segmentation results.

[0055] Therefore, the present application proposes an image segmentation method, device, computer device, and storage medium that can solve the above technical problems.

[0056] The image segmentation method provided by the embodiments of the present application can be applied to, for exampleFigure 1 in the application environment shown. The application environment includes a computer device, which can be a server, and its internal structure diagram can be as Figure 1 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data for image segmentation. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements an image segmentation method. The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0057] In an exemplary embodiment, as Figure 2 shown, an image segmentation method is provided. Taking the computer device in Figure 1 as an example, the following S201 to S202 are included. Among them:

[0058] S201, perform feature extraction on the image to be segmented to obtain the target feature map of the image to be segmented; the target feature map includes a first target feature map and a second target feature map; the first target feature map and the second target feature map are feature maps obtained by performing feature extraction on the image to be segmented from different angles.

[0059] Optionally, the image to be segmented can be a medical image, for example, a nuclear magnetic resonance image, a computed tomography image, etc.; it can also be an environmental image, for example, a sky image, a mountain image, etc.; it can also be an animal image, etc.

[0060] Optionally, the image to be segmented can be a two-dimensional image or a three-dimensional image.

[0061] In the embodiments of the present application, the medical image to be segmented can be input into an encoder, and the encoder is used to convert the image to be segmented into a fixed-length vector representation, that is, an image embedding vector. The purpose of the image embedding vector is to convert the high-dimensional image to be segmented into lower-dimensional data that is easier to process while retaining as much original image information as possible.

[0062] Further, feature extraction is performed on the image embedding vector to obtain a first target feature map; and the image embedding vector is transposed, and feature extraction is performed on the transposed image embedding vector to obtain a second target feature map different from the first target feature map.

[0063] In a possible implementation, feature extraction can also be directly performed on the image to be segmented to obtain a first target feature map, the image to be segmented is transposed to obtain a transposed image to be segmented, and feature extraction is performed on the transposed image to be segmented to obtain a second target feature map different from the first target feature map.

[0064] In another possible implementation, different feature extraction methods or different feature extraction networks can be used to perform feature extraction on the image to be segmented to obtain a first target feature map and a second target feature map. For example, principal component analysis, histogram of oriented gradients, feature point detection, descriptors; convolutional neural network, long short-term memory network, recurrent neural network, etc.

[0065] S202, perform segmentation processing on the image to be segmented based on the target feature map to obtain a region-of-interest image in the image to be segmented.

[0066] In the embodiment of the present application, the first target feature map and the second target feature map can be fused to obtain a fused feature map, and segmentation is performed on the medical image to be segmented based on the fused feature map to obtain a region-of-interest image in the medical image to be segmented.

[0067] In a possible implementation, the medical image to be segmented can be first segmented using the first target feature map to obtain a first candidate region-of-interest image in the medical image to be segmented, the medical image to be segmented is segmented using the second target feature map to obtain a second candidate region-of-interest image in the medical image to be segmented, and the first candidate region-of-interest image and the second candidate region-of-interest image are fused and processed to obtain a region-of-interest image.

[0068] In the above image segmentation method, feature extraction is performed on the image to be segmented to obtain the target feature map of the image to be segmented. Based on the target feature map, segmentation processing is performed on the image to be segmented to obtain the region-of-interest image in the image to be segmented. The target feature map includes a first target feature map and a second target feature map. The first target feature map and the second target feature map are feature maps obtained by performing feature extraction on the image to be segmented from different perspectives. Since the image features obtained by existing image segmentation methods are single, the image segmentation effect is poor. In the embodiments of the present application, the first target feature map and the second target feature map are obtained by performing feature extraction on the image to be segmented from different perspectives, and the first target feature map and the second target feature map are used to perform segmentation processing on the image to be segmented, which can improve the segmentation accuracy of the image to be segmented. Moreover, the method proposed in the present application can not only segment two-dimensional images, but also be applicable to the segmentation of three-dimensional images, improving the universality of image segmentation.

[0069] In an exemplary embodiment, performing feature extraction on the image to be segmented to obtain the target feature map of the image to be segmented includes: performing embedding processing on the image to be segmented to obtain the image embedding vector of the image to be segmented; performing feature extraction on the image embedding vector to obtain the first target feature map of the image to be segmented; transposing the image embedding vector to obtain the transposed image embedding vector; and performing feature extraction on the transposed image embedding vector to obtain the second target feature map of the image to be segmented.

[0070] In the embodiments of the present application, as Figure 3 shown, taking the image to be segmented as a three-dimensional image as an example, the image to be segmented can be input into an autoencoder to perform embedding processing on the image to be segmented to obtain the image embedding vector. " " represents an image embedding vector, where N is the number of embeddings, L is the embedding length, and D is the number of operations, that is, the depth.

[0071] The first target feature map extraction branch is mainly responsible for learning the spatial correlation of the three-dimensional image on the horizontal and vertical planes, that is, the relationship between pixels within the same depth slice. Specifically, for the 3D embedding vector with a depth of D, the embedding of each depth slice is input into the first target feature map extraction branch as an independent sequence, and the first target feature map extraction branch is applied within the range of to learn and abstract the spatial correlation.

[0072] The second target feature map extraction branch focuses on the correlation of the 3D image in the depth direction, that is, the relationship between different depth slices. Its goal is to capture the feature changes along the depth direction, which is crucial for understanding the 3D structure, such as the position and morphology of organs, tissues, and lesions. During specific operations, the image embedding vector is input into the second target feature map extraction branch. First, the image embedding vector is transposed to obtain the transposed image embedding vector " ". The second target feature map extraction branch is applied within the range of to learn and abstract the depth correlation.

[0073] In a possible implementation, taking the image to be segmented as a 2D image as an example, the image to be segmented can be input into an autoencoder for embedding processing to obtain an image embedding vector. The image embedding vector represented by " " is input into the first target feature map extraction branch and the second target feature map extraction branch. The first target feature map extraction branch extracts features from the image embedding vector to obtain the first target feature map. The first target feature map extraction branch first transposes the image embedding vector to obtain the transposed image embedding vector " ", and then extracts features from the transposed image embedding vector to obtain the second target feature map.

[0074] In the embodiment of the present application, the image to be segmented is subjected to embedding processing to obtain the image embedding vector of the image to be segmented; features are extracted from the image embedding vector to obtain the first target feature map of the image to be segmented; the image embedding vector is transposed to obtain the transposed image embedding vector; features are extracted from the transposed image embedding vector to obtain the second target feature map of the image to be segmented, so as to obtain the features of the image to be segmented from multiple angles, laying a foundation for subsequent image segmentation processing based on the first target feature map and the second target feature map.

[0075] Figure 4 FIG. is a schematic flowchart of a method for determining a target spatial feature map in an embodiment. As shown in Figure 4 , the embodiment of the present application relates to a possible implementation manner of obtaining the target feature map of the medical image to be segmented according to the image embedding vector and the attention model in the image segmentation model, including the following steps:

[0076] S401, perform spatial feature extraction on the image embedding vector to obtain the initial spatial feature map of the image to be segmented.

[0077] Unlike two-dimensional images, many medical images are three-dimensional, such as magnetic resonance images and computed tomography images. Since doctors usually need to identify the correlations between slices to make relevant decisions. The segmented region of interest images of these three-dimensional medical images are crucial in clinical use. Although the MedSAM model can be applied to each slice of the medical image to obtain the final segmentation, it does not consider the spatial correlation and / or depth correlation of the medical image.

[0078] In this embodiment, the image embedding vector can be input into the first attention module, and the interaction of the first attention module is applied to the image embedding vector, focusing on learning the spatial correlation within each slice to learn and abstract the spatial association information, that is, using the first attention module to extract spatial information and semantic information to obtain the initial spatial feature map. Optionally, the first attention module can be a multi-head attention module, a cross-attention module, a squeeze-and-excitation (SE) module, a convolutional block attention module (CBAM), an efficient channel attention (ECA) module, a non-local module, a global context (GC) module, a SimAM module, etc.

[0079] In a possible implementation, as Figure 5 shown, a first regularization layer can also be set before the first attention module. The image embedding vector is input into the first regularization layer to obtain the first feature map, and the first feature map is input into the first attention module to extract spatial information and semantic information to obtain the initial spatial feature map. Using the first regularization layer for data processing can improve the convergence speed and generalization ability of the image segmentation model and reduce overfitting. For example, the image embedding vector is input into the first regularization layer for feature extraction, and different features in the image embedding vector are mapped to the same value range to obtain the first feature map, even if the mean of each feature of the image embedding vector is 0 and the variance is 1.

[0080] In another possible implementation, an image enhancement layer or the like can also be set before the first attention module, and the embodiments of the present application do not limit this.

[0081] S402. Input the initial spatial feature map into the spatial adaptor for adaptive learning to obtain the target spatial feature map.

[0082] In this embodiment, the initial spatial feature map is input into a spatial adapter for adaptive learning. The spatial adapter focuses on capturing the spatial correlations of the image to be segmented in the horizontal and vertical dimensions, and obtains the target spatial feature map. The spatial adapter can enhance the performance of the image segmentation model in processing two-dimensional spatial features through fine-tuning, such as identifying the spatial patterns of tissue boundaries and anatomical structures.

[0083] Continue as Figure 5 shown, the above embedding process, feature extraction, and subsequent segmentation processing based on the target feature map can be integrated into an image segmentation model, and the spatial adapter is set after the spatial feature extraction. As Figure 6 shown, the spatial adapter includes an up-projection, an activation function, and a down-projection. The down-projection uses a simple multi-layer perceptron to compress the input initial spatial feature map to a smaller dimension, and the up-projection uses another multi-layer perceptron to expand the features compressed by the down-projection back to its original dimension to obtain the target spatial feature map. The activation function can be a ReLU function, sigmoid function, Tanh function, Leaky function, PReLU function, ELU function, Maxout function, selu function, etc.

[0084] In the embodiment of the present application, spatial feature extraction is performed on the image embedding vector to obtain the initial spatial feature map of the image to be segmented, and the initial spatial feature map is input into the spatial adapter for adaptive learning to obtain the target spatial feature map. The embodiment of the present application uses the spatial adapter to perform adaptive learning on the initial spatial feature map and further extracts features to obtain the target spatial feature map. The design of the spatial adapter can enable most of the pre-trained parameters in the image segmentation model to be fixed during the training of the entire image segmentation model. By only fine-tuning the parameters in the spatial adapter, the same segmentation effect as adjusting all parameters can be achieved, that is, the Adapeter technology of efficient fine-tuning is used to adjust the parameters, which greatly reduces the calculation and storage costs and solves the problem that it is difficult to train large image segmentation models.

[0085] Figure 7 is a schematic flowchart of a method for determining the target depth feature map in an embodiment. As Figure 7 shown, the embodiment of the present application relates to another possible implementation manner of how to obtain the target feature map of the medical image to be segmented according to the image embedding vector and the attention model in the image segmentation model, including the following steps:

[0086] S701, perform depth feature extraction on the transposed image embedding vector to obtain the initial depth feature map of the image to be segmented.

[0087] In this embodiment, continue as above Figure 5As shown, the above embedding process, transpose process, feature extraction, and subsequent segmentation process based on the target feature map can be integrated into an image segmentation model, and the depth adapter is set after the depth feature extraction. First, transpose the image embedding vector to obtain the transposed image embedding vector, and the transposed image embedding vector can be input into the second attention module, and the interaction is applied to the , focusing on learning the depth correlation between different slices, that is, using the second attention module to extract depth information and semantic information to obtain the initial depth feature map.

[0088] Among them, the first attention module and the second attention module can share parameters.

[0089] Optionally, the second attention module can be a multi-head attention module, a cross-attention module, a squeeze-and-excitation (SE) module, a convolutional block attention module (CBAM), an efficient channel attention (ECA) module, a non-local module, a global context (GC) module, a SimAM module, etc.

[0090] In a possible implementation, a second regularization layer can also be set before the first attention module, input the transposed image embedding vector into the second regularization layer to obtain a second feature map, and input the second feature map into the second attention module to extract depth information and semantic information to obtain the initial depth feature map.

[0091] In another possible implementation, an image enhancement layer can also be set before the second attention module, etc., and the embodiments of the present application do not limit this.

[0092] Optionally, multiple first attention modules and second attention modules can also be set. By stacking multiple such modules, the understanding ability of the image segmentation model for three-dimensional images can be further enhanced. This method is particularly important for three-dimensional medical image segmentation tasks because these tasks usually require accurate volume localization and structural segmentation. Through the interaction of such spatial and depth branches, the image segmentation model can better learn how to distinguish different anatomical structures and pathological changes.

[0093] S702, input the initial depth feature map into the depth adapter for adaptive learning to obtain the target depth feature map.

[0094] In this embodiment, the initial depth feature map is input into a depth adapter for adaptive learning. The depth adapter focuses on correlating the features along the depth direction, enabling the image segmentation model to understand the changes between consecutive slices, which is crucial for reconstructing three-dimensional structures and volume information.

[0095] Similarly, the depth adapter can also include the Figure 6 up-projection, activation function, and down-projection shown above. The down-projection uses a simple multi-layer perceptron to compress the input initial depth feature map to a smaller dimension, while the up-projection uses another multi-layer perceptron to expand the features compressed by the down-projection back to their original dimension to obtain the target depth feature map. The activation function can be a ReLU function, sigmoid function, Tanh function, Leaky function, PReLU function, ELU function, Maxout function, selu function, etc.

[0096] In summary, the modular design of the above spatial adapter and depth adapter enables the adapter to be easily added to the architecture of an existing image segmentation model without major modification to the entire image segmentation model. Moreover, the spatial adapter and depth adapter only introduce a small number of parameters, making the fine-tuning process more efficient. They can learn important features for specific tasks without destroying the general feature representation of the pre-trained image segmentation model, enabling the image segmentation model to quickly adapt to new tasks.

[0097] The above spatial adapter and depth adapter provide a customized fine-tuning strategy for the image segmentation task, enabling the image segmentation model to better understand and process complex data. Through this method, the image segmentation model can learn new tasks while retaining the useful knowledge and features learned from large-scale datasets.

[0098] In the embodiment of this application, depth feature extraction is performed on the transposed image embedding vector to obtain the initial depth feature map of the image to be segmented. The initial depth feature map is input into the depth adapter for adaptive learning to obtain the target depth feature map. The application embodiment uses the spatial adapter to perform adaptive learning on the initial spatial feature map and further extracts features to obtain the target spatial feature map. The design of the spatial adapter enables most of the pre-trained parameters in the image segmentation model to be fixed during the training of the entire image segmentation model. By only fine-tuning the parameters in the depth adapter, the same segmentation effect as adjusting all parameters can be achieved. This method mainly integrates medical domain-specific knowledge into the image segmentation model by means of a simple and effective Adapter technology, which can show good image segmentation performance, greatly reducing the computational and storage costs and solving the problem of difficult training of large image segmentation models.

[0099] Figure 8It is a schematic flowchart of a method for determining an image of a region of interest in an embodiment. As Figure 8 shown, the embodiment of the present application relates to a possible implementation manner of segmenting an image to be segmented based on a target feature map to obtain an image of a region of interest in the image to be segmented, including the following steps:

[0100] S801, transpose the second target feature map to obtain the transposed second target feature map;

[0101] S802, fuse the first target feature map and the transposed second target feature map to obtain a fused feature map;

[0102] S803, perform segmentation processing on the image to be segmented based on the fused feature map to obtain an image of a region of interest.

[0103] In the embodiment of the present application, since the second target feature map is obtained based on the transposed image embedding vector obtained by transposing the image embedding vector, if it is necessary to fuse the first target feature map and the second target feature map, it is necessary to first transpose the second target feature map back to its original shape to obtain the transposed second target feature map, and then fuse the first target feature map and the transposed second target feature map to obtain a fused feature map.

[0104] Optionally, the Alpha fusion algorithm, the pyramid fusion algorithm, and the Poisson fusion algorithm can be used to fuse the first target feature map and the transposed second target feature map.

[0105] In the embodiment of the present application, the second target feature map is transposed to obtain the transposed second target feature map, the first target feature map and the transposed second target feature map are fused to obtain a fused feature map, and the image to be segmented is segmented based on the fused feature map to obtain an image of a region of interest. The fused feature map of the embodiment of the present application combines different features of the image to be segmented, thereby obtaining a comprehensively integrated feature image, thereby improving the accuracy of image segmentation based on the fused feature map.

[0106] In one embodiment, performing segmentation processing on the image to be segmented based on the fused feature map to obtain an image of a region of interest includes: obtaining interactive prompt information; performing decoding processing on the image to be segmented based on the interactive prompt information and the fused feature map, that is, performing segmentation processing on the medical image to be segmented to obtain an image of a region of interest.

[0107] Among them, the interactive prompt information includes at least one of line prompt information, point prompt information, and box prompt information. Optionally, the line prompt information can receive at least one pair of lines drawn by the user. Each pair of lines includes a line for positive prompt and a line for negative prompt. The line for positive prompt is obtained by drawing on the region of interest, while the line for negative prompt is obtained by drawing outside the region of interest. A line consists of the coordinates of a series of sampled points, and overlapping sampled points are included only once. All the sampled points are concatenated into a vector, and each sampled point has a prompt label indicating whether the sampled point is positive or negative. For example, the line prompt information includes the line for positive prompt composed of the sampled points of the heart and the line for negative prompt composed of the sampled points outside the heart. The sampled points for positive prompt are labeled as 1, and the sampled points for negative prompt are labeled as -1.

[0108] In this embodiment, as described above Figure 3 As shown, the fused feature map and the interactive prompt information can be input into the decoder, and the decoder segments the medical image to be segmented according to the interactive prompt information and the fused feature map to obtain the region of interest image.

[0109] In the embodiment of the present application, interactive prompt information is obtained; based on the interactive prompt information and the fused feature map, the image to be segmented is segmented to obtain the region of interest image. The embodiment of the present application uses line prompt information, point prompt information, and box prompt information to achieve accurate segmentation of irregularly shaped images such as long strips, so as to solve the problem of generating false positive and false negative masks in the image segmentation region and improve the accuracy of image segmentation.

[0110] Figure 9 It is a schematic flowchart of a method for determining the region of interest image in another embodiment. As Figure 9 shown, the embodiment of the present application relates to a possible implementation manner of how to segment the medical image to be segmented according to the target feature map, the decoder, and the interactive prompt information to obtain the region of interest image, including the following steps:

[0111] S901, associate the image embedding vector and the interactive prompt information to obtain associated information.

[0112] Optionally, the third attention module can also be various modules shown in S401 above.

[0113] In this embodiment, as described above Figure 3 As shown, the image embedding vector and the interactive prompt information can be input into the third attention module, and the third attention module contextually associates the image embedding vector and the interactive prompt information to obtain associated information. The associated information can indicate whether each pixel point in the image embedding vector is a pixel point including the region of interest or a pixel point not including the region of interest.

[0114] S902, perform segmentation processing on the image to be segmented based on the association information and the fused feature map to obtain a region of interest image.

[0115] In this embodiment, since the association information combines the image embedding vector and the interactive prompt information, each pixel point in the image embedding vector in the association information can indicate whether it is a pixel point including the region of interest or not including the region of interest. Based on the association information, the decoder can be assisted to segment the medical image to be segmented according to the fused feature map, so as to improve the accuracy of the segmentation of the image to be segmented, that is, to obtain a more accurate region of interest image.

[0116] In the embodiment of the present application, the image embedding vector and the interactive prompt information are associated to obtain association information, and segmentation processing is performed on the image to be segmented based on the association information and the fused feature map to obtain a region of interest image. In the embodiment of the present application, the image embedding vector and the interactive prompt information are associated to obtain association information, and segmentation processing is performed on the medical image to be segmented based on the association information and the fused feature map, so as to minimize the occurrence of the situation where the image segmentation performance decreases due to incorrect prompts in the image segmentation.

[0117] It should be understood that although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0118] Based on the same inventive concept, the embodiment of the present application also provides an image segmentation device for implementing the above-mentioned image segmentation method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following image segmentation device can refer to the limitations on the image segmentation method in the above text, and will not be repeated here.

[0119] In an exemplary embodiment, as Figure 10 shown, an image segmentation device is provided, including: an extraction module 11 and a segmentation module 12, where:

[0120] An extraction module 11 is configured to perform feature extraction on the image to be segmented, so as to obtain a target feature map of the image to be segmented; the target feature map includes a first target feature map and a second target feature map; the first target feature map and the second target feature map are feature maps obtained by performing feature extraction on the image to be segmented from different perspectives;

[0121] A segmentation module 12 is configured to perform segmentation processing on the image to be segmented based on the target feature map, so as to obtain a region-of-interest image in the image to be segmented.

[0122] In one embodiment, the extraction module 11 is specifically configured to perform embedding processing on the image to be segmented to obtain an image embedding vector of the image to be segmented; perform feature extraction on the image embedding vector to obtain a first target feature map of the image to be segmented; transpose the image embedding vector to obtain a transposed image embedding vector; perform feature extraction on the transposed image embedding vector to obtain a second target feature map of the image to be segmented.

[0123] In one embodiment, the extraction module 11 is specifically configured to perform spatial feature extraction on the image embedding vector to obtain an initial spatial feature map of the image to be segmented; input the initial spatial feature map into a spatial adaptor for adaptive learning to obtain a target spatial feature map.

[0124] In one embodiment, the extraction module 11 is specifically configured to perform depth feature extraction on the transposed image embedding vector to obtain an initial depth feature map of the image to be segmented; input the initial depth feature map into a depth adaptor for adaptive learning to obtain a target depth feature map.

[0125] In one embodiment, the segmentation module 12 is specifically configured to transpose the second target feature map to obtain a transposed second target feature map; fuse the first target feature map and the transposed second target feature map to obtain a fused feature map; perform segmentation processing on the image to be segmented based on the fused feature map to obtain a region-of-interest image.

[0126] In one embodiment, the segmentation module 12 is specifically configured to obtain interaction prompt information; perform segmentation processing on the image to be segmented based on the interaction prompt information and the fused feature map to obtain a region-of-interest image.

[0127] In one embodiment, the segmentation module 12 is specifically configured to associate the image embedding vector with the interaction prompt information to obtain associated information; perform segmentation processing on the image to be segmented based on the associated information and the fused feature map to obtain a region-of-interest image.

[0128] Each module in the above image segmentation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of a computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0129] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structural diagram can be as Figure 11 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements an image segmentation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0130] Those skilled in the art can understand that Figure 11 the structure shown in

[0131] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0132] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method embodiment are implemented.

[0133] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps of the above method embodiment.

[0134] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the data need to comply with regulations.

[0135] Those of ordinary skill in the art can understand that all or part of the processes in implementing the above method embodiments can be completed by hardware instructed by a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0137] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An image segmentation method, characterized in that: The method comprises: Perform embedding processing on the image to be segmented to obtain an image embedding vector D×N×L of the image to be segmented, where N is the number of embeddings, L is the embedding length, and D is the depth; For the embedding vector N×L corresponding to each depth in the image embedding vector, the embedding vector N×L is subjected to spatial feature extraction to obtain an initial spatial feature map of the image to be segmented; the initial spatial feature map is input into a spatial adapter for adaptive learning to obtain a target spatial feature map; the spatial adapter is used to learn the spatial correlation of the three-dimensional image to be segmented in the horizontal and vertical directions, that is, the relationship between pixels in the same depth slice; Transposing the image embedding vector D×N×L to obtain a transposed image embedding vector N×D×L; For the embedding vectors D×L corresponding to each embedding number in the transposed image embedding vector, perform depth feature extraction on the embedding vectors D×L to obtain an initial depth feature map of the image to be segmented; input the initial depth feature map into a depth adapter for adaptive learning to obtain a target depth feature map; the depth adapter is used to focus on the correlation of the three-dimensional image to be segmented in the depth direction, that is, the relationship between different depth slices, and capture feature changes along the depth direction; The image to be segmented is segmented based on the target spatial feature map and the target depth feature map to obtain an image of a region of interest in the image to be segmented.

2. The method according to claim 1, characterized in that The step of performing segmentation processing on the image to be segmented based on the target spatial feature map and the target depth feature map to obtain an image of a region of interest in the image to be segmented includes: Transposing the target depth feature map to obtain a transposed target depth feature map; Fusing the target spatial feature map and the transposed target depth feature map to obtain a fused feature map; The image to be segmented is segmented based on the fused feature map to obtain the region of interest image.

3. The method according to claim 2, characterized in that The step of performing segmentation processing on the image to be segmented based on the fused feature map to obtain the region of interest image includes: Get interactive prompt information; The image to be segmented is segmented based on the interactive prompt information and the fused feature map to obtain the region of interest image.

4. The method according to claim 3, characterized in that The step of performing segmentation processing on the image to be segmented based on the interactive prompt information and the fused feature map to obtain the region of interest image includes: Associating the image embedding vector with the interactive prompt information to obtain associated information; The image to be segmented is segmented based on the association information and the fused feature map to obtain the region of interest image.

5. An image segmentation device, characterized in that: The device comprises: An extraction module is used to perform embedding processing on the image to be segmented to obtain an image embedding vector D×N×L of the image to be segmented, where N is the number of embeddings, L is the embedding length, and D is the depth; For the embedding vector N×L corresponding to each depth in the image embedding vector, the embedding vector N×L is subjected to spatial feature extraction to obtain an initial spatial feature map of the image to be segmented; the initial spatial feature map is input into a spatial adapter for adaptive learning to obtain a target spatial feature map; the spatial adapter is used to learn the spatial correlation of the three-dimensional image to be segmented in the horizontal and vertical directions, that is, the relationship between pixels in the same depth slice; Transposing the image embedding vector D×N×L to obtain a transposed image embedding vector N×D×L; For the embedding vectors D×L corresponding to each embedding number in the transposed image embedding vector, perform depth feature extraction on the embedding vectors D×L to obtain an initial depth feature map of the image to be segmented; input the initial depth feature map into a depth adapter for adaptive learning to obtain a target depth feature map; the depth adapter is used to focus on the correlation of the three-dimensional image to be segmented in the depth direction, that is, the relationship between different depth slices, and capture feature changes along the depth direction; The segmentation module is used to perform segmentation processing on the image to be segmented based on the target spatial feature map and the target depth feature map to obtain an image of a region of interest in the image to be segmented.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Image semantic segmentation method and device, equipment and storage medium

    CN113807354A

  • Video target segmentation method and device, computer equipment and storage medium

    CN116385947A

  • Image segmentation method and device, electronic equipment and storage medium

    CN116758092A