Three-dimensional organ image segmentation methods, devices and computer equipment
By using a feature calibration model based on spatial and semantic attention mechanisms, the recognition difficulties caused by scale inconsistency in 3D organ image segmentation were solved, achieving accurate segmentation of liver tumors of different sizes and improving the segmentation effect.
Patent Information
- Application Number
- CN202110512121.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-11
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-05-11
AI Technical Summary
Existing technologies struggle to accurately identify liver tumors of different scales in 3D organ image segmentation, especially small tumors, and large tumors often have incomplete margins, resulting in low identification completeness and accuracy.
An organ feature calibration model based on spatial and semantic attention mechanisms is adopted to calibrate and train 3D organ feature maps at different scales. Multiple feature maps at different scales are extracted using a feature pyramid network, and calibration processing is performed using spatial and semantic correlation information to enhance the spatial details and target semantic information of the feature maps.
It improves the accuracy and completeness of 3D organ image segmentation, enabling more precise identification of target objects such as liver tumors of different sizes, and enhancing the organ region segmentation effect.
Smart Images

Figure CN113192085B_ABST
Abstract
Description
Technical Field
[0001] This application relates primarily to the field of image processing technology, and more specifically to a three-dimensional organ image segmentation method, apparatus, and computer equipment. Background Technology
[0002] With the development of various image processing technologies, research on the processing and analysis of medical images is increasing. Taking liver tumor recognition as an example, image segmentation technology is currently commonly used to extract the liver tumor region from the acquired liver image, thereby assisting in improving the accuracy and efficiency of subsequent image processing.
[0003] In practical applications, due to the different scales of liver tumors of different types and stages, small tumors are easily overlooked in liver feature maps of the same scale during image segmentation. Furthermore, the edges of large tumors obtained through segmentation may be incomplete, resulting in low completeness and accuracy in the identification of liver tumor regions. Summary of the Invention
[0004] In view of this, this application provides a three-dimensional organ image segmentation method, the method comprising:
[0005] Obtain the 3D organ image to be segmented;
[0006] Feature extraction is performed on the three-dimensional organ images to obtain multiple three-dimensional organ feature maps at different scales;
[0007] The multiple three-dimensional organ feature maps at different scales are input into the organ feature calibration model, and the target three-dimensional organ feature map is output. The organ feature calibration model is obtained by calibrating and training three-dimensional organ features of samples at different scales based on spatial attention mechanism and semantic attention mechanism.
[0008] Using the target three-dimensional organ feature map, the three-dimensional organ image is segmented, and the organ region segmentation result of the three-dimensional organ image is output.
[0009] Optionally, the step of inputting the multiple three-dimensional organ feature maps at different scales into the organ feature calibration model and outputting the target three-dimensional organ feature map includes:
[0010] By utilizing the spatial correlation information between two three-dimensional organ feature maps at adjacent scales, the three-dimensional organ feature map at a smaller scale is calibrated to obtain the organ feature map to be determined at the corresponding scale.
[0011] By utilizing the semantic correlation information between two candidate organ feature maps at adjacent scales, the candidate organ feature map at a larger scale is calibrated to obtain a candidate organ feature map at the corresponding scale.
[0012] The candidate organ feature maps obtained at multiple different scales are fused to obtain the target three-dimensional organ feature map.
[0013] Optionally, the step of using the spatial correlation information between two three-dimensional organ feature maps of adjacent scales to calibrate the organ feature map at a smaller scale to obtain a target organ feature map at the corresponding scale includes:
[0014] The two three-dimensional organ feature maps at adjacent scales are processed to obtain a spatial attention map for the three-dimensional organ feature map at the smaller scale;
[0015] Using the spatial attention map, the corresponding smaller-scale three-dimensional organ feature map is calibrated to obtain the organ feature map of the corresponding scale.
[0016] Optionally, the processing of two three-dimensional organ feature maps at adjacent scales to obtain a spatial attention map for the three-dimensional organ feature map at a smaller scale includes:
[0017] The two three-dimensional organ feature maps at adjacent scales are merged to obtain a merged feature map;
[0018] The merged feature map is input into a spatial attention network, which outputs a spatial attention map of the smaller-scale 3D organ feature map among the two adjacent 3D organ feature maps.
[0019] Optionally, the step of using the spatial attention map to calibrate the corresponding smaller-scale three-dimensional organ feature map to obtain the organ feature map of the corresponding scale includes:
[0020] The spatial attention map is downsampled and normalized to obtain a calibrated organ feature map;
[0021] In the two three-dimensional organ feature maps of adjacent scales, the smaller-scale three-dimensional organ feature map is converted to obtain an organ feature map to be calibrated that matches the format of the corresponding spatial attention map.
[0022] The feature map of the organ to be calibrated is multiplied with the corresponding feature map of the calibrated organ to obtain the feature map of the organ to be determined at the corresponding scale.
[0023] Optionally, the step of using semantic correlation information between two candidate organ feature maps at adjacent scales to calibrate the candidate organ feature map at a larger scale to obtain a candidate organ feature map at the corresponding scale includes:
[0024] The two feature maps of the organs to be determined at adjacent scales are processed to obtain a semantic attention vector for the feature map of the organs to be determined at a larger scale.
[0025] Using the semantic attention vector, the corresponding larger-scale candidate organ feature maps are calibrated to obtain candidate organ feature maps of the corresponding scale.
[0026] Optionally, the step of processing the two organ feature maps of adjacent scales to obtain a semantic attention vector for the organ feature map of a larger scale includes:
[0027] For the two undetermined organ feature maps obtained at adjacent scales, max pooling and average pooling are performed on the channel dimension respectively. The processed organ feature vectors are then merged to obtain the semantic organ feature vector.
[0028] The semantic organ feature vector is subjected to regression processing to obtain a semantic attention vector for the feature map of the undetermined organ at a larger scale.
[0029] Optionally, the step of using the semantic attention vector to calibrate the corresponding larger-scale candidate organ feature map to obtain a candidate organ feature map of the corresponding scale includes:
[0030] The semantic attention vector is normalized to obtain the calibrated semantic organ feature vector;
[0031] The larger-scale feature map of the organ to be determined is converted to a format that matches the format of the corresponding semantic attention vector.
[0032] The feature map of the organ to be calibrated is multiplied by the feature vector of the calibrated semantic organ to obtain the candidate organ feature map at the corresponding scale.
[0033] This application also proposes a three-dimensional organ image segmentation device, the device comprising:
[0034] The image acquisition module is used to acquire three-dimensional organ images to be segmented;
[0035] The feature extraction module is used to extract features from the three-dimensional organ image to obtain three-dimensional organ feature maps at multiple different scales;
[0036] The feature calibration module is used to input the multiple three-dimensional organ feature maps of different scales into the organ feature calibration model and output the target three-dimensional organ feature map; wherein, the organ feature calibration model is obtained by calibrating and training the three-dimensional organ features of samples at different scales based on spatial attention mechanism and semantic attention mechanism;
[0037] The image segmentation module is used to segment the three-dimensional organ image using the feature map of the target three-dimensional organ, and output the organ region segmentation result of the three-dimensional organ image.
[0038] This application also proposes a computer device, the computer device comprising:
[0039] Communication module;
[0040] The memory is used to store programs that implement the three-dimensional organ image segmentation method described above;
[0041] A processor is used to load and execute the program stored in the memory to implement the steps of the three-dimensional organ image segmentation method described above.
[0042] Therefore, this application proposes a three-dimensional organ image segmentation method, apparatus, and computer device. In order to accurately identify target objects of different sizes contained in a three-dimensional organ image, such as liver tumors of different sizes, the three-dimensional organ image can be feature extracted to obtain multiple three-dimensional organ feature maps of different scales. This application proposes to use the spatial correlation information and semantic correlation information between features at different levels to calibrate the feature maps at the corresponding scales, so as to enhance the representation of target objects (such as liver tumors) in contexts at different scales. Specifically, three-dimensional organ feature maps of different scales can be input, and an organ feature calibration model trained based on spatial attention mechanism and semantic attention mechanism can be used to accurately and completely identify the edges and categories of target objects in the target three-dimensional feature map using more complete and accurate spatial detail information and target semantic information. Based on this, the three-dimensional organ image can be segmented, which greatly improves the organ region segmentation effect, that is, accurately identifies target objects of different sizes contained in the three-dimensional organ image, such as tumors of different sizes. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0044] Figure 1 This is a flowchart illustrating an optional example of the three-dimensional organ image segmentation method proposed in this application;
[0045] Figure 2 In the three-dimensional organ image segmentation method proposed in this application, features of three-dimensional organs at different scales are obtained. Figure 1 A schematic diagram of the optional feature extraction model;
[0046] Figure 3 A flowchart illustrating another alternative example of the three-dimensional organ image segmentation method proposed in this application;
[0047] Figure 4 A flowchart illustrating another alternative example of the three-dimensional organ image segmentation method proposed in this application;
[0048] Figure 5 A flowchart illustrating another alternative example of the three-dimensional organ image segmentation method proposed in this application;
[0049] Figure 6 This is a schematic diagram of an optional example of obtaining feature maps of organs to be determined based on spatial attention mechanism in the three-dimensional organ image segmentation method proposed in this application.
[0050] Figure 7 This is a flowchart illustrating another possible example of obtaining feature maps of organs to be determined based on spatial attention mechanism in the three-dimensional organ image segmentation method proposed in this application.
[0051] Figure 8 This is a schematic diagram of an optional example of obtaining candidate organ feature maps based on a semantic attention mechanism in the three-dimensional organ image segmentation method proposed in this application.
[0052] Figure 9 This is a flowchart illustrating another possible example of obtaining candidate organ feature maps based on a semantic attention mechanism in the three-dimensional organ image segmentation method proposed in this application.
[0053] Figure 10 A schematic diagram of an optional example of the three-dimensional organ image segmentation apparatus proposed in this application;
[0054] Figure 11 This is a schematic diagram of another optional example of the three-dimensional organ image segmentation device proposed in this application;
[0055] Figure 12 This is a schematic diagram of the hardware structure of an optional example of a computer device suitable for the three-dimensional organ image segmentation method and apparatus proposed in this application. Detailed Implementation
[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. For ease of description, only the parts related to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. That is to say, all other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0057] It should be understood that the terms "system," "apparatus," "unit," and / or "module" used in this application are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0058] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "a," and / or "the" are not specifically singular and may include the plural. Generally, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements. An element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes the element.
[0059] In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more. The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, features defined with "first" and "second" can explicitly or implicitly include one or more of that feature.
[0060] Furthermore, flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, the steps can be processed in reverse order or simultaneously. Additionally, other operations can be added to these processes, or one or more steps can be removed from them.
[0061] As described in the background section, image segmentation models such as convolutional neural networks utilize multiple layers of convolutional kernels of different scales to extract image features at different levels, thereby enabling the identification of organ regions (such as liver tumors) of corresponding sizes on different convolutional layers. However, this image segmentation method directly fuses information from multiple levels without considering the characteristics of different levels and their correlations, which will affect the final organ identification effect, i.e., reduce the accuracy of three-dimensional organ image segmentation.
[0062] To address the aforementioned issues, and with the application of technologies such as deep learning, machine learning, and computer vision (including Artificial Intelligence, AI) in image processing, this application proposes a three-dimensional organ image segmentation based on an attention mechanism (AM). The attention mechanism can be intuitively explained using human visual mechanisms; it can be considered a resource allocation mechanism that allocates resources based on the importance of the attention object (such as a feature extracted from an image), allocating more resources to important objects and less important or less desirable objects. This application does not elaborate on the specific working principle of the attention mechanism.
[0063] Based on this, in the embodiments of this application, the spatial and semantic correlations between organ features proposed by different convolutional layers can be obtained. The weights of different organ features can then be adjusted according to these correlations to enhance the representation of contextual information of the object to be identified (such as a liver tumor) at different scales. Specifically, this increases the weight of the network on the region of interest (i.e., the region where the object to be identified is located, such as the liver tumor region) from both channel and spatial dimensions, while decreasing the weight of non-regions of interest (such as non-tumor regions in a liver image). This allows for more effective and accurate extraction of features from regions of interest at different scales, thereby improving the 3D organ image segmentation effect and ultimately enhancing the accuracy of liver tumor identification. It should be noted that the 3D organ image segmentation method proposed in this application is not limited to the liver tumor identification scenario. The implementation process for other organ identification scenarios is similar, and will not be detailed in detail here.
[0064] Based on the above description of the technical concept of the three-dimensional organ image segmentation method proposed in this application, the following will use the application example of liver tumor recognition to explain in detail the three-dimensional organ image segmentation method proposed in this application, which includes but is not limited to the method steps described in the embodiments below.
[0065] Reference Figure 1The diagram below illustrates an optional example of the three-dimensional organ image segmentation method proposed in this application. This method can be applied to computer devices, which can be servers or terminal devices with certain data processing capabilities. The server can be a standalone physical server, a server cluster integrating multiple physical servers, or a cloud server with cloud computing capabilities. The terminal device can include, but is not limited to, smartphones, tablets, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), augmented reality (AR) devices, virtual reality (VR) devices, robots, desktop computers, etc. This application does not limit the type of computer device; it can be determined as appropriate.
[0066] like Figure 1 As shown, the three-dimensional organ image segmentation method proposed in this embodiment may include, but is not limited to, the following steps:
[0067] Step S11: Obtain the three-dimensional organ image to be segmented;
[0068] In this application, the three-dimensional organ image to be segmented can be obtained by scanning the object to be detected (such as a living organism with an organ to be identified, such as a sick patient). Since the objects to be detected are different in different application scenarios, the corresponding scanned body parts are different, and the content of the obtained three-dimensional organ image to be segmented will also change accordingly. This application does not limit the content of the three-dimensional organ image to be segmented or the method of obtaining it.
[0069] It is understood that the above-mentioned three-dimensional organ images are acquired by independent image acquisition devices in the manner described above, but not limited to it, and then sent to computer devices; alternatively, the image acquisition device of the computer device itself can perform three-dimensional image scanning of the object to be detected to obtain the corresponding three-dimensional organ images to be segmented. The specific implementation process depends on the situation and will not be described in detail in this application.
[0070] Step S12: Extract features from the three-dimensional organ images to obtain three-dimensional organ feature maps at multiple different scales;
[0071] In scenarios such as liver tumor recognition, since liver tumors vary in size, this application proposes using feature extraction networks with multiple scales to extract features from the acquired 3D organ images (i.e., downsampling at different resolutions) to obtain 3D organ feature maps at corresponding scales that contain features of liver tumors of different sizes in order to accurately identify liver tumors of different sizes contained in the 3D organ images. Specifically, a convolutional neural network can be used to encode the input 3D organ image and extract features at different depths of the network. In this way, shallow features have smaller receptive fields and widths, focusing more on local low-level texture features; while deep features have larger receptive fields and widths, focusing more on high-level semantic information. Thus, this application uses convolutional neural networks with convolutional kernels of different scales to extract features from 3D organ images, which can obtain feature maps at different scales containing multi-dimensional information.
[0072] In order to make the obtained feature map contain more comprehensive and accurate feature information, this application can choose a pyramid feature extraction network composed of fully convolutional neural networks, namely Feature Pyramid Networks (FPN), to extract features from three-dimensional organ images and obtain higher quality feature maps. This application does not limit the specific network structure of the feature pyramid network.
[0073] Reference Figure 2 The diagram shows the structure of a feature pyramid network. The feature pyramid network can include an encoder and a decoder, and the encoder and decoder can be connected by a bottleneck layer. Figure 2 The network structure shown includes an encoder that can include a series of convolutional and pooling layers to continuously reduce the feature scale; and a decoder that corresponds to the encoder, which also consists of a series of convolutional and deconvolutional layers of the corresponding scale to continuously expand the feature scale. The outputs of each scale are combined together to output a multi-scale three-dimensional organ feature map.
[0074] As can be seen, compared to other feature extraction networks such as RCNN (Region-Convolutional Neural Network), this feature pyramid network implements lateral connections between convolutional layers (downsampling) and deconvolutional layers (upsampling) at the same level. This ensures that the output 3D organ feature maps at this level contain more detailed and accurate semantic and positional information. This allows for the detection of targets (such as organs like the liver) at different scales of feature maps (i.e., 3D organ feature maps at different scales). Furthermore, since the feature maps output by each convolutional layer are derived from the fusion of features from the current layer and higher levels, the resulting feature maps at each scale possess sufficient feature representation capabilities, such as detecting spatial information and high-level semantic information in the image, which helps improve image segmentation performance. This lateral connection can be implemented using 1×1×1 convolutions, but is not limited to this implementation method and can be determined depending on the situation.
[0075] For example, this application uses a feature extraction network with five convolutional layers of different scales as an example for illustration. Figure 2 As shown, the kernel sizes of these five convolutional layers and their corresponding deconvolutional layers, from top to bottom, can be 128×128×128×32, 64×64×64×64, 32×32×32×128, 16×16×16×256, and 8×8×8×320, respectively. Based on the working principle of the feature pyramid network, layers with convolutional kernels of different sizes extract features from the input 3D organ image, resulting in 3D organ feature maps of corresponding scales. The specific implementation process is not detailed here. The bottleneck layer connecting the encoder and decoder can have a scale of 4×4×4×320, which, along with the deconvolutional layer, enables the detection and extraction of high-level semantic information. This application does not detail how the feature pyramid network implements feature extraction from the input image.
[0076] It should be noted that the size of each convolutional kernel in the feature pyramid network is not limited to... Figure 2 The dimensions shown, and the number of convolutional and deconvolutional layers in the feature pyramid network, can be determined according to actual needs, including but not limited to... Figure 2 The number of levels shown.
[0077] In practical applications of this application, the initial feature extraction model described above can be used to train the sample 3D organ images until the training constraints are met (such as the loss value of each extracted feature map being less than the loss threshold or convergence, etc. This application does not impose restrictions on the content of these conditions and can be determined as appropriate), thus obtaining a feature extraction model with a corresponding network structure. In this way, after actually acquiring the 3D organ image to be segmented, the pre-trained feature extraction model can be directly called, and the 3D organ image can be input into the feature extraction model to output multiple 3D organ feature maps at different scales. Compared with the online adjustment of feature extraction network parameters, this improves the efficiency of feature extraction processing.
[0078] In some embodiments, this application may also calculate the loss value of the obtained three-dimensional organ feature map, and adjust the network parameters of the feature extraction model accordingly to further improve the output accuracy of the feature extraction model, thereby improving the image segmentation accuracy. It should be noted that this application does not describe in detail the specific implementation process of the feature extraction model training based on the feature pyramid network, and includes but is not limited to the training implementation and application process described above.
[0079] Step S13: Input multiple three-dimensional organ feature maps of different scales into the organ feature calibration model and output the target three-dimensional organ feature map;
[0080] As described above regarding the process of determining the technical concept of this application, this application aims to achieve accurate segmentation of objects of different sizes (such as liver tumors) in three-dimensional organ images by considering the characteristics and correlations of image features at different levels while fusing multi-level feature information. Therefore, this application proposes to pre-calibrate and train three-dimensional organ features of samples at different scales based on spatial attention and semantic attention mechanisms to obtain an organ feature calibration model. The specific training implementation method is not limited in this application.
[0081] In this embodiment, since a feature pyramid network can be used to extract features from three-dimensional organ images to obtain multi-scale three-dimensional organ feature maps, and given the input image processing characteristics of this feature pyramid network, high-level feature maps focus on global high-level semantic information, while low-level feature maps focus on local low-level textures. Therefore, in the training and implementation process of the organ feature calibration model, as shown in this application... Figure 3 As shown, by calculating the spatial correlation between high-level features and low-level features, semantic information in high-level feature maps can be introduced into low-level feature maps; and by calculating the semantic correlation between low-level features and high-level features, the spatial relationship of low-level features can be introduced into high-level features, thereby ensuring that the final candidate 3D organ feature maps of multiple different scales can include the spatial and semantic information of the object to be identified in more detail and accurately.
[0082] In the training process of the above model, a supervision signal (such as the acquired three-dimensional organ image) can be added to the feature map obtained in each processing step to calculate the corresponding loss value, so as to adjust the network parameters, such as the network parameters of the feature extraction network, and improve the output reliability and accuracy of the feature extraction network.
[0083] Specifically, following the above description, after calibrating the extracted 3D organ feature maps at different scales based on the spatial attention mechanism to obtain multiple candidate 3D organ feature maps at corresponding scales, a preset loss function can be used to calculate the loss value of each candidate 3D organ feature map in order to adjust the network parameters accordingly. Similarly, after further calibrating the obtained multiple candidate 3D organ feature maps at different scales based on the semantic attention mechanism to obtain multiple candidate 3D organ feature maps at corresponding scales, a preset loss function can be used to calculate the loss value of each candidate 3D organ feature map in order to adjust the network parameters accordingly. This application does not elaborate on the type of loss function used for model training or the specific implementation process of model training based on the loss function.
[0084] Subsequently, this application can fuse multiple candidate 3D organ feature maps of different scales together, such as by using a convolutional layer to achieve the fusion processing of multiple candidate 3D organ feature maps of different scales. This allows the resulting target 3D organ feature map to include not only basic target features such as color and texture features, but also spatial detail features and target semantic features, which helps to achieve more accurate segmentation and recognition of target objects.
[0085] Step S14: Using the feature map of the target three-dimensional organ, the three-dimensional organ image is segmented, and the organ region segmentation result of the three-dimensional organ image is output.
[0086] As analyzed above, for multi-layered three-dimensional organ feature maps at different scales, this application introduces spatial and semantic correlations between features at different levels. Based on these correlations, the feature weights of the three-dimensional organ features directly extracted at the corresponding scales are adjusted, thereby enhancing the representation of the object to be identified (such as liver tumors) in contexts at different scales. In this way, by fusing multiple candidate three-dimensional organ feature maps at different scales, the characteristics of the organ features themselves, as well as the detected spatial detail information and target semantic information, can be utilized to more accurately and completely identify the edges of the target object, thereby improving the organ region segmentation effect of the three-dimensional organ image, that is, accurately identifying target objects of different sizes contained in the three-dimensional organ image, such as tumors of different sizes.
[0087] Reference Figure 4This is a flowchart illustrating another optional example of the three-dimensional organ image segmentation method proposed in this application. This embodiment can be an optional refinement of the three-dimensional organ image segmentation method described in the above embodiments, but it is not limited to the refinement implementation method described in this embodiment. Referring to the above... Figure 3 The schematic diagram of the processing flow of the three-dimensional organ image segmentation method shown in this embodiment may include:
[0088] Step S41: Obtain the three-dimensional organ image to be segmented;
[0089] Step S42: Extract features from the three-dimensional organ images to obtain three-dimensional organ feature maps at multiple different scales;
[0090] The implementation process of steps S41 and S42 can be referred to the description of the corresponding parts of the above embodiments, and will not be repeated here.
[0091] Step S43: Using the spatial correlation information between two three-dimensional organ feature maps of adjacent scales, the three-dimensional organ feature map of the smaller scale is calibrated to obtain the organ feature map of the corresponding scale.
[0092] Refer to above Figure 3 The network structure of the organ feature calibration model shown in the figure has different spatial detail information and target semantic information detected during the acquisition process of any two three-dimensional organ feature maps at different scales output by the feature extraction network. Specifically, the three-dimensional organ feature maps at relatively high levels pay more attention to target semantic information, while the three-dimensional organ feature maps at low levels pay more attention to spatial detail information.
[0093] As can be seen, higher-level 3D organ feature maps (i.e., smaller-scale feature maps) have relatively less spatial detail information. Therefore, this embodiment proposes to utilize the spatial correlation information between these two adjacent 3D organ feature maps to perform spatial information calibration processing on the higher-level 3D organ feature map, thereby enriching the spatial detail information of the corresponding scale 3D organ feature map and facilitating more accurate determination of the target location. For ease of description, this embodiment can refer to the calibrated 3D organ feature map as the corresponding scale undetermined organ feature map.
[0094] It should be noted that this application does not limit the implementation method of how to use the spatial attention mechanism to obtain the spatial correlation information between two three-dimensional organ feature maps of adjacent scales. It can be determined according to the working principle of the spatial attention mechanism. The embodiments of this application will not be described in detail here.
[0095] Step S44: Using the semantic correlation information between two candidate organ feature maps at adjacent scales, the candidate organ feature map at a larger scale is calibrated to obtain the candidate organ feature map at the corresponding scale.
[0096] Unlike the spatial information calibration process described in step S43 above, since the target semantic information of the lower-level three-dimensional organ feature map (i.e., the feature map at a larger scale) is relatively limited, in order to combine the target semantic information and more accurately identify the target region (such as the liver tumor region) in the three-dimensional organ feature map, this embodiment will continue to obtain the semantic correlation information between two candidate organ feature maps at adjacent scales. Then, by using the semantic correlation information, the lower-level three-dimensional organ feature map is semantically calibrated to obtain a candidate organ feature map containing more detailed and accurate target semantic information.
[0097] It should be understood that, for the 3D organ feature maps of different scales extracted directly from the acquired 3D organ images, this application sequentially uses the spatial correlation and semantic correlation of features at adjacent levels for calibration processing, so that the final candidate organ feature maps at each scale all contain reliable and accurate spatial detail information and target semantic information.
[0098] Furthermore, it should be noted that during the calibration process of the extracted 3D organ feature maps at different scales, depending on the needs, the semantic correlation information between two 3D organ feature maps at adjacent scales can be used first to calibrate the larger-scale 3D organ feature map. Then, spatial calibration can be performed on the calibrated feature maps at adjacent scales in the manner described in step S43 above. The implementation process is similar and will not be detailed in this application. Therefore, this application does not restrict the execution order of calibration based on spatial correlation versus calibration based on semantic correlation in the above organ feature calibration model; it can be determined as needed.
[0099] Step S45: The obtained candidate organ feature maps at multiple different scales are fused to obtain the target three-dimensional organ feature map.
[0100] Step S46: Using the feature map of the target three-dimensional organ, the three-dimensional organ image is segmented, and the organ region segmentation result of the three-dimensional organ image is output.
[0101] As described above, this application utilizes the obtained target three-dimensional organ feature map to achieve complete and accurate identification of the edges of target objects of various sizes. Based on the identification result, image segmentation can be performed to obtain the organ region segmentation result of the three-dimensional organ image, that is, to accurately identify target objects of different sizes in the three-dimensional organ image, such as liver tumors of different sizes. The implementation process of image segmentation processing using feature maps will not be described in detail in this application.
[0102] Therefore, in this embodiment of the application, feature extraction is performed on the collected three-dimensional organ feature maps to obtain three-dimensional organ feature maps of different scales. Considering that there are differences in the spatial detail information and target semantic information detected by the input network image at the corresponding level during the acquisition of three-dimensional organ feature maps of different scales, such as the loss of some spatial detail information or target semantic information, directly using the extracted three-dimensional organ feature maps of the same scale for image segmentation processing will reduce the segmentation effect of the three-dimensional organ image, and may fail to reliably and completely identify target objects of different sizes.
[0103] To address the aforementioned issues, this application proposes obtaining spatial correlation information between feature maps of adjacent scales. For feature maps output by higher-level networks that lose many details due to reduced resolution, this spatial correlation information is used to calibrate feature maps lacking spatial detail, thereby improving the spatial detail information contained in the feature map at that scale. Similarly, the semantic correlation information between feature maps of adjacent scales obtained after this calibration is used to calibrate feature maps lacking target semantic information. This ensures that the final feature maps of different scales contain relatively complete and accurate spatial detail and target semantic information. Thus, fusing these multi-scale feature maps into a single target 3D organ feature map for image segmentation can reliably and accurately identify target objects such as liver tumors at different scales in 3D organ images.
[0104] Reference Figure 5 This is a flowchart illustrating another optional example of the three-dimensional organ image segmentation method proposed in this application. This embodiment can be a further refined implementation of the three-dimensional organ image segmentation method described in the above embodiments, mainly describing a refined implementation method for the feature map calibration process, but it is not limited to the refined implementation method described in this embodiment. Figure 5 As shown, this refined implementation method may include:
[0105] Step S51: Obtain the three-dimensional organ image to be segmented;
[0106] Step S52: Input the three-dimensional organ image into the pyramid feature extraction model and output multiple three-dimensional organ feature maps at different scales.
[0107] Based on the description of the corresponding parts of the above embodiments, the pyramid feature extraction model is obtained by training the sample three-dimensional organ images based on the feature pyramid network. The specific training implementation process will not be described in detail in this embodiment of the application.
[0108] It should be noted that for feature extraction models that extract multiple 3D organ feature maps at different scales from 3D organ images, including but not limited to pyramid feature extraction models, other types of feature extraction networks can also be used for modeling according to actual needs. These will not be described in detail here.
[0109] Step S53: Process the two three-dimensional organ feature maps of adjacent scales to obtain a spatial attention map for the three-dimensional organ feature map of the smaller scale.
[0110] Step S54: Using the spatial attention map, the corresponding smaller-scale three-dimensional organ feature map is calibrated to obtain the corresponding scale of the undetermined organ feature map.
[0111] The main purpose of attention mechanisms in computer vision is to make the system focus its attention on areas of interest, such as the target objects to be identified in the three-dimensional organ images acquired in this application, specifically tumors of various sizes in a three-dimensional liver image.
[0112] Specifically, in this embodiment, soft attention based on a spatial attention mechanism can be used to process two 3D organ feature maps at adjacent scales. This involves performing corresponding spatial transformations on the spatial domain information of the two 3D organ feature maps at adjacent scales to extract key information (i.e., information about the region of interest, such as the location of a liver tumor). It should be noted that this application can specifically combine the spatial domain information of the two 3D organ feature maps at adjacent scales to obtain an organ region (i.e., the region determined by 3D coordinates) in a 3D organ feature map at a smaller scale, resulting in a feature map that more accurately represents the target object to be identified, denoted as a spatial attention map. It should be noted that this application does not limit the specific implementation method for obtaining the spatial attention map.
[0113] As analyzed above, the spatial attention map obtained in this application not only achieves more accurate localization and identification of the target object's location, but also introduces target semantic information from adjacent larger-scale three-dimensional organ feature maps. Therefore, this embodiment uses the spatial attention map to calibrate the directly extracted smaller-scale three-dimensional organ feature maps, which can improve the noise of the underlying network introduced during the extraction of the smaller-scale three-dimensional organ feature maps, as well as the spatial detail information lost by the higher-level network due to the reduced resolution. Thus, the calibrated undetermined organ feature map at this scale can more clearly reflect the more complete spatial detail information and target semantic information of the target object. The specific implementation process of the above calibration process will not be described in detail in this embodiment.
[0114] Step S55: Process the two undetermined organ feature maps obtained at adjacent scales to obtain a semantic attention vector for the undetermined organ feature map at a larger scale.
[0115] Step S56: Using the semantic attention vector, the corresponding larger-scale candidate organ feature map is calibrated to obtain the corresponding scale candidate organ feature map.
[0116] Compared with the calibration of spatial domain key information of the smaller-scale 3D organ feature maps output by each relatively lower-level network described above, considering that the larger-scale 3D organ feature maps output by the relatively higher-level network focus more on global semantic information and ignore other aspects of information, this embodiment uses two undetermined organ feature maps of adjacent scales to obtain the semantic attention vector of the undetermined organ feature map of the larger scale, so that the semantic attention vector can be introduced into the spatial detail information detected by the lower-level network.
[0117] Therefore, this embodiment of the application utilizes the semantic attention vector to calibrate the corresponding larger-scale feature maps of the organs to be determined. This not only improves the semantic information in the feature maps output by the lower-level network based on the target semantic information detected by the higher-level network, but also ensures that the candidate organ feature maps obtained after calibration can contain more accurate spatial detail information.
[0118] Step S57: The obtained candidate organ feature maps at different scales are fused to obtain the target three-dimensional organ feature map.
[0119] Step S58: Using the feature map of the target three-dimensional organ, the three-dimensional organ image is segmented, and the organ region segmentation result of the three-dimensional organ image is output.
[0120] In summary, in this embodiment, after extracting features from 3D organ images using a pyramid feature extraction model to obtain multiple 3D organ feature maps at different scales, to address the issue that different network layers may not be able to simultaneously detect complete and detailed spatial details and target semantic information in the image, and to more accurately and reliably identify target objects (such as livers and liver tumors) of different sizes in 3D organ images, this paper proposes to utilize the feature advantages contained in the feature maps output by adjacent network layers, such as spatial details or target semantic information, and process them based on corresponding attention mechanisms to obtain calibration information for calibrating the feature maps of the corresponding scale output by the network at this level. This enriches the spatial details and target semantic information contained in the feature map, thereby making the spatial and semantic features contained in the calibrated candidate organ feature map more accurate and detailed. Based on this, the 3D organ image is segmented, which greatly improves the organ region segmentation effect of the 3D organ image, accurately and completely identifies the target objects in the 3D organ image, and meets the target object localization and recognition requirements of current application scenarios.
[0121] Reference Figure 6 This is a flowchart illustrating another optional example of the three-dimensional organ image segmentation method proposed in this application. This embodiment mainly elaborates on the implementation process in the above embodiment of using the spatial correlation information between two three-dimensional organ feature maps of adjacent scales to calibrate a smaller-scale three-dimensional organ feature map and obtain a feature map of the organ to be determined at the corresponding scale. For other implementation steps of the three-dimensional organ image segmentation method, please refer to the description of the corresponding parts of the above embodiments, which will not be repeated in this embodiment. Figure 6 As shown, the refined implementation method proposed in this embodiment may include:
[0122] Step S61: Merge the feature maps of two three-dimensional organs at adjacent scales to obtain a merged feature map;
[0123] Step S62: Input the merged feature map into the spatial attention network and output the spatial attention map of the smaller scale of the three-dimensional organ feature map among the two three-dimensional organ feature maps of adjacent scales.
[0124] Reference Figure 7 The flowchart shown is a schematic diagram of a three-dimensional organ feature map calibration method. Combined with the technical concept of realizing three-dimensional organ feature map calibration based on spatial attention mechanism described in the above embodiments, it can be seen that in the process of obtaining spatial attention map, this application embodiment hopes to not only determine the spatial detail information of the three-dimensional organ feature map of the corresponding scale output by the network at this level, but also introduce the target semantic information of the three-dimensional organ feature map output by the adjacent lower-level networks.
[0125] Based on this, for a 3D organ feature map of the corresponding scale output by any level of the network, when obtaining a spatial attention map for calibrating the 3D organ feature map, the embodiments of this application can upsample the 3D organ feature map output by the network at this level to obtain a 3D organ feature map output by an adjacent lower-level network, and merge it with the 3D organ feature map output by the network at this level (such as by using, but not limited to, the Concat function) to obtain a merged feature map with more channel dimensions. Thus, the merged feature map contains more accurate and complete spatial detail information and target semantic information relative to the 3D organ feature map output by the network at this level.
[0126] Subsequently, this application inputs a merged feature map containing the 3D organ feature maps output from two adjacent network levels into the spatial attention network. Compared to directly inputting the 3D organ feature map output from the current network level into the spatial attention network, this method of acquiring spatial attention images allows the target semantic information contained in the 3D feature maps output from higher-level networks to be introduced into the 3D organ feature maps output from lower-level networks. Simultaneously, it fully utilizes the more accurate and complete spatial features contained in the 3D organ feature maps output from lower-level networks to accurately identify and locate the target object region in the 3D organ feature maps output from higher-level networks. In other words, the spatial attention map obtained in this embodiment represents more complete and accurate spatial detail information of the target object. It should be noted that this application does not limit the specific network structure of the spatial attention network; it can be determined based on the working principle of the spatial attention mechanism.
[0127] For example, in the pyramid extraction model, different network layers encode the input image at corresponding resolutions, and the output 3D organ feature maps at the corresponding scales can be denoted as Fi, where i can represent the number of network layers in the pyramid extraction model, i.e., i∈[0,L-1], and L can represent the number of network layers, such as... Figure 2 and Figure 3 The feature pyramid extraction network structure shown has encoders and decoders that each contain five layers, L=5. These five layers are denoted as Level0, Level1, Level2, Level3, and Level4, respectively. The corresponding three-dimensional organ feature maps Fi∈R are output by different levels of the network. Ci ×Di×Hi×Wi That is, the feature map of the i-th scale can be represented by a spatial transformation matrix, in which the elements are channel C, depth D, height H, and width W, respectively. The specific values can be determined according to the size of the convolution kernel of the convolutional network at this level. This application does not describe in detail the acquisition process of the three-dimensional organ features at each scale.
[0128] Based on this, in the process of obtaining the spatial attention map, this embodiment inputs the merged feature map, which contains the outputs of two adjacent network layers and is obtained through merging, into the pooling layer, and performs max pooling on it respectively (e.g., Figure 6 The Max processing shown) and average pooling processing (as shown) Figure 6 The Avg processing procedure shown is used to extract features from the two parallel, adjacent-scale 3D organ feature maps in the spatial dimension. The two processed 3D organ feature maps are then stitched together to form a single feature map T. i spa The specific implementation process will not be described in detail in this embodiment.
[0129] Next, the concatenated feature map is input into a fully connected layer composed of a multilayer perceptron (MLP) for processing. The resulting feature map is denoted as a spatial attention map, such as... Figure 7 As shown, the multilayer perceptron (MLP) in this example can be composed of a three-layer convolutional network, i.e., a three-layer perceptron, but it is not limited to this network structure and can be determined as needed.
[0130] Step S63: The spatial attention map is downsampled and normalized to obtain the calibrated organ feature map;
[0131] Following the above description, this application proposes a spatial attention mechanism to process the merged feature map after combining the features of three-dimensional organs at two adjacent scales, and regress it into an attention map A. i m Let A be the calibration organ feature map, used to calibrate the smaller-scale 3D organ feature map output by the high-level network. Therefore, in this calibration process, to enable operations between the two feature maps, regression processing can be performed on the obtained spatial attention map. Specifically, the spatial attention map can be downsampled to reduce the number of pixels it contains, and then the sigmoid activation function is used to normalize the downsampled feature map to obtain the calibration organ feature map A. i m This application does not restrict the specific implementation process.
[0132] Among them, the attention map obtained from the regression is the calibration organ feature map A. i m It can be defined as:
[0133]
[0134] In formula (1), f spa () can represent a three-layer convolutional network, i.e., a three-layer perceptron. θ can indicate the network parameters of the perceptron. The specific acquisition process of the spatial attention map can be determined based on the working principle of the spatial attention mechanism. This application will not be described in detail here. It should be noted that since this application performs spatial attention processing on the merged feature map after merging two three-dimensional organ feature maps of adjacent scales, it introduces the target semantic information detected by the larger-scale three-dimensional organ feature map output by the high-level network, which improves the accuracy of the feature weights contained in the obtained calibrated organ feature map. It can be used as a calibration reference feature map for the smaller-scale three-dimensional organ feature map to improve the spatial detail information and target semantic information in the smaller-scale three-dimensional organ feature map, which helps to completely and accurately identify the edge of the target object and improve the image segmentation effect.
[0135] Step S64: For two three-dimensional organ feature maps at adjacent scales, perform format conversion processing on the smaller-scale three-dimensional organ feature map to obtain an organ feature map to be calibrated that matches the format of the corresponding spatial attention map.
[0136] Step S65: Perform feature product operation between the feature map of the organ to be calibrated and the corresponding feature map of the calibrated organ to obtain the feature map of the organ to be determined at the corresponding scale.
[0137] In practical applications, due to format differences between the calibrated organ feature map obtained after normalization and the 3D organ feature map at the same scale directly output by the feature extraction model, direct computation between these two feature maps is not possible. Therefore, this embodiment proposes to first perform format conversion processing on the 3D organ feature map at the same scale, i.e., the smaller-scale 3D organ feature map among two adjacent scales. For example, a 1×1×1 convolutional layer can be used to process the smaller-scale 3D organ feature map. Then, the resulting organ feature map to be calibrated and the calibrated organ feature map are multiplied together. The corresponding scale of the calibrated feature map of the organ to be determined is obtained. i spa The specific calculation process will not be detailed in this application.
[0138] Therefore, the process of calculating the spatial correlation between 3D organ feature maps of any two adjacent scales, introducing high-level semantic concepts into low-level features, and performing feature calibration on the smaller-scale 3D organ feature maps to obtain the feature map of the organ to be determined can be achieved by the calculation method represented by the following formula:
[0139]
[0140] In the above formula (1), F i spa This can represent the calibrated attention features of the 3D organ feature map at the i-th scale, i.e., the features contained in the aforementioned undetermined organ feature map, which represents the most informative features in the 3D organ feature map at that scale, f. spa (F i ,F i-1 ) can be expressed as F i With F i-1 The relationship is as above. Figure 7 As shown, this can be obtained through regression processing using a perceptron consisting of three convolutional networks, but it is not limited to this acquisition method; g spa (F i ) can be used to implement F i The feature mapping, i.e., the format conversion process described in step S64 above, yields the feature map of the organ to be calibrated; C is the regularization factor, 1 / C(f spa (Fi ,F i-1 )) can represent the processing procedure described in step S63 above, therefore, That is, the above-mentioned calibration organ feature map.
[0141] In summary, compared to directly applying spatial attention processing to the 3D organ feature map at this scale for calibration, the method proposed in this embodiment calculates the spatial correlation of features contained in 3D organ feature maps at adjacent scales. This allows the target semantic information from the smaller-scale 3D organ feature map output by the higher-level network to be introduced into the larger-scale 3D organ feature map output by the lower-level network. This improves the completeness and accuracy of the spatial detail information and target semantic information contained in the resulting calibrated organ feature map. Thus, calibrating the smaller-scale 3D organ feature map can reliably and accurately identify the edges and categories of target objects, which helps to improve image segmentation results.
[0142] Reference Figure 8 This is a flowchart illustrating another optional example of the three-dimensional organ image segmentation method proposed in this application. This embodiment mainly elaborates on the process described above, which utilizes the semantic correlation information between two candidate organ feature maps of adjacent scales to calibrate a larger-scale candidate organ feature map and obtain a candidate organ feature map of the corresponding scale. For other implementation steps of the three-dimensional organ image segmentation method, please refer to the descriptions in the corresponding parts of the above embodiments; this embodiment will not repeat them here. Figure 8 As shown, the refined implementation method proposed in this embodiment may include:
[0143] Step S81: For the two undetermined organ feature maps obtained at adjacent scales, perform max pooling and average pooling on the channel dimension respectively, and merge the processed organ feature vectors to obtain semantic organ feature vectors.
[0144] Referring to the description of the spatial correlation calculation process above, in the process of calculating the semantic correlation information of these two three-dimensional organ feature maps, that is, for two undetermined organ feature maps at adjacent scales (such as F... i spa and In the process of processing to obtain semantic attention vectors for feature maps of organs of a larger scale, such as... Figure 9 As shown, in the pooling layer processing stage, max pooling can be performed on each input feature map of the organ to be determined (e.g., ...). Figure 9 The Max processing shown) and average pooling processing (as shown) Figure 9The Avg processing procedure shown merges the max pooling results of the feature maps of the organs to be determined at two adjacent scales to obtain a max pooling feature vector; and merges the average pooling results of these two feature maps to obtain an average pooling feature vector. Finally, these two feature vectors are concatenated into a single feature vector, denoted as the semantic organ feature vector T. i sem .
[0145] It should be noted that the process of performing max pooling and average pooling on the 3D organ feature map at each scale based on the semantic attention mechanism can be determined according to the working principle of the semantic attention mechanism, and will not be described in detail here.
[0146] Step S82: Perform regression processing on the semantic organ feature vector to obtain the semantic attention vector for the feature map of the undetermined organ at a larger scale.
[0147] Step S83: Normalize the semantic attention vector to obtain the calibrated semantic organ feature vector;
[0148] Similar to the spatial correlation analysis process described above, this embodiment, based on the semantic attention mechanism, can also regress an attention vector A after feature extraction processing of the feature maps of the organs to be determined at two adjacent scales in parallel along the channel dimension. i v It can be defined as This application can use the attention vector A i v This is called the calibration semantic organ feature vector.
[0149] In the embodiments of this application, such as Figure 9 As shown, in the process of obtaining the above semantic attention vector, a perceptron f composed of three convolutional networks can still be used. sem Regression processing is performed on the semantic organ feature vectors, but it is not limited to this structure of the perceptron and can be determined as appropriate.
[0150] Subsequently, this application can use the sigmoid activation function to further normalize the semantic attention vector output by the perceptron, obtaining the calibrated semantic organ feature vector A. i v It should be noted that this application does not elaborate on the type of the sigmoid activation function or the principle of its normalization process.
[0151] Step S84: For the two undetermined organ feature maps obtained at adjacent scales, perform format conversion processing on the larger scale undetermined organ feature map to obtain an uncalibrated undetermined organ feature map that matches the format of the corresponding semantic attention vector.
[0152] Step S85: Perform a product operation between the feature map of the organ to be calibrated and the feature vector of the calibrated semantic organ to obtain the candidate organ feature map at the corresponding scale.
[0153] In large-scale feature maps of undetermined organs Before calibration, a 1×1×1 convolutional network can be used to perform format conversion, so that the resulting feature map of the organ to be calibrated matches the format of the semantic attention vector. The two can then be further multiplied to achieve the desired feature map. Feature calibration processing.
[0154] Therefore, this application embodiment is based on a semantic attention mechanism to process the feature map of the organ to be determined at the i-th scale. The calibration process yields candidate organ feature maps at the corresponding scales. It can be defined as: Specifically, based on the semantic relevance analysis process described above, the embodiments of this application can calculate the candidate organ feature map at any scale according to the following formula:
[0155]
[0156] The meanings of the three parts on the right side of the equation in formula (3) are similar to those of the corresponding parts in formula (2), and will not be elaborated upon in this embodiment.
[0157] Therefore, this application, after calibrating the features of smaller-scale three-dimensional organs based on spatial attention mechanism to obtain the feature maps of the organs to be determined at the corresponding scale, further processes the feature maps of the organs to be determined at two adjacent scales based on semantic attention mechanism. That is, the spatial detail information of the lower level is introduced into the higher level features to calibrate the feature maps of the organs to be determined at the larger scale. This ensures that the edge and category of the target object of the corresponding size can be determined based on complete and detailed spatial detail information and target semantic information in the final candidate organ feature maps of each scale, thereby improving the accuracy of target object region recognition. In this way, the target three-dimensional organ feature map obtained by fusing candidate organ feature maps of different scales can be used to perform image segmentation on the three-dimensional organ image, and can reliably and accurately identify target objects of different sizes contained in the three-dimensional organ image, such as liver tumors of different sizes.
[0158] Based on the calculation process of spatial and semantic correlation between feature maps of adjacent scales described in the above embodiment, in the training process of the organ feature calibration model based on this, after calculating the spatial correlation of the input sample 3D organ feature maps of different scales in the above manner, deep supervised training can be performed on the obtained feature maps of the undetermined organ at each scale. For example, using a preset loss function, the loss value of the feature map of the undetermined organ at that scale can be calculated, and the network parameters can be adjusted accordingly to improve the accuracy of feature extraction and spatial correlation calculation between feature maps of adjacent scales. The specific training implementation process will not be described in detail in this embodiment.
[0159] Similarly, following the semantic relevance calculation method described above, based on the semantic attention mechanism, the semantic relevance between undetermined feature maps at adjacent scales is calculated to calibrate the undetermined feature maps at larger scales. After obtaining the candidate organ feature maps, deep supervised training can still be used to adjust network parameters, thereby improving the accuracy of feature extraction, spatial relevance information between feature maps at adjacent scales, and semantic relevance information between feature maps at adjacent scales. This improves the 3D organ image segmentation effect and enables accurate and reliable identification of target objects of different sizes. This application does not elaborate on the training process of the organ feature calibration model based on the above-mentioned relevance calculation method.
[0160] In some other embodiments proposed in this application, the aforementioned feature extraction network, spatial attention network, and semantic attention network can also be used to construct an initial image segmentation network. Following the network processing procedures and training methods described above, this initial image segmentation network can be trained using different sample 3D organ images to obtain a 3D organ image segmentation model. It is understood that this 3D organ image segmentation model may include sub-models such as the feature extraction model described above (e.g., the pyramid feature extraction model) and the organ feature calibration model. This application does not elaborate on the training implementation process of this 3D organ image segmentation model; however, the training process of each sub-model described above can be referenced, but is not limited to, based on the technical concept proposed in this application.
[0161] It should be understood that for any model trained above, in practical applications, the actual collected 3D organ images or obtained feature maps can be input into the model for processing. Based on the model output results, the network parameters of the model can be further optimized (such as performing one or more model iterations) to further improve the output accuracy of the model in the current application scenario. The specific optimization process is not detailed in this application.
[0162] Reference Figure 10This is a schematic diagram of an optional example of the three-dimensional organ image segmentation apparatus proposed in this application. This apparatus can be applied to the computer equipment described above. The type of computer equipment can be determined as appropriate, and this embodiment does not impose any limitations. Figure 10 As shown, the device may include:
[0163] Image acquisition module 101 is used to acquire three-dimensional organ images to be segmented;
[0164] The feature extraction module 102 is used to extract features from the three-dimensional organ image to obtain multiple three-dimensional organ feature maps at different scales;
[0165] The feature calibration module 103 is used to input the multiple three-dimensional organ feature maps of different scales into the organ feature calibration model and output the target three-dimensional organ feature map.
[0166] The organ feature calibration model is obtained by calibrating and training three-dimensional organ features of samples at different scales based on spatial attention and semantic attention mechanisms.
[0167] The image segmentation module 104 is used to segment the three-dimensional organ image using the target three-dimensional organ feature map and output the organ region segmentation result of the three-dimensional organ image.
[0168] In some embodiments, the feature calibration module 103 may include:
[0169] The first calibration processing unit is used to calibrate the three-dimensional organ feature map at a smaller scale by utilizing the spatial correlation information between two three-dimensional organ feature maps at adjacent scales, so as to obtain the organ feature map to be determined at the corresponding scale.
[0170] The second calibration processing unit is used to calibrate the feature map of the larger scale by utilizing the semantic correlation information between two feature maps of the organ to be determined at adjacent scales, so as to obtain the candidate organ feature map at the corresponding scale.
[0171] In one possible implementation, such as Figure 11 As shown, the first calibration processing unit described above may include:
[0172] Spatial attention map generation unit 1031 is used to process two three-dimensional organ feature maps at adjacent scales to obtain a spatial attention map for the three-dimensional organ feature map at a smaller scale.
[0173] The first calibration unit 1032 is used to calibrate the corresponding smaller-scale three-dimensional organ feature map using the spatial attention map to obtain the organ feature map of the corresponding scale.
[0174] Optionally, the spatial attention map obtaining unit 1031 described above may include:
[0175] The first merging processing unit is used to merge two three-dimensional organ feature maps of adjacent scales to obtain a merged feature map;
[0176] A spatial attention processing unit is used to input the merged feature map into a spatial attention network and output a spatial attention map of the smaller-scale three-dimensional organ feature map among the two three-dimensional organ feature maps of adjacent scales.
[0177] Accordingly, the first calibration unit 1032 may include:
[0178] The first normalization processing unit is used to downsample and normalize the spatial attention map to obtain a calibrated organ feature map;
[0179] The first format conversion processing unit is used to perform format conversion processing on the smaller-scale three-dimensional organ feature map among the two three-dimensional organ feature maps of adjacent scales to obtain an organ feature map to be calibrated that matches the format of the corresponding spatial attention map.
[0180] The spatial calibration unit is used to perform feature product operation on the feature map of the organ to be calibrated and the corresponding feature map of the calibrated organ to obtain the feature map of the organ to be determined at the corresponding scale.
[0181] In some other embodiments, as described above Figure 11 As shown, the second calibration processing unit described above may include:
[0182] The semantic attention vector acquisition unit 1033 is used to process the two proposed organ feature maps at adjacent scales to obtain a semantic attention vector for the proposed organ feature map at a larger scale.
[0183] The second calibration unit 1034 is used to perform calibration processing on the corresponding larger-scale candidate organ feature map using the semantic attention vector to obtain the candidate organ feature map of the corresponding scale.
[0184] Optionally, the semantic attention vector acquisition unit 1033 mentioned above may include:
[0185] The semantic organ feature vector acquisition unit is used to perform max pooling and average pooling on the channel dimension of the two undetermined organ feature maps at adjacent scales, and merge the processed organ feature vectors to obtain the semantic organ feature vector.
[0186] The regression processing unit is used to perform regression processing on the semantic organ feature vector to obtain a semantic attention vector for the feature map of the undetermined organ at a larger scale.
[0187] Accordingly, the second calibration unit 1034 may include:
[0188] The second normalization processing unit is used to normalize the semantic attention vector to obtain a calibrated semantic organ feature vector.
[0189] The second format conversion processing unit is used to perform format conversion processing on the larger-scale organ feature map among the two organ feature maps of adjacent scales obtained, so as to obtain a calibrated organ feature map that matches the format of the corresponding semantic attention vector.
[0190] The semantic calibration unit is used to perform a product operation on the feature map of the organ to be calibrated and the calibration semantic organ feature vector to obtain a candidate organ feature map of the corresponding scale.
[0191] Based on the above analysis, the feature calibration module 103 may further include:
[0192] The fusion processing unit 1035 is used to perform fusion processing on the obtained candidate organ feature maps of multiple different scales to obtain the target three-dimensional organ feature map.
[0193] It should be noted that the various modules and units in the above-mentioned device embodiments can all be stored in the memory as program modules. The processor executes the above-mentioned program modules stored in the memory to realize the corresponding functions. The functions realized by each program module and its combination, as well as the technical effects achieved, can be referred to the description of the corresponding part of the above-mentioned method embodiments. This embodiment will not repeat them here.
[0194] This application also provides a computer-readable storage medium on which a computer program can be stored. The computer program can be called and loaded by a processor to implement the various steps of the three-dimensional organ image segmentation method described in the above embodiments. The specific implementation process can be referred to the description of the corresponding part of the above embodiments, and will not be repeated in this embodiment.
[0195] Reference Figure 12 The above is a schematic diagram of the hardware structure of an optional example of a computer device suitable for the three-dimensional organ image segmentation method and apparatus proposed in this application, such as... Figure 12 As shown, the computer device may include: a communication module 121, a memory 122, and a processor 123, wherein:
[0196] The number of communication module 121, memory 122 and processor 123 can all be at least one, and communication module 121, memory 122 and processor 123 can all be connected to a communication bus to realize data interaction between them through the communication bus. The specific implementation process can be determined according to the needs of the specific application scenario, and will not be described in detail in this application.
[0197] The communication module 121 may include a communication module capable of data interaction using a wireless communication network, such as a WIFI module, a 5G / 6G (fifth-generation mobile communication network / sixth-generation mobile communication network) module, a GPRS module, etc. The communication module 121 may also include a communication interface for data interaction between internal components of a computer device, such as a USB interface, a serial / parallel port, etc. This application does not limit the specific contents included in the communication module 121.
[0198] In this embodiment, memory 122 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device. Processor 123 may be a central processing unit (CPU), application-specific integrated circuit (ASIC), digital signal processor (DSP), application-specific integrated circuit (ASIC), off-the-shelf programmable gate array (FPGA), or other programmable logic device.
[0199] In practical applications of this embodiment, the memory 122 can be used to store programs that implement the three-dimensional organ image segmentation method described in any of the above method embodiments; the processor 23 can load and execute the programs stored in the memory 122 to implement the various steps of the three-dimensional organ image segmentation method proposed in any of the above method embodiments of this application. The specific implementation process can be referred to the description of the corresponding part of the corresponding embodiment above, and will not be repeated here.
[0200] It should be understood that, Figure 12 The structure of the computer device shown does not constitute a limitation on the computer device in the embodiments of this application. In practical applications, the computer device may include more than Figure 12 The number or number of components shown, or combinations of certain components, may be determined according to the product type of the computer device. For example, if the computer device is a terminal device listed above, the computer device may also include at least one device such as a touch sensing unit for sensing touch events on a touch display panel, a keyboard, a mouse, an image acquisition device (such as a camera), a microphone, etc.; or at least one output device such as a monitor, a speaker, a vibration mechanism, a lamp, etc., which will not be listed here.
[0201] In the case where the computer device is the aforementioned terminal device, the terminal device can collect a scan of the object to be detected to obtain a three-dimensional organ image to be segmented. Then, according to the three-dimensional organ image segmentation method described above, the three-dimensional organ image can be segmented to identify target objects of different sizes, such as liver tumors of different sizes. The category of the target object can be determined according to the specific application scenario, including but not limited to the application scenario of liver tumor recognition. In some other embodiments, the terminal device can also receive three-dimensional organ images collected and sent by other devices and perform three-dimensional organ image segmentation processing in accordance with the manner described in the above embodiments. This application does not limit this and can be determined as appropriate.
[0202] In the case where the computer device is a server, a terminal device with three-dimensional image acquisition function can typically acquire the three-dimensional organ image to be segmented and send it to the server. The server then segments the three-dimensional organ image according to the three-dimensional organ image segmentation method described in the above embodiments to obtain organ region segmentation results that meet the application requirements, such as the segmentation results of liver tumors of different sizes. These results are then fed back to a preset terminal for display, assisting in the diagnosis of diseases of the object to be detected and the determination of treatment plans. The specific implementation process will not be detailed in this application.
[0203] Finally, it should be noted that the various embodiments in this specification are described in a progressive or parallel manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus and computer equipment disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
[0204] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A three-dimensional organ image segmentation method, the method comprising: Obtain the 3D organ image to be segmented; Feature extraction is performed on the three-dimensional organ images to obtain multiple three-dimensional organ feature maps at different scales; The multiple three-dimensional organ feature maps at different scales are input into the organ feature calibration model, and the target three-dimensional organ feature map is output. The organ feature calibration model is obtained by calibrating and training the three-dimensional organ features of samples at different scales based on spatial attention mechanism and semantic attention mechanism. Using the target three-dimensional organ feature map, the three-dimensional organ image is segmented, and the organ region segmentation result of the three-dimensional organ image is output. The step of inputting the multiple three-dimensional organ feature maps of different scales into the organ feature calibration model and outputting the target three-dimensional organ feature map includes: The two three-dimensional organ feature maps at adjacent scales are merged to obtain a merged feature map; The merged feature map is input into a spatial attention network, which outputs a spatial attention map of the smaller-scale three-dimensional organ feature map among the two adjacent three-dimensional organ feature maps. Using the spatial attention map, the corresponding smaller-scale three-dimensional organ feature map is calibrated to obtain the organ feature map of the corresponding scale. By utilizing the semantic correlation information between two candidate organ feature maps at adjacent scales, the candidate organ feature map at a larger scale is calibrated to obtain a candidate organ feature map at the corresponding scale. The candidate organ feature maps obtained at multiple different scales are fused to obtain the target three-dimensional organ feature map.
2. The method according to claim 1, wherein the step of using the spatial attention map to calibrate the corresponding smaller-scale three-dimensional organ feature map to obtain a feature map of the organ to be determined at the corresponding scale includes: The spatial attention map is downsampled and normalized to obtain a calibrated organ feature map; In the two three-dimensional organ feature maps of adjacent scales, the smaller-scale three-dimensional organ feature map is converted to obtain an organ feature map to be calibrated that matches the format of the corresponding spatial attention map. The feature map of the organ to be calibrated is multiplied with the corresponding feature map of the calibrated organ to obtain the feature map of the organ to be determined at the corresponding scale.
3. The method according to claim 1 or 2, wherein the step of using semantic correlation information between two candidate organ feature maps of adjacent scales to calibrate the candidate organ feature map of a larger scale to obtain a candidate organ feature map of the corresponding scale includes: The two feature maps of the organs to be determined at adjacent scales are processed to obtain a semantic attention vector for the feature map of the organs to be determined at a larger scale. Using the semantic attention vector, the corresponding larger-scale candidate organ feature maps are calibrated to obtain candidate organ feature maps of the corresponding scale.
4. The method according to claim 3, wherein processing the two undetermined organ feature maps at adjacent scales to obtain a semantic attention vector for the undetermined organ feature map at a larger scale includes: For the two undetermined organ feature maps obtained at adjacent scales, max pooling and average pooling are performed on the channel dimension respectively. The processed organ feature vectors are then merged to obtain the semantic organ feature vector. The semantic organ feature vector is subjected to regression processing to obtain a semantic attention vector for the feature map of the undetermined organ at a larger scale.
5. The method according to claim 3, wherein the step of using the semantic attention vector to calibrate the corresponding larger-scale candidate organ feature map to obtain a candidate organ feature map of the corresponding scale includes: The semantic attention vector is normalized to obtain the calibrated semantic organ feature vector; The larger-scale feature map of the organ to be determined is converted into a format to obtain a feature map of the organ to be calibrated that matches the format of the corresponding semantic attention vector. The feature map of the organ to be calibrated is multiplied by the feature vector of the calibrated semantic organ to obtain the candidate organ feature map at the corresponding scale.
6. A three-dimensional organ image segmentation device, the device comprising: The image acquisition module is used to acquire three-dimensional organ images to be segmented; The feature extraction module is used to extract features from the three-dimensional organ image to obtain three-dimensional organ feature maps at multiple different scales; The feature calibration module is used to input the multiple three-dimensional organ feature maps of different scales into the organ feature calibration model and output the target three-dimensional organ feature map; wherein, the organ feature calibration model is obtained by calibrating and training the three-dimensional organ features of samples at different scales based on spatial attention mechanism and semantic attention mechanism; The image segmentation module is used to segment the three-dimensional organ image using the feature map of the target three-dimensional organ and output the organ region segmentation result of the three-dimensional organ image. The feature calibration module includes: The first calibration processing unit is used to merge two three-dimensional organ feature maps of adjacent scales to obtain a merged feature map; input the merged feature map into a spatial attention network, and output a spatial attention map of the smaller-scale three-dimensional organ feature map among the two three-dimensional organ feature maps of adjacent scales; use the spatial attention map to calibrate the corresponding smaller-scale three-dimensional organ feature map to obtain a feature map of the undetermined organ at the corresponding scale. The second calibration processing unit is used to calibrate the feature map of the larger scale by utilizing the semantic correlation information between two feature maps of the organ to be determined at adjacent scales, so as to obtain the candidate organ feature map at the corresponding scale. The fusion processing unit is used to fuse the obtained candidate organ feature maps of multiple different scales to obtain the target three-dimensional organ feature map.
7. A computer device, the computer device comprising: Communication module; A memory for storing a program that implements the three-dimensional organ image segmentation method as described in any one of claims 1 to 5; A processor is configured to load and execute the program stored in the memory to implement the steps of the three-dimensional organ image segmentation method as described in any one of claims 1 to 5.