A medical image organ segmentation method based on semantic context relationship and geometric perception prior
By introducing semantic context and geometric priors into 3D medical images and combining them with deep convolutional networks, the inaccuracy of organ segmentation in existing technologies is solved, achieving more accurate organ segmentation results.
Patent Information
- Application Number
- CN202411242874.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-09-05
AI Technical Summary
Existing technologies lack modeling of 3D spatial context information in 3D medical image segmentation, resulting in broken regions and incorrect topological structures in organ segmentation, making it impossible to correctly restore the shape and structure of organs.
We employ a method based on semantic context and geometric perception priors. By extracting geometric information as constraints, we combine it with a deep 3D convolutional coding and decoding network to inject spatial context semantic relationships, perform feature fusion, and use activation functions for organ segmentation.
It improves the accuracy and consistency of image segmentation, enhances the model's understanding of image geometric features, and enables more accurate identification and segmentation of complex 3D medical images.
Smart Images

Figure CN119251491B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing, and in particular to a method for organ segmentation in medical images based on semantic context and geometric perception priors. Background Technology
[0002] Automated multi-organ segmentation is crucial for advancing computer-aided medical image analysis and disease diagnosis. Its aim is to accurately assign voxels from medical 3D volumetric data to multiple organ regions and correctly reconstruct the organ's geometry. In practice, automated multi-organ segmentation not only reduces the workload of manual annotation by doctors and radiologists but also opens up possibilities for clinical applications such as surgery, radiotherapy, and morphological evaluation. Thanks to the powerful generalization capabilities of deep neural networks, encoder-decoder models based on convolutional neural networks and visual transformers (ViT) have been widely applied to medical image segmentation tasks and have achieved significant performance results.
[0003] However, the lack of modeling for 3D spatial context information limits the ability of these models to capture complex 3D structures. Furthermore, existing models tend to implicitly learn the implied contextual features using local perception and global dependencies, but these models cannot correctly recover the shape and structure of organs, resulting in discontinuous regions and incorrect topological structures in 3D volume segmentation. In practical applications, pathological diagnosis typically requires consideration of the semantic structure of organs and surrounding tissues. On the other hand, prior anatomical knowledge of abdominal organs is also crucial for radiologists to correctly label voxels. Summary of the Invention
[0004] To overcome the shortcomings of the prior art in correctly segmenting organs, this invention provides a medical image organ segmentation method based on semantic context and geometric perception priors.
[0005] To achieve the above-mentioned technical effects, the technical solution of the present invention is as follows:
[0006] A medical image organ segmentation method based on semantic context and geometric perception priors includes the following steps:
[0007] Geometric information is extracted from multi-organ medical image data with true segmentation values and used as geometric prior constraints.
[0008] The multi-organ medical image data to be segmented is preprocessed and then input into a three-dimensional convolutional encoder-decoder network to extract multi-level features;
[0009] Calculate spatial context semantic relationships for multi-organ medical imaging data, and inject the spatial context semantic relationships into the multi-level features;
[0010] The geometric prior constraints and the multi-level features are fused through the encoding-decoding network, and organ segmentation is performed on the multi-organ medical image data to be segmented based on the activation function to obtain the segmentation result.
[0011] This invention also proposes a medical image organ segmentation system based on semantic context and geometric perception priors, applying the medical image organ segmentation method based on semantic context and geometric perception priors proposed in this invention. The system includes:
[0012] Geometric Information Constraint Module: Used to extract geometric information from multi-organ medical image data with segmentation truth values as geometric prior constraints;
[0013] Multi-level feature extraction module: used to extract multi-level features from input medical organ image data through a three-dimensional convolutional encoder-decoder network;
[0014] Semantic relation prior propagation module: used to calculate the spatial context semantic relations of the input medical organ image data, and inject the spatial context semantic relations into the multi-level features step by step;
[0015] Organ segmentation module: It performs feature fusion of geometric prior constraints and multi-level features through an encoding-decoding network, and performs organ segmentation on the multi-organ medical image data to be segmented based on the activation function to obtain the segmentation result.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0017] This invention extracts geometric information from existing segmented images as geometric prior constraints, ensuring the consistency of geometric information and improving the model's understanding of image geometric features, thus making the final image processing results more accurate. Within the framework of a deep 3D convolutional coding and decoding model, the accuracy of voxel spatial contextual semantic relationships is considered. By analyzing the spatial relationships between voxels, the model's understanding of spatial structure is enhanced. Furthermore, spatial contextual semantic relationships are injected into multi-level features, enhancing the model's ability to express features, thereby providing more accurate semantic segmentation results during image segmentation. Attached Figure Description
[0018] Figure 1 This is a flowchart of the medical image organ segmentation method based on semantic context and geometric perception prior in Example 1.
[0019] Figure 2 This is an architecture diagram of the medical image organ segmentation system based on semantic context and geometric perception priors in Example 2. Detailed Implementation
[0020] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the invention.
[0021] It will be understood by those skilled in the art that some well-known descriptions may be omitted in the accompanying drawings.
[0022] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0023] Example 1
[0024] This embodiment proposes a medical image organ segmentation method based on semantic context and geometric perception priors, such as... Figure 1 The diagram shown is a flowchart of the medical image organ segmentation method based on semantic context and geometric perception priors in this embodiment.
[0025] The medical image organ segmentation method based on semantic context and geometric perception prior proposed in this embodiment includes the following steps:
[0026] Geometric information is extracted from multi-organ medical image data with true segmentation values and used as geometric prior constraints.
[0027] The multi-organ medical image data to be segmented is preprocessed and then input into a three-dimensional convolutional encoder-decoder network to extract multi-level features;
[0028] Calculate spatial context semantic relationships for multi-organ medical imaging data, and inject the spatial context semantic relationships into the multi-level features;
[0029] The geometric prior constraints and the multi-level features are fused through the encoding-decoding network, and organ segmentation is performed on the multi-organ medical image data to be segmented based on the activation function to obtain the segmentation result.
[0030] In this embodiment, geometric information is first extracted from the existing segmented multi-organ medical image data as a geometric prior constraint for subsequent model prediction, ensuring the consistency of geometric information and improving the model's understanding of image geometric features. Subsequently, within the framework of a deep 3D convolutional coding and decoding model, spatial context semantic relationships are extracted, and the spatial relationships between voxels are analyzed to enhance the model's understanding of spatial structure. This enables the model to identify and understand the implicit spatial layers and structures in the image, demonstrating higher accuracy and efficiency. Then, spatial context semantic relationships are injected into multi-level features. This multi-level feature fusion strategy allows the model to capture multi-scale information from local to global perspectives and to provide more refined and accurate segmentation based on different levels of semantic information when performing tasks such as image segmentation. This results in more precise semantic segmentation results. Finally, the spatial context semantic relationships and geometric prior constraints are fused to segment the data to be segmented, yielding the final result. By comprehensively considering geometric and semantic information, the model can effectively reduce errors and improve segmentation quality when processing highly complex and variable image data.
[0031] In one alternative embodiment, the geometric information includes geometric contours and geometric normals.
[0032] Specifically, for a given 3D medical image dataset with voxel-by-voxel annotation and the corresponding voxel-by-voxel-annotated voxel-level multi-organ segmentation ground truth map, data augmentation is performed by random voxel cropping, and voxel data is normalized by truncation values before geometric contour and geometric normal division is performed.
[0033] In this embodiment, multiple organs can be accurately divided using geometric contours; geometric normals can reflect the directionality and structural information of the organ surface, generating more realistic surface textures and shadow effects in the model, thus improving the visualization quality of the image.
[0034] In one optional embodiment, the geometric contour segmentation step includes: annotating the segmentation ground truth of multi-organ medical image data, determining the boundary voxels of each organ on an organ-by-organ basis, and obtaining the geometric contour.
[0035] Specifically, by using the ground truth annotations of voxels in multi-organ medical imaging data, the boundary voxels of each organ are determined organ by organ. From axial, sagittal, and coronal three-plane views, the 3D segmentation annotations are sliced into 2D slices. In each view, the boundaries of the 2D semantic mask are extracted slice by slice, and the corresponding slice index number is recorded. Then, based on the recorded 3D coordinates, all boundary points on the axial, sagittal, and coronal planes are collected, and the 2D slices are fused into 3D volume data to generate the geometric contour of the 3D organ.
[0036] In this implementation, by combining boundary information from different perspectives, a more accurate and comprehensive three-dimensional reconstruction can be achieved. Multi-view fusion helps to overcome the spatial information that may be ignored by a single perspective, thereby generating a more accurate three-dimensional model of an organ or lesion.
[0037] In an optional embodiment, the step of dividing the geometric normal includes:
[0038] Based on the geometric contour, for any single organ, a search sphere of a specific radius space is selected, and the nearest point of the voxel point on the current organ contour to the center of the search sphere is retrieved to obtain the nearest neighbor voxel.
[0039] Based on the nearest neighbor voxels, the surface normal vector is estimated using covariance analysis to obtain the geometric normal.
[0040] Furthermore, to ensure that the normal direction faces outward from the voxel, it is necessary to further analyze the semantic information of the voxel block to which the normal is pointing to determine whether to reverse the surface normal.
[0041] In this embodiment, by using a search sphere based on the geometric contour to retrieve nearest neighbor voxels and employing covariance analysis to estimate the surface normal vector, the normal vector of the organ or lesion surface can be determined more accurately. This reduces errors caused by local surface unevenness or noise, resulting in more accurate and reliable normal vectors. Furthermore, by analyzing the semantic information of the voxel block to which the normal vector points, it is ensured that the normal direction is correctly oriented outward from the voxel, thus resolving the problem of inconsistent normal estimation directions in traditional methods.
[0042] In an optional embodiment, the step of computing spatial context semantic relations includes: setting a domain range for voxel blocks in multi-organ medical image data, sequentially recording the semantic labels of neighboring voxels within the neighborhood range of each voxel block, and constructing voxel-level spatial context semantic relations.
[0043] As an example, voxel points are selected sequentially as centers, and the semantic labels of 26 three-dimensional neighbor voxels in a 3×3×3 grid around the center voxel are recorded sequentially as prior truth information of multi-scale semantic relationships.
[0044] In this embodiment, local semantic understanding is enhanced:
[0045] By analyzing the semantic labels of the neighboring voxels of the central voxel, the model can obtain detailed information about the surrounding environment of each voxel, which helps to better understand the semantic and structural features of local regions, help to understand complex or closely connected multi-organ regions, and improve its ability to discriminate new samples.
[0046] In an optional embodiment, the step of injecting spatial context semantic relations into the multi-level features includes: injecting context semantic relations into the multi-level features level by level through the multi-level bypass output structure and feature pyramid in the encoding-decoding network.
[0047] Further optionally, the step of injecting contextual semantic relationships into multi-level features step by step includes: predicting semantic relationships at different levels using a multi-level bypass output structure; and propagating the predicted high-level feature layer semantic relationship information to low-level feature representations along the feature pyramid structure by introducing recurrent connections and three-dimensional convolutions, the expression of which is:
[0048]
[0049] in, This represents the prediction result of the semantic context relationship at level i; Represents the i-th level semantic context relation prediction network; Conv i Represents a 3D convolutional layer for the i-th level feature; CAT represents the feature concatenation operation; l represents the number of feature layers. This represents the feature of the (i-1)th layer.
[0050] In this embodiment, by injecting contextual semantic relationships into multi-level features, the hierarchy and richness of feature expression are enhanced, enabling the model to more accurately capture structural details at different scales and depths when processing complex medical images. Utilizing recurrent connections and 3D convolution techniques effectively transfers high-level semantic information to low-level feature representations, helping to improve the model's accuracy in organ segmentation. Furthermore, the feature pyramid structure optimizes the information flow from high to low levels, ensuring effective integration of information between different levels. This structure allows high-level global information to guide low-level local processing, improving the overall network's ability to process complex images.
[0051] In an optional embodiment, the feature fusion step includes: first, fusing the geometric prior constraints into the first-level features of the multi-level features; then, fusing the multi-level features injected with spatial context semantic relationships and the first-level features fused with geometric prior constraints using a three-dimensional convolutional layer, the expression of which is as follows:
[0052]
[0053] in, This represents a feature that integrates spatial context semantic relations and geometric prior constraints; Conv(·) represents a 3D convolutional layer; CAT represents a feature concatenation operation. This represents the prediction result of the semantic context relationship at level i; This represents a first-level feature with geometric prior constraints.
[0054] In this embodiment, by combining geometric priors and spatial context information, the model can more comprehensively understand and analyze objects and structures in three-dimensional space. Utilizing three-dimensional convolutional layers to fuse features and geometric information from different levels helps the model maintain high recognition and segmentation accuracy even when faced with image noise, blurring, or other quality issues. Furthermore, it enhances the robustness of features and improves the reliability of the model in practical applications.
[0055] In an optional embodiment, the step of performing organ segmentation on the multi-organ medical image data to be segmented based on the activation function includes:
[0056] Construct a semantic segmentation prediction head using convolutional layers and activation functions;
[0057] The probability or score of multiple organs at each voxel is calculated using the semantic segmentation prediction head and multi-level features, and organ classification is achieved by setting a threshold; the expression is as follows:
[0058]
[0059] in, P represents the probability or score of multiple organs at each voxel point. (seg) This is the semantic segmentation prediction head.
[0060] In this embodiment, a specially designed semantic segmentation prediction head can effectively calculate the probability or score of each voxel corresponding to multiple organs, making the classification process more accurate and reducing classification errors. This is particularly true when organs are similar or adjacent in shape, significantly improving segmentation accuracy. Furthermore, operations are performed directly on the feature map through convolutional layers, requiring fewer computational resources compared to fully connected layers. This makes the entire segmentation process more efficient, enabling rapid image processing and analysis.
[0061] Example 2
[0062] This embodiment proposes a medical image organ segmentation system based on semantic context and geometric perception priors, applying the medical image organ segmentation method based on semantic context and geometric perception priors proposed in Embodiment 1. For example... Figure 2 The diagram shown is an architecture diagram of the medical image organ segmentation system based on semantic context and geometric perception priors in this embodiment.
[0063] This embodiment proposes a medical image organ segmentation system based on semantic context and geometric perception priors, which includes:
[0064] Geometric Information Constraint Module: Used to extract geometric information from multi-organ medical image data with segmentation truth values as geometric prior constraints;
[0065] Multi-level feature extraction module: used to extract multi-level features from input medical organ image data through a three-dimensional convolutional encoder-decoder network;
[0066] Semantic relation prior propagation module: used to calculate the spatial context semantic relations of the input medical organ image data, and inject the spatial context semantic relations into the multi-level features step by step;
[0067] Organ segmentation module: It performs feature fusion of geometric prior constraints and multi-level features through an encoding-decoding network, and performs organ segmentation on the multi-organ medical image data to be segmented based on the activation function to obtain the segmentation result.
[0068] It is understood that the system in this embodiment corresponds to the method in Embodiment 1 above, and the options in Embodiment 1 above are also applicable to this embodiment, so they will not be described again here.
[0069] The terminology used in the accompanying drawings is for illustrative purposes only and should not be construed as limiting the scope of this patent.
[0070] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for organ segmentation in medical images based on semantic context and geometric perception priors, characterized in that, Includes the following steps: Geometric information is extracted from multi-organ medical image data with true segmentation values and used as geometric prior constraints. The multi-organ medical image data to be segmented is preprocessed and then input into a three-dimensional convolutional encoder-decoder network to extract multi-level features; Calculate spatial context semantic relationships for multi-organ medical imaging data, and inject the spatial context semantic relationships into the multi-level features; The geometric prior constraints and the multi-level features are fused through the encoding-decoding network, and organ segmentation is performed on the multi-organ medical image data to be segmented based on the activation function to obtain the segmentation result. The step of injecting spatial context semantic relations into the multi-level features includes: injecting context semantic relations into the multi-level features step by step through the multi-level bypass output structure and feature pyramid in the encoding-decoding network; The step of injecting contextual semantic relationships into multi-level features includes: predicting semantic relationships at different levels using a multi-level bypass output structure; and propagating the predicted high-level feature layer semantic relationship information along the feature pyramid structure to the low-level feature representation by introducing recurrent connections and three-dimensional convolutions. The expression is as follows: in, Indicates the first i Level semantic context relationship prediction results; Indicates the first i Level semantic context relation prediction network; Indicates the first i Three-dimensional convolutional layers with level features; CAT Indicates a feature cascade operation; Indicates the number of feature layers. Indicates the first Layer features; The feature fusion steps include: first, fusing the geometric prior constraints into the first-level features of the multi-level features; then, fusing the multi-level features injected with spatial context semantic relationships and the first-level features fused with geometric prior constraints using a three-dimensional convolutional layer, as shown in the following expression: in, This indicates a feature that integrates spatial context semantic relations and geometric prior constraints. Represents a three-dimensional convolutional layer; CAT Indicates a feature cascade operation; Indicates the first i Level semantic context relationship prediction results; This represents a first-level feature with geometric prior constraints.
2. The medical image organ segmentation method based on semantic context and geometric perception prior as described in claim 1, characterized in that, The geometric information includes the geometric profile and geometric normal.
3. The medical image organ segmentation method based on semantic context and geometric perception prior as described in claim 2, characterized in that, The geometric contour segmentation steps include: annotating the segmentation ground truth of multi-organ medical image data, determining the boundary voxels of each organ, and obtaining the geometric contour.
4. The medical image organ segmentation method based on semantic context and geometric perception prior as described in claim 2, characterized in that, The steps for dividing the geometric normals include: Based on the geometric contour, for any single organ, a search sphere of a specific radius space is selected, and the nearest point of the voxel point on the current organ contour to the center of the search sphere is retrieved to obtain the nearest neighbor voxel. Based on the nearest neighbor voxels, the surface normal vector is estimated using covariance analysis to obtain the geometric normal.
5. The medical image organ segmentation method based on semantic context and geometric perception prior as described in claim 1, characterized in that, The steps for calculating spatial context semantic relationships include: defining the domain range for voxel blocks in multi-organ medical imaging data, sequentially recording the semantic labels of neighboring voxels within the neighborhood range of each voxel block, and constructing voxel-level spatial context semantic relationships.
6. The medical image organ segmentation method based on semantic context and geometric perception prior as described in claim 1, characterized in that, The steps for organ segmentation of multi-organ medical image data to be segmented based on activation functions include: Construct a semantic segmentation prediction head using convolutional layers and activation functions; The probability or score of multiple organs at each voxel is calculated using the semantic segmentation prediction head and multi-level features, and organ classification is achieved by setting a threshold; the expression is as follows: in, The probability or score of multiple organs at each voxel point. This is the semantic segmentation prediction head.
7. A medical image organ segmentation system based on semantic context and geometric prior, applied to the medical image organ segmentation method based on semantic context and geometric prior as described in any one of claims 1 to 6, characterized in that, The system includes: Geometric Information Constraint Module: Used to extract geometric information from multi-organ medical image data with segmentation truth values as geometric prior constraints; Multi-level feature extraction module: used to extract multi-level features from input medical organ image data through a three-dimensional convolutional encoder-decoder network; Semantic relation prior propagation module: used to calculate the spatial context semantic relations of the input medical organ image data, and inject the spatial context semantic relations into the multi-level features step by step; Organ segmentation module: It performs feature fusion of geometric prior constraints and multi-level features through an encoding-decoding network, and performs organ segmentation on the multi-organ medical image data to be segmented based on the activation function to obtain the segmentation result.
Citation Information
Patent Citations
Automatic organ segmentation method based on laryngoscope medical image
CN114841320A
Processing image data sets
US20210358101A1