Liver surface landmark segmentation method based on double-flow grid convolutional neural network
Through the dual-stream grid convolutional neural network TSMCN, the problem of inaccurate segmentation of the liver surface anatomical structure in existing methods is solved, accurate segmentation of key areas of the liver is achieved, and the accuracy of AR navigation and surgical efficiency are improved.
Patent Information
- Application Number
- CN202510875580.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-10
AI Technical Summary
Existing automated segmentation methods based on 3D mesh processing have difficulty accurately segmenting key anatomical structures on the liver surface, such as the hepatic ridge and falciform ligament, resulting in inaccurate AR navigation registration during surgery. Existing methods cannot adapt to the diversity of liver shape and appearance when extracting features from a single perspective, and lack an effective feature fusion mechanism, resulting in information loss and insufficient segmentation performance.
A two-stream grid convolutional neural network (TSMCN) is used, which includes an edge processing stream (E-stream) and a point processing stream (P-stream). It extracts geometric representations from grid edges and coordinate points, respectively, and fuses features through a fine-grained aggregate attention mechanism (FGA) module to generate segmentation probabilities for key anatomical regions. The model is optimized using weighted cross entropy and Dice loss functions.
It achieves accurate automatic segmentation of key anatomical areas on the liver surface, improves the accuracy and stability of AR navigation, reduces the time for manual labeling, improves surgical efficiency and segmentation accuracy, and enhances adaptability to changes in liver shape.
Smart Images

Figure CN120765671A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a liver surface landmark segmentation method based on a dual-stream grid convolutional neural network. Background Art
[0002] Laparoscopic liver resection based on augmented reality (AR) has been gradually applied in clinical practice. AR navigation provides important technical support and guarantee for precise treatment. This work requires the alignment of the preoperative 3D liver model (mesh) with the intraoperative 2D laparoscopic image, which largely relies on the use of important anatomical structures such as the hepatic ridge and falciform ligament as registration constraints. At present, registration mainly relies on surgeons to manually annotate these anatomical structures on the 3D mesh image of the liver, which is time-consuming and error-prone, hindering the application of AR navigation in surgery. It can be said that the automatic segmentation of the hepatic ridge and falciform ligament is the registration constraint necessary to achieve accurate 3D-2D fusion, and has become a key driving factor for AR applications.
[0003] Automatically segmenting key anatomical structures from a liver mesh requires a comprehensive understanding of the mesh's global geometry (spatial relationships) and local topology (mesh cell face shapes). Existing segmentation methods, such as the point-based method PointNet++ and mesh-specific methods (Mesh Convolutional Neural Networks, MeshCNN), primarily focus on simple geometric properties of the mesh from a single perspective (mesh vertices or edges) and train single-stream networks for automatic segmentation. These methods struggle to adapt to the vast variations in liver shape and appearance, as well as limited data availability.
[0004] 1. Single feature extraction perspective: Most existing methods only extract features from a single perspective, such as extracting features based on local topological properties of mesh edges or relying solely on the coordinate position information of point cloud data. These methods are unable to fully capture the complex and changing geometry and topology of the liver and are unable to fully adapt to the diversity of liver shapes and anatomical regions.
[0005] 2. Lack of an effective feature fusion mechanism: The simple processing method based on local attributes easily loses key information during the feature fusion process, fails to balance the contributions of different features, and struggles to form an accurate and comprehensive representation of the liver's anatomical structure.
[0006] 3. Key structural information is easily lost: During the processing process, some methods may misjudge and discard information that is crucial for the segmentation of key anatomical areas of the liver (such as the falciform ligament and hepatic ridge), resulting in the inability to accurately reflect the true morphology and position of these key structures, affecting the accuracy of the segmentation results;
[0007] 4. Limited complex task processing capability: due to the above limitations in feature extraction and fusion, the existing method is difficult to consider multiple factors when facing the segmentation task of the complex anatomical structure of the liver, and cannot effectively deal with complex tasks. The overall segmentation performance is poor, and it is difficult to meet the actual application requirements.
[0008] However, different basic attributes reveal completely different geometric information, such as spatial relationship and local topology, which are crucial for liver surface landmark segmentation task. The existing automatic segmentation method based on 3D mesh processing is difficult to realize accurate landmark segmentation due to the large shape variation of the liver and limited data. SUMMARY
[0009] The purpose of the present application is to overcome the defects of the prior art and provide a liver surface landmark segmentation method based on a two-stream mesh convolutional neural network, which can realize accurate automatic segmentation of key anatomical regions of the liver surface, and is beneficial to improve the effect of augmented reality (AR) guided liver surgery.
[0010] The purpose of the present application can be achieved by the following technical solutions: a liver surface landmark segmentation method based on a two-stream mesh convolutional neural network, comprising the following steps:
[0011] S1, constructing a data set containing liver mesh data and corresponding annotations;
[0012] S2, building a two-stream mesh convolutional neural network (Two-stream MeshCNN, TSMCN) containing two parallel streams of edge processing stream (E-stream) and point processing stream (P-stream) for extracting geometric representation from mesh edges and coordinate points respectively;
[0013] S3, training and optimizing TSMCN using the data set to obtain a liver image segmentation model;
[0014] S4, inputting the current liver mesh data into the liver image segmentation model to output the segmentation result of the corresponding key anatomical region.
[0015] Further, the step S1 specifically acquires a plurality of liver mesh data from a public liver data set and performs corresponding annotations.
[0016] Further, the step S1 specifically annotates the key anatomical regions of the liver in the liver mesh data.
[0017] Further, the step S2 comprises the following steps:
[0018] S21, constructing an edge processing stream E-stream to capture the local topology of triangular patches from all edges in the liver mesh, and outputting an edge feature map;
[0019] The point processing stream P-stream is constructed with the coordinates of each edge center point as input to capture the spatial relationship complementary to the edge features, and outputs a point feature map;
[0020] S22, based on the attention mechanism of fine-grained aggregation (FGA), an FGA module is constructed to adjust the fused edge feature map and point feature map, realize cross-view feature aggregation, and output the accumulated aggregated features;
[0021] S23, a decoder branch is constructed to decode the accumulated aggregated features to generate the segmentation probability of the key anatomical region.
[0022] Further, the working process of the edge processing stream E-stream includes: based on all edges of the liver mesh, the input edge features are defined according to the topological properties (dihedral angle, internal angle, edge length ratio) to generate an input matrix with a size of Mx5, M is the number of mesh edges;
[0023] The input matrix is converted to a new feature space with a fixed dimension of Mx8 by using an input transformer composed of MeshConv;
[0024] The MeshSE layer is introduced to enhance the expression ability of single-view features by learning the channel relationship:
[0025]
[0026] wherein, is the enhanced feature, are the input and output features respectively, c represents the number of channels, Sigmoid(·), FC(·), Pool(·) and ⊙ operations represent the application of Sigmoid activation function, full connection layer, global average pooling process and Hadamard product operation respectively;
[0027] The enhanced feature is input into K residual MeshConv modules to extract more representative geometric features:
[0028]
[0029] In addition, combined with the point features output by the point processing stream P-stream, a queue for mesh pooling is constructed to change the original edge features according to the compressed mesh structure, and three scale pooling operations are performed to make the E-stream transition from capturing local information to learning more reliable global topological information.
[0030] Furthermore, the working process of the point processing stream P-stream includes: taking the coordinates of each edge center point as input, the size is M×3, and using the input transformer composed of MeshConv to convert the input into a fixed dimension of M×8;
[0031] MeshSE and K residual MeshConv blocks are used to enhance the expression of spatial geometric information in point features, and mesh reconstruction is combined to aggregate and pool point features:
[0032]
[0033] Among them, Mesh 1 The liver mesh is reconstructed by ReMesh. After reconstruction, the point features and edge features Continue to share the same mesh structure Mesh 1 , thereby extracting complementary single-view geometric representations at different levels.
[0034] Furthermore, the working process of the FGA module includes: firstly generating a compact feature descriptor D using 1×1 convolution e ∈M×C / 4 and D p ∈M×C / 4;
[0035] Subsequently, two different response maps W are calculated e ∈M×C and W p ∈M×C, used for adjustment and fusion respectively and
[0036]
[0037] in, Shows the characteristics of cumulative aggregation, represents element-wise multiplication;
[0038] Deploy two FGA instances, namely FGA (1) and FGA (2) , using dual MeshConv types with different kernel sizes, including 1×1 convolutional layer and 1×5 convolutional layer, where 1×1 convolutional layer is used for main feature extraction and 1×5 convolutional layer is used for extracting fine-grained features. The outputs of the two levels of FGA are combined and the response map W is generated through splicing, convolution and Softmax layer operations. e and W p .
[0039] Furthermore, step S3 includes the following steps:
[0040] S31, divide the data set into training set, validation set and test set;
[0041] S32, based on the training set, and in combination with the loss function, training optimization is performed on the TSMCN, and then a training model with minimum validation loss is selected for testing and evaluation to screen a liver image segmentation model.
[0042] Further, the loss function in step S32 includes a weighted cross-entropy loss and a Dice loss.
[0043] Further, the liver image segmentation model is screened by calculating the average 3D Chamfer distance on the test set in step S32.
[0044]
[0045] wherein CD ave is the average 3D Chamfer distance, N is the number of samples in the test set, CD is the 3D Chamfer distance, which is used to measure the distance error between the segmented region and the real anatomical region, and v and w represent the vertices of the predicted and real anatomical regions, respectively.
[0046] Compared with the prior art, the present application has the following advantages:
[0047] The present application builds a double-flow grid convolutional neural network TSMCN, which includes two parallel streams of edge processing stream E-stream and point processing stream P-stream, for extracting geometric representations from grid edges and coordinate points, respectively; and then a data set is constructed to train and optimize the TSMCN to obtain a liver image segmentation model. Among them, the two parallel streams of edge processing stream E-stream and point processing stream P-stream extract topological information from the edges of the liver grid unit surface and capture spatial relationships from the vertex coordinates, i.e. E-stream extracts topological information from the edges of the liver grid, and P-stream captures spatial relationships of coordinate points. This design breaks through the limitations of traditional single-stream networks and can learn geometric features from two key perspectives at the same time, providing multi-dimensional information support for accurate segmentation and facilitating accurate and automatic segmentation of key anatomical regions on the liver surface.
[0048] In the double-flow mesh convolutional neural network TSMCN built by the application, the E-stream aims to capture the local topology of triangular patches from all edges in the liver mesh, the input transformer composed of MeshConv is used to transform the input matrix into a new feature space with fixed dimensions M*8, after the input transformer, a mesh channel attention layer MeshSE is introduced to enhance the single-view feature representation ability by learning the channel relationship, then the enhanced features are input into K residual MeshConv modules to stably extract more representative geometric features, the residual MeshConv block stabilizes the extraction of more representative geometric features in the E-stream, which is beneficial to subsequent mesh pooling.In addition, for the improvement of mesh pooling, the complementary features from the P-stream are integrated to construct a queue for mesh pooling, the original edge features are changed according to the compressed mesh structure, which can make the E-stream transition from capturing local information to learning more reliable global topology information, thereby enhancing the overall understanding of the liver mesh.
[0049] In the double-flow mesh convolutional neural network TSMCN built by the application, the P-stream takes the coordinates of each edge center point as input, aiming to capture the spatial relationship complementary to the edge features, the input transformer composed of MeshConv is used to first convert the input into a fixed dimension M*8, and then perform higher-level spatial geometric feature extraction.Subsequently, another MeshSE and K residual MeshConv blocks are used to further enhance the spatial geometry representation in the point features.In addition, the point features are aggregated and pooled while the mesh is reconstructed, after the mesh reconstruction, the point features and the edge features continue to share the same mesh structure, so that complementary single-view geometric representations can be extracted at different levels.
[0050] In the double-flow mesh convolutional neural network TSMCN built by the application, a fine-grained aggregation (FGA) based attention mechanism specially designed for mesh data is designed, and an FGA module is constructed, which first uses a 1*1 convolution to generate a compact feature descriptor, which can reduce the computational demand while preserving significant features, then, two different response maps are calculated to adjust and fuse edge features and point features, realizing cross-view feature aggregation, in addition, two FGA instances are designed and deployed to model the cross-view correspondence, both FGA use double MeshConv types with different kernel sizes, specifically, a 1*1 convolution layer is used for main feature extraction, and a 1*5 convolution layer is used for secondary fine-grained feature extraction.Then, through a series of operations such as connection, convolution and application of softmax layer, the FGA outputs of two levels are integrated into a response map.The FGA-based attention mechanism not only balances the contributions of different view features, but also generates more reliable pooled meshes, preserves the structure critical to the task, and enhances the effectiveness of subsequent single-view feature extraction.
[0051] During model training, the present invention uses a loss function that combines weighted cross entropy and Dice loss. The Dice coefficient can reliably reflect the degree of overlap between the segmented points or edges in the grid and the annotated area. In addition, the training model is evaluated by calculating the average 3D Chamfer distance on the test set (used to measure the distance error between the segmented area and the true anatomical area). This allows the optimal training model to be screened and used as the liver image segmentation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 Schematic diagram of the method flow of the present invention;
[0053] Figure 2 Schematic diagram of the application process of the embodiment;
[0054] Figure 3 Schematic diagram of the TSMCN automatic segmentation process in the embodiment. DETAILED DESCRIPTION
[0055] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0056] Example
[0057] Current automated segmentation methods based on 3D mesh processing struggle to achieve accurate landmark segmentation due to the large shape variations of the liver and limited data. Existing methods based on point or edge features, such as PointNet++ and MeshCNN, train single-stream networks from a single perspective and are unable to adapt to the complexities of the liver. This proposal aims to provide a method for segmenting key anatomical landmarks on the liver surface. This method is used to achieve accurate automatic segmentation of key anatomical regions on the liver surface, facilitating intraoperative navigation and position monitoring, thereby improving the effectiveness of augmented reality (AR)-guided liver surgery.
[0058] like Figure 1 As shown, a liver surface landmark segmentation method based on a two-stream grid convolutional neural network includes the following steps:
[0059] S1. Construct a dataset containing liver mesh data and corresponding annotations;
[0060] S2. Build a two-stream grid convolutional neural network TSMCN, which includes two parallel streams: edge processing stream E-stream and point processing stream P-stream, which are used to extract geometric representations from grid edges and coordinate points respectively;
[0061] S3. Use the dataset to train and optimize TSMCN to obtain a liver image segmentation model;
[0062] S4. Input the current liver grid data into the liver image segmentation model, and output the segmentation results of the corresponding key anatomical regions.
[0063] This embodiment applies the above solution, such as Figure 2 As shown in the figure, it covers multiple key links such as dataset processing, model construction and training, aiming to achieve accurate segmentation of key anatomical areas on the liver surface. The main contents are:
[0064] 1. Data Collection and Processing
[0065] This example collects data from three publicly available liver datasets, including a total of 200 liver mesh data from 3Dircadb (15 samples), LiTS (108 samples), and Amos (77 samples), and performs manual annotation.
[0066] A liver surface model was first extracted using 3D Slicer and processed in MeshLab for manifold simplification, compression, and closure. Key anatomical regions (the falciform ligament and hepatic crest) were annotated in Blender by two computational scientists. Their annotations were cross-validated for consistency and subsequently reviewed and approved by two experienced clinical experts.
[0067] The resulting liver mesh data contains 3,000 to 20,000 edges, which are divided into 100 for training (i.e., training set), 30 for validation (i.e., validation set), and 70 for testing (i.e., test set).
[0068] This embodiment uses the Dice coefficient (%) and 3D Chamfer distance (CD, unit: mm) to evaluate segmentation performance. The Dice coefficient (%) reflects the degree of overlap between the segmented points or edges in the grid and the annotated area, which can directly measure the accuracy of segmentation. The 3D Chamfer distance (CD, unit: mm) is used to measure the distance error between the segmented area and the true anatomical area:
[0069]
[0070] Among them, v and w represent the vertices of the predicted and true anatomical regions, respectively. During the model training and optimization process, the average 3D Chamfer distance (CD ave ) to evaluate the performance of the training model:
[0071]
[0072] Where N is the number of samples in the test set.
[0073] 2. Model Construction
[0074] As Figure 3 shown in Fig. 2, the overall architecture of TSMCN contains two parallel streams, i.e., E-stream and P-stream, which extract geometric representations from mesh edges and vertices, respectively. These complementary features are adaptively fused by a fine-grained aggregation attention mechanism to generate a task-specific pooling mesh that preserves key topological and spatial structures. The fused multi-level features are then passed to the decoder branch to generate the segmentation probability of key anatomical regions. Specifically, it includes:
[0075] Building E-stream: Based on all edges of the liver mesh, MeshCNN defines the edge features of the input according to the topological properties formed by the four adjacent edges (i.e., dihedral angle, internal angle, and edge length ratio), which form two triangles, generating an input matrix of size Mx5 (M is the number of mesh edges). Then, through the input transformer composed of MeshConv, it is transformed into a new feature space with fixed dimensions Mx8.
[0076] After the input transformer, a mesh channel attention layer MeshSE is introduced to enhance the single-view feature representation ability by learning the channel relationship, and the formula is:
[0077]
[0078] wherein are the input and output features, respectively, c represents the number of channels, Sigmoid(·), FC(·), Pool(·) and ⊙ operations represent the application of Sigmoid activation function, full connection layer, global average pooling process and Hadamard product operation, respectively. Then, the enhanced feature is input into K (K=2 in this embodiment) residual MeshConv modules to stably extract more representative geometric features:
[0079]
[0080] Here, the regularization and activation function are omitted, which represents the refined features provided by the MeshSE component. The residual MeshConv block stabilizes the extraction of more representative geometric features in the E-stream so as to be used for subsequent mesh pooling.
[0081] Mesh pooling improvement: MeshCNN defines mesh pooling as a series of "edge collapse" operations, which are prioritized in descending order of edge feature norm, thereby generating a new mesh structure with fewer edges. Traditional single-stream MeshCNN considers discarding unnecessary edge information from a single perspective. However, it is challenging to determine the importance of features solely relying on low-level local topology, especially in landmark segmentation tasks such as the ligamentum teres and liver hills, which differ in spatial location and local topology. Therefore, the present scheme proposes to integrate complementary features from P-stream to construct a queue for mesh pooling, which changes the original edge features according to the compressed mesh structure: traditional MeshCNN based on single perspective for mesh pooling is easy to discard key information. TSMCN integrates complementary features from P-stream to construct a mesh pooling queue, the formula is:
[0082]
[0083] wherein, indicates the feature fusion of E-Stream and P-Stream using FGA attention. Mesh1 represents the new mesh structure obtained by ReMesh. Based on this, Pool e The original edge features are pooled, and the receptive field is expanded before entering the next operation. In this embodiment, TSMCN involves three scale pooling operations, enabling E-stream to transition from capturing local information to learning more reliable global topology information, thereby enhancing the overall understanding of the liver mesh.
[0084] Constructing P-stream: Although E-stream is good at learning discriminative local topology from edge features, it is initially not sensitive enough to the location information that distinguishes individual triangular elements. To complement E-stream, the present scheme further designs P-stream, taking the coordinates of each edge center point as input (size Mx3), aiming to capture spatial relationships complementary to edge features. Similarly, the input transformer composed of MeshConv is used to convert the input to a fixed dimension (Mx8) for subsequent higher-level spatial geometric feature extraction. MeshSE and K (K=2 in this embodiment) residual MeshConv blocks are adopted to further enhance the expression of spatial geometric information in point features.
[0085] Generally, the pooling module of the point cloud aggregates features based on adjacent points. However, due to the disorder of points and edges in the grid structure, the traditional point cloud pooling module may cause spatial misalignment with the edge features, making the feature fusion complex. To this end, in view of the problem that the traditional point cloud pooling module is spatially misaligned with the edge features when processing the grid structure, the scheme innovatively combines grid reconstruction to aggregate and pool point features, that is, to aggregate and pool point features while reconstructing the grid, and the formula is:
[0086]
[0087] where Mesh 1 is the newly reconstructed liver grid, and after this grid reconstruction, the point features and the edge features share the same grid structure, realizing the extraction of complementary single-view geometric representations at different levels.
[0088] In the E-stream and the P-stream, the scheme respectively uses MeshConv, MeshSE and residual MeshConv blocks to perform feature conversion, enhancement and stable extraction. At the same time, in view of the characteristics of the grid structure, the point feature pooling method is improved to cooperate with the edge features, which can effectively improve the feature processing capability.
[0089] Design an attention mechanism based on fine-grained aggregation FGA: in order to effectively process the two different input feature maps from the E-stream and from the P-stream. The scheme is designed for the two different input feature maps from the E-stream and the P-stream. First, 1x1 convolution is used to generate compact feature descriptors D e ∈M×C / 4 and D p ∈M×C / 4 to retain significant features and reduce computational complexity. Then, two response maps W e ∈M×C and W p ∈M×C are calculated to modulate and fuse and cross-view feature aggregation, that is, the implementation of Atten(·) in ReMesh, and the fusion formula is:
[0090]
[0091] where, shows the accumulated aggregated features, denotes element-wise multiplication.
[0092] The FGA module is good at generating response maps that can reflect cross-view relationships at the granularity levels of points (or edges) and triangular elements. The embodiment deploys two FGA instances FGA (1) and FGA(2) , to model these cross-view correspondences. Both FGAs utilize a dual MeshConv type with different kernel sizes (1x1 convolutional layers for main feature extraction and 1x5 convolutional layers for fine-grained feature extraction), and the outputs of the two-level FGAs are combined by concatenation, convolution, and Softmax layer operations to generate the response map W e and W p , thus balancing the contributions of different view features, generating a more reliable pooling grid that preserves task-critical structures and enhances the effectiveness of subsequent single-view feature extraction.
[0093] The above FGA-based feature fusion strategy can fuse the features of the E-stream and the P-stream, and then generate a new grid after pooling based on the new features, avoiding the collapse of the key landmark area where the landmark is located. In addition, based on the new grid, the E-stream and the P-stream continue to process single-view features, so that the feature learning does not interfere with each other.
[0094] III. Model training
[0095] In this embodiment, the model is built using the PyTorch framework and trained on an RTX 8000 GPU. The grid data input to the MeshCNN is standardized to a fixed 20000 edges, and three downsampling steps are set to reduce the number of edges to 3000, 2250 and 1750, respectively. A loss function combining weighted cross-entropy and Dice loss is used, with a background weight of 1 and a landmark weight of 10. The Adam optimizer is used with a learning rate of 0.01, a batch size of 4, and 600 epochs of training. The model with the smallest validation loss is selected for testing, and the average performance of three different random seed runs is taken as the final optimal training model, which is used as the liver image segmentation model. In actual application, only the current liver grid data needs to be input into the liver image segmentation model to obtain the corresponding segmentation result of the key anatomical region of the liver surface.
[0096] In summary, the TSMCN proposed in this scheme contains two parallel streams, E-stream and P-stream, which respectively extract topological information from the edges of the liver grid unit surface and capture spatial relationships from the vertex coordinates, then adaptively fuse single-view features through the attention mechanism based on fine-grained aggregation FGA to generate a pooling grid that preserves key structures, and then realize precise segmentation of key anatomical regions such as the liver ridge and the falciform ligament. The TSMCN automatic segmentation can reduce the manual workload of preoperative planning and improve the consistency of anatomical landmark segmentation, providing strong support for AR-assisted liver surgery, especially in falciform ligament segmentation, which is crucial for the alignment of 3D preoperative grid and intraoperative AR navigation. Compared with existing segmentation methods, the scheme has the following obvious advantages:
[0097] 1. Improved segmentation accuracy: TSMCN can effectively learn and integrate complementary information between spatial relationships and local topology. Compared with existing methods, it has higher accuracy in segmenting key anatomical structures on the liver surface, providing more accurate information for surgical planning.
[0098] 2. Enhanced robustness: Through the dual-flow structure and FGA attention mechanism, it has stronger adaptability to different liver meshes and can maintain stable segmentation performance under changes in liver shape and appearance.
[0099] 3. Promote the development of AR guided surgery and improve surgical efficiency: Achieving accurate segmentation of key anatomical areas on the preoperative liver surface effectively improves 3D-2D registration efficiency, reduces the time required for manual labeling, and improves the accuracy and stability of AR navigation, thereby contributing to improved surgical efficiency.
Claims
1. A liver surface landmark segmentation method based on a two-stream grid convolutional neural network, characterized in that: The following steps are involved: S1. Construct a dataset containing liver mesh data and corresponding annotations; S2. Build a two-stream grid convolutional neural network TSMCN, which includes two parallel streams: edge processing stream E-stream and point processing stream P-stream, which are used to extract geometric representations from grid edges and coordinate points respectively; S3. Use the dataset to train and optimize TSMCN to obtain a liver image segmentation model; S4. Input the current liver grid data into the liver image segmentation model, and output the segmentation results of the corresponding key anatomical regions.
2. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 1, characterized in that: The step S1 specifically obtains a plurality of liver grid data from a public liver dataset and labels them accordingly.
3. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 1, characterized in that: The step S1 specifically involves marking key anatomical areas of the liver in the liver mesh data.
4. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 1, characterized in that: The step S2 comprises the following steps: S21, constructing an edge processing stream E-stream to capture the local topology of triangular patches from all edges in the liver mesh and output an edge feature map; Construct a point processing stream P-stream, which takes the coordinates of each edge center point as input to capture the spatial relationship complementary to the edge features and outputs a point feature map; S22. Based on the attention mechanism of fine-grained aggregation FGA, an FGA module is constructed to adjust and fuse edge feature maps and point feature maps, realize cross-view feature aggregation, and output the accumulated aggregated features; S23. Construct a decoder branch to decode the accumulated aggregated features and generate segmentation probabilities of key anatomical regions.
5. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 4, characterized in that: The working process of the edge processing stream E-stream includes: based on all edges of the liver mesh, defining the input edge features according to the topological attributes, and generating an input matrix of size M×5, where M is the number of mesh edges; Using the input transformer composed of MeshConv, the input matrix is transformed into a new feature space of fixed dimension M×8; The MeshSE layer is introduced to enhance the expressive power of single-view features by learning channel relationships: in, For the enhanced features, are the input and output features, c represents the number of channels, Sigmoid(·), FC(·), Pool(·) and ⊙ operations represent the application of Sigmoid activation function, fully connected layer, global average pooling process and Hadamard product operation respectively; The enhanced features Input into K residual MeshConv modules to extract more representative geometric features: In addition, the point features output by the point processing stream P-stream are combined to construct a queue for grid pooling. The original edge features are changed according to the compressed grid structure, and pooling operations at three scales are performed, allowing E-stream to transition from capturing local information to learning more reliable global topology information.
6. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 5, characterized in that: The working process of the point processing stream P-stream includes: taking the coordinates of each edge center point as input, the size is M×3, and using the input transformer composed of MeshConv to convert the input into a fixed dimension of M×8; MeshSE and K residual MeshConv blocks are used to enhance the expression of spatial geometric information in point features, and mesh reconstruction is combined to aggregate and pool point features: Among them, Mesh 1 The liver mesh is reconstructed by ReMesh. After reconstruction, the point features and edge features Continue to share the same mesh structure Mesh 1 , thereby extracting complementary single-view geometric representations at different levels.
7. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 6, characterized in that: The working process of the FGA module includes: first, using 1×1 convolution to generate a compact feature descriptor D e ∈M×C / 4 and D p ∈M×C / 4; Subsequently, two different response maps W are calculated e ∈M×C and W p ∈M×C, used for adjustment and fusion respectively and in, Shows the characteristics of cumulative aggregation, represents element-wise multiplication; Deploy two FGA instances, namely FGA (1) and FGA (2) , using dual MeshConv types with different kernel sizes, including 1×1 convolutional layer and 1×5 convolutional layer, where 1×1 convolutional layer is used for main feature extraction and 1×5 convolutional layer is used for extracting fine-grained features. The outputs of the two levels of FGA are combined and the response map W is generated through splicing, convolution and Softmax layer operations. e and W p .
8. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 1, characterized in that: The step S3 comprises the following steps: S31, divide the data set into training set, validation set and test set; S32. Based on the training set and combined with the loss function, the TSMCN training is optimized, and then the training model with the smallest verification loss is selected for testing and evaluation to screen out the liver image segmentation model.
9. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 8, characterized in that: The loss function in step S32 includes weighted cross entropy loss and Dice loss.
10. The method for liver surface landmark segmentation based on a dual-stream grid convolutional neural network according to claim 8, characterized in that: The step S32 specifically screens and obtains the liver image segmentation model by calculating the average 3D Chamfer distance on the test set: Among them, CD ave is the average 3D Chamfer distance, N is the number of samples in the test set, CD is the 3D Chamfer distance, which is used to measure the distance error between the segmented region and the true anatomical region, and v and w represent the vertices of the predicted and true anatomical regions, respectively.