A couinaud liver segment segmentation method and system based on vascular topology guidance
By constructing a vascular topology map and combining it with a graph convolutional neural network, the problems of insufficient accuracy and generalization ability of Couinaud liver segment segmentation in the existing technology are solved, and high-precision liver segment segmentation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU PUJIAN MEDICAL TECH CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies for CT image-based Couinaud liver segmentation struggle to accurately distinguish anatomically defined liver segment boundaries by vascular pathways, leading to segmentation results that do not match clinical reality. Furthermore, the model exhibits weak generalization ability and significant fluctuations in segmentation accuracy.
By acquiring and preprocessing 3D CT image data, liver blood vessels are segmented, a vascular topology map is constructed, topological features are extracted using a graph convolutional neural network, and multi-center training is performed using a cross-attention mechanism and a global memory bank to optimize the model and improve segmentation accuracy and generalization ability.
This method enables the direct derivation of liver segmentation results conforming to the Couinaud segmentation standard from the spatial distribution of blood vessels, thereby improving the anatomical accuracy of the segmentation and the generalization ability of the model.
Smart Images

Figure CN121685559B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical technology, specifically a Couinaud liver segmentation method and system based on vascular topology guidance. Background Technology
[0002] In the field of automated liver segmentation based on CT images, existing technologies can achieve relatively accurate extraction of the overall liver contour. However, Couinaud segmentation, which is crucial for precise surgical planning in liver surgery, still faces significant challenges. Traditional methods mainly rely on the density or texture information of the liver parenchyma, making it difficult to distinguish the anatomically defined boundaries of liver segments by vascular pathways, resulting in segmentation results that are seriously inconsistent with clinical reality. Existing methods that incorporate vascular information are mostly limited to simple distance mapping or rule-based post-processing, failing to fully explore the complex three-dimensional topology of the vascular system and its profound anatomical relationship with liver segmentation. Therefore, the models have weak generalization ability, and the segmentation accuracy fluctuates greatly among different individuals or under different imaging conditions. Summary of the Invention
[0003] The purpose of this invention is to provide a Couinaud liver segmentation method and system based on vascular topology guidance to overcome the shortcomings of the prior art. It can directly derive results that conform to the Couinaud segmentation standard from the spatial distribution of blood vessels, thereby improving the anatomical accuracy of segmentation and the generalization ability of the model.
[0004] One embodiment of this application provides a Couinaud liver segmentation method based on vascular topology guidance, the method comprising:
[0005] Acquire and preprocess 3D CT image data, and segment liver blood vessels in the images;
[0006] Based on the segmentation results of the liver blood vessels, the topological structure, key points and geometric features of the liver blood vessels are extracted to construct a blood vessel topology map to characterize the spatial distribution and hierarchical relationship of blood vessels.
[0007] The vascular topology map is input into a graph convolutional neural network for high-dimensional structured encoding to extract the topological features of the blood vessels. The topological features are then fused with the image features extracted by the liver segmentation backbone network using a cross-attention mechanism.
[0008] By combining the structural contrast loss function and the global memory, the fused network is trained and optimized at multiple centers to improve the model's generalization ability and boundary segmentation accuracy. The final output is a multi-class liver segmentation result that conforms to the Couinaud segmentation anatomy rules.
[0009] Another embodiment of this application provides a Couinaud liver segmentation system based on vascular topology guidance, the system comprising:
[0010] The acquisition module is used to acquire and preprocess 3D CT image data and segment liver blood vessels in the images;
[0011] The extraction module is used to extract the topological structure, key points and geometric features of the liver blood vessels based on the segmentation results of the liver blood vessels, so as to construct a blood vessel topology map to characterize the spatial distribution and hierarchical relationship of blood vessels.
[0012] The fusion module is used to input the vascular topology map into the graph convolutional neural network for high-dimensional structured encoding, extract the topological features of the blood vessels, and use a cross-attention mechanism to fuse the topological features with the image features extracted by the liver segmentation backbone network.
[0013] The segmentation module is used to perform multi-center training and optimization of the fused network by combining the structural contrast loss function and the global memory, so as to improve the model's generalization ability and boundary segmentation accuracy, and finally output multi-class liver segmentation results that conform to the Couinaud segmentation anatomy rules.
[0014] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.
[0015] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.
[0016] Compared with existing technologies, this invention provides a Couinaud liver segmentation method based on vascular topology guidance. It acquires and preprocesses 3D CT image data and segments the liver vessels in the images. Based on the segmentation results, a vascular topology map is constructed to characterize the spatial distribution and hierarchical relationship of the vessels. The vascular topology map is input into a graph convolutional neural network for high-dimensional structured encoding to extract the topological features of the vessels. This topological feature is then fused with the image features extracted by the liver segmentation backbone network using a cross-attention mechanism. The fused network is then optimized through multi-center training using a structure contrast loss function and a global memory database. Finally, it outputs multi-class liver segmentation results that conform to the Couinaud segmentation anatomical rules, thereby enabling the direct derivation of results conforming to the Couinaud segmentation standard from the spatial distribution of blood vessels, improving the anatomical accuracy of the segmentation and the generalization ability of the model. Attached Figure Description
[0017] Figure 1 A hardware structure block diagram of a computer terminal for a Couinaud liver segmentation method based on vascular topology guidance provided in an embodiment of the present invention;
[0018] Figure 2 A schematic flowchart of a Couinaud liver segmentation method based on vascular topology guidance provided in an embodiment of the present invention;
[0019] Figure 3 This is a schematic diagram of a Couinaud liver segmentation system based on vascular topology guidance, provided as an embodiment of the present invention. Detailed Implementation
[0020] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0021] This invention first provides a Couinaud liver segmentation method based on vascular topology guidance, which can be applied to electronic devices, such as computer terminals, specifically ordinary computers.
[0022] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for a Couinaud liver segmentation method based on vascular topology guidance, provided as an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0023] The non-volatile storage medium can store an operating system and a computer program. This computer program includes program instructions that, when executed, cause the processor to perform any Couinaud liver segmentation method guided by vascular topology.
[0024] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0025] The internal memory provides an environment for the execution of computer programs in non-volatile storage media. When executed by the processor, the computer program can enable the processor to execute any Couinaud liver segmentation method based on vascular topology guidance.
[0026] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0027] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0028] See Figure 2 The present invention provides a Couinaud liver segmentation method based on vascular topology guidance, which may include the following steps:
[0029] S201: Acquire and preprocess 3D CT image data, and segment liver blood vessels in the images;
[0030] Specifically, it can acquire 3D CT image data containing the liver and generate raw DICOM format image sequences;
[0031] This step forms the data source foundation for liver segmentation. Its core is acquiring CT image sequences in a clinically standardized format to ensure compatibility and data integrity for subsequent processing. The specific implementation method is as follows:
[0032] The acquired 3D CT image data came from routine clinical liver enhancement scans, covering the entire liver region (from the top of the diaphragm to the lower edge of the liver). The scanning equipment was a commonly used clinical spiral CT scanner, and the scanning mode was dual-phase enhancement scanning (arterial phase and portal venous phase) to ensure clear differentiation of the liver parenchyma, portal vein, and hepatic venous system. The image data was stored in the DICOM 3.0 standard format, with each DICOM file corresponding to one CT slice. The number of slices in the sequence depends on the scan range, usually 80-150 slices (e.g., 100 slices when the vertical diameter of the liver region is 10cm and the slice thickness is 1mm).
[0033] The original DICOM format image sequence contains rich metadata and pixel data. The metadata is stored in the DICOM header file. The core fields include basic patient information (anonymized patient ID, gender, age), scanning parameters (tube voltage 120kV, tube current 250mA, slice thickness 1mm, slice spacing 1mm, matrix 512×512, pixel spacing 0.9766mm×0.9766mm), image enhancement phase (arterial phase marker "Phase:Arterial", portal venous phase marker "Phase:PortalVenous"), window width and window level (soft tissue window: window width 400HU, window level 50HU), etc. These parameters directly determine the image resolution and tissue contrast and are important bases for subsequent preprocessing.
[0034] Pixel data is in 16-bit integer format, ranging from -1024 HU (air) to +3071 HU (bone). The HU value of liver parenchyma in the portal venous phase ranges from 40 to 80 HU, while the HU value of the portal vein and hepatic veins, due to the presence of contrast agent, ranges from 100 to 150 HU. This grayscale difference forms the basis for subsequent vascular segmentation. During acquisition, the system receives image data in batches through the DICOM service interface of the hospital's PACS system (supporting DICCOM C-MOVE / C-STORE protocols), automatically verifying file integrity (checking slice sequence continuity and metadata consistency). If a slice is missing or the file is corrupted, a retransmission request is immediately triggered to ensure that the generated original DICOM image sequence is complete and error-free, and can be directly used for subsequent preprocessing.
[0035] The original DICOM image sequence is preprocessed, including resampling the anisotropic original voxels to isotropic resolution, standardizing the image intensity using the Z-Score method, and applying a nonlocal mean filtering algorithm to reduce image noise, thereby generating preprocessed standardized 3D CT image volume data.
[0036] This step eliminates differences in image acquisition and reduces noise interference through standardization, providing high-quality input for the segmentation network. The specific implementation method is as follows:
[0037] The core of anisotropic voxel resampling is to convert the non-uniform voxels of the original CT image into isotropic voxels (i.e., voxel side lengths are consistent in the x, y, and z axes). The in-plane voxel resolution of the original CT image is typically 0.9766mm × 0.9766mm (matrix 512 × 512, scan field 500mm), and the z-axis slice thickness may be 1mm (isotropic) or 5mm (anisotropic). For anisotropic images with a slice thickness of 5mm, resampling to an isotropic resolution of 1mm × 1mm × 1mm is required. Resampling employs a trilinear interpolation algorithm. This algorithm achieves a smooth transition by calculating the weighted average gray value of the eight adjacent original voxels surrounding the target voxel. The weights are inversely proportional to the distance from the target voxel to each original voxel, ensuring that the spatial structure of the resampled image is not distorted. For example, a slice with an original slice thickness of 5mm will generate five consecutive slices in the z-axis direction after resampling, with a smooth transition in gray values between adjacent slices, avoiding step artifacts.
[0038] Z-Score image intensity standardization is used to eliminate grayscale differences caused by different scanning equipment and parameters, making the HU value distribution of the liver and blood vessels more uniform. The calculation formula is: I'=(I-μ) / σ, where I is the original pixel HU value, μ is the average HU value of the liver region, σ is the standard deviation of the HU value of the liver region, and I' is the standardized pixel value. When calculating μ and σ, the liver region is first initially segmented by a simple thresholding method (HU value -20 to 200) to eliminate interference from background, skeleton, etc. Then, the mean and standard deviation of the HU value of the region are statistically analyzed. For example, in a CT sequence, the liver region has μ=60HU and σ=25HU. If the original HU value of a pixel is 85, then after standardization, I'=(85-60) / 25=1.0; if the HU value of another pixel is 35, then I'=(35-60) / 25=-1.0. After standardization, the pixel values are concentrated in the range of [-3,3], which preserves the grayscale differences of the tissue while eliminating global intensity shift.
[0039] Nonlocal mean filtering is used to reduce image noise (such as scanning quantum noise) while preserving key structures such as liver edges and vascular details. The core of this algorithm is to find similar pixel blocks in the image and then perform weighted averaging to reduce noise in the current pixel. The filtering parameters are set as follows: search window size 21×21×21 voxels (corresponding to a 21mm×21mm×21mm range), similarity window size 7×7×7 voxels (corresponding to a 7mm×7mm×7mm range), and smoothing coefficient h=0.1 (controlling the filtering strength; a smaller h indicates stronger filtering). For example, for noisy pixels at the liver edge, the algorithm finds multiple liver parenchyma pixel blocks with similar grayscale distributions within the search window, reduces noise through weighted averaging, and avoids blurring edge details due to the small similarity window. The signal-to-noise ratio of the filtered image is improved by more than 30%, and the contrast between blood vessels and liver parenchyma is clearer.
[0040] The preprocessed standardized 3D CT image volume data is in 3D tensor format with dimensions [H,W,D] (e.g., 512×512×100), pixel value range [-3,3], and voxel resolution of 1mm×1mm×1mm. This preserves the spatial structure and tissue contrast of the original image while eliminating equipment differences and noise interference, fully meeting the input requirements of the subsequent segmentation network.
[0041] Standardized 3D CT image volume data is input into a pre-trained liver and blood vessel segmentation network. This network adopts the 3DU-Net++ architecture and generates liver region masks and blood vessel system segmentation masks by segmenting the liver parenchyma on arterial phase images and segmenting the portal vein and hepatic vein system in the liver on portal venous phase images.
[0042] This step is crucial for obtaining the basic segmentation results of the liver and blood vessels. Through targeted dual-phase image segmentation, binary masks for the liver parenchyma and vascular system are obtained separately. The specific implementation method is as follows:
[0043] The pre-trained liver and blood vessel segmentation network is based on an improved 3DU-Net++ architecture. Compared to the traditional 3DU-Net, it adds dense skip connections and a deep supervision mechanism, which can better capture multi-scale features and improve the segmentation accuracy of small blood vessels and liver edges. The network input is standardized 3D CT image volume data (arterial phase or portal venous phase), and the output is a binary segmentation mask (pixel value 1 represents the target area, and 0 represents the background). The arterial phase image is used to segment the liver parenchyma (the liver parenchyma is uniformly enhanced in the arterial phase and has a clear boundary with the surrounding tissue), and the portal venous phase image is used to segment the portal vein and hepatic venous system (the blood vessels in the portal venous phase are significantly different in grayscale from the liver parenchyma due to the contrast agent filling).
[0044] The training process for the network for arterial phase liver parenchyma segmentation: 1000 clinical arterial phase CT images were used as the training set. Image annotation was completed jointly by two senior radiologists, and the annotation range covered the entire liver parenchyma (including liver segments but excluding blood vessels). The network training parameters were: batch size 2, initial learning rate 0.001 (decreasing to 0.9 times the original value every 500 training epochs), loss function using Dice loss + cross-entropy loss (weight ratio 1:1), optimizer AdamW, and training iterations for 2000 epochs. The Dice coefficient for liver segmentation on the test set reached over 0.97. After inputting standardized arterial phase images, the network extracts multi-scale features through downsampling by the encoder (4 convolutional blocks, each containing two 3×3×3 convolutional layers, a BatchNorm layer, and a ReLU activation function). The decoder upsamples through transposed convolutions and fuses the deep and shallow features of the encoder with dense skip connections. Finally, it outputs a liver region mask. In the mask, the pixel value of the liver parenchyma is 1, and the pixel value of the background (such as the stomach, intestines, and muscles) is 0. For example, a three-dimensional mask has dimensions of 512×512×100, with 1.2 million pixels having a value of 1, corresponding to the liver parenchyma region.
[0045] The training process for the portal venous phase vascular system segmentation network: 1000 clinical portal venous phase CT images were used as the training set. The annotation range included the main portal vein and its tertiary branches, and the main hepatic vein and its major branches (left hepatic vein, middle hepatic vein, and right hepatic vein), also annotated by two radiologists. The network architecture was consistent with the liver segmentation network, and the training parameters were adjusted as follows: the loss function adopted was Focal loss + Dice loss (weight ratio 1:1) to address the class imbalance problem caused by the low proportion of vascular pixels (approximately 5%). The training iterations were 2000 times, and the Dice coefficient for vascular segmentation on the test set reached over 0.92. After inputting standardized portal venous phase images, the network captured the tubular structure of blood vessels through multi-scale features, distinguishing between the portal vein (densely branched peripherally) and the hepatic veins (converging towards the center), and outputting a vascular system segmentation mask. In the mask, the pixel value of blood vessels was 1, while the pixel value of liver parenchyma and background was 0. For example, 300,000 pixels in the mask had a value of 1, corresponding to all annotated branches of the portal vein and hepatic veins.
[0046] The generated liver region mask and vascular system segmentation mask are both three-dimensional binary tensors with the same dimension as the input image volume data. The vascular mask will be further optimized through post-processing to lay the foundation for topology extraction.
[0047] The vascular system segmentation mask is post-processed, and the main vascular structure is preserved by three-dimensional connected domain analysis. The centerline skeleton of the blood vessels is extracted by morphological thinning algorithm to generate the final liver blood vessel segmentation result and skeleton line.
[0048] This step refines the vessel segmentation mask through post-processing, removes noise interference, and extracts the core skeleton structure, providing a precise structural foundation for vessel topology map construction. The specific implementation method is as follows:
[0049] The core of 3D connected component analysis is to remove small noise fragments from the vascular mask while retaining the main vessels and branches with anatomical significance. A connected component is defined as a set of pixels that are adjacent to each other (6-neighborhood: top / bottom, left / right, front / back) and have a pixel value of 1. During the analysis, all pixels of the vascular mask are traversed first, all connected components are marked, and the volume of each connected component is calculated (number of pixels × 1 mm³). A volume threshold of 100 mm³ (corresponding to 100 voxels) is set, retaining connected components with a volume ≥ 100 mm³ and removing small fragments with a volume < 100 mm³ (mostly scanning noise or vascular artifacts). For example, if a vascular mask contains 20 connected components, of which 15 have a volume ≥ 100 mm³ (corresponding to the main portal vein, main hepatic vein, and major branches), and 5 have a volume < 100 mm³ (noise fragments), after connected component analysis, only 15 effective connected components are retained, generating an optimized vascular mask to ensure that subsequent topology analysis is not interfered with by noise.
[0050] The morphological thinning algorithm employs the Zhang-Suen fast parallel thinning algorithm to extract a single-pixel-width centerline skeleton from the optimized vascular mask. This algorithm iteratively removes pixels at the vessel edges while retaining centerline pixels. The iteration terminates when no pixels meet the deletion criteria or when the number of iterations reaches 10 (sufficient for complete thinning). The deletion criteria are: the pixel is an edge pixel (the number of pixels with a value of 0 in its 6-neighborhood is ≥1), the pixel is a non-endpoint pixel (the number of pixels with a value of 1 in its 6-neighborhood is ≥2 and ≤6), and the deletion does not disrupt the connectivity of the vessel. These conditions ensure that the thinned centerline is continuous, unbroken, and 1 pixel wide. For example, the mask width of the portal vein trunk is 8 pixels. After thinning, a continuous centerline is generated with a coordinate sequence of [(x1,y1,z1),(x2,y2,z2),...,(xn,yn,zn)], and the spacing between adjacent coordinates is 1 mm, accurately reflecting the direction and bifurcation structure of the vessel.
[0051] The final liver vessel segmentation result consists of two parts: an optimized binary mask of the vascular system (preserving the main trunk and effective branches) and a centerline skeleton (coordinate sequence format) with a width of one pixel. The vessel mask is used for subsequent keypoint extraction, and the centerline skeleton is used to construct the edge structure of the vascular topology map. Together, they constitute the basic data for vascular topology analysis, ensuring that the subsequently constructed topology map can accurately reflect the spatial distribution and hierarchical relationship of the vessels.
[0052] S202, Based on the segmentation results of the liver blood vessels, extract the topological structure, key points and geometric features of the liver blood vessels to construct a blood vessel topology map to characterize the spatial distribution and hierarchical relationship of blood vessels;
[0053] Specifically, based on the liver blood vessel segmentation results and skeleton lines, a 3D corner detection algorithm can be used to identify the bifurcation points and endpoints of the blood vessel skeleton, define these points as nodes of the topology graph, and generate a set of key points of the blood vessels.
[0054] This step is the core starting point for constructing the vascular topology map. By accurately identifying the key nodes of the vascular skeleton, it provides core elements for subsequent topological relationship construction. The specific implementation method is as follows:
[0055] The liver vessel segmentation result is a binary mask (vessel region is 1, background is 0), and the skeleton line is a continuous curve with a single pixel width, accurately reflecting the spatial orientation and branching structure of the vessels. The Harris3D corner detection algorithm is used for 3D corner detection. This algorithm identifies points in the vascular skeleton where curvature changes abruptly or branches intersect by calculating the autocorrelation matrix eigenvalues of local voxel regions. Its core principle is that corner regions show significant grayscale changes in multiple directions, and both eigenvalues of the autocorrelation matrix are large, while edge regions have only one large eigenvalue, and flat regions have both small eigenvalues.
[0056] The algorithm parameters are set as follows: the local neighborhood window size is 3×3×3 voxels (corresponding to a range of 3mm×3mm×3mm) to capture local structural changes in blood vessels; the response function threshold is set to 0.2, and points above the threshold are identified as candidate corner points; the non-maximum suppression window size is 5×5×5 voxels to remove redundant points among the candidate corner points and ensure uniform node distribution. During the recognition process, a bifurcation point is defined as a node in the vascular skeleton with ≥3 connected branches (i.e., a pixel with a degree ≥3 in the skeleton line), such as the junction of the left and right branches of the portal vein trunk, which is connected to 3 branches and has a degree of 3, and is identified as a bifurcation point; an endpoint is defined as a node with 1 connected branch (a pixel with a degree of 1), such as the terminal point of the hepatic vein, which is connected to only 1 branch and is identified as an endpoint.
[0057] To ensure accurate identification, candidate nodes need to be verified: A bifurcation point must have at least three non-collinear skeletal segments in its neighborhood, with an angle ≥30° between the segments (to avoid misidentifying slightly curved vessel points as bifurcation points); an endpoint must have only one skeletal segment in its neighborhood, and there are no other skeletal points on the extension of that segment and the endpoint (to avoid misidentifying the midpoint of a short branch as an endpoint). For example, if a candidate point has three skeletal segments in its neighborhood with angles of 45°, 60°, and 75°, satisfying the bifurcation point condition, it is identified as the bifurcation point of the left branch of the portal vein; another candidate point is connected to only one segment, and its extension has no other skeletal points, thus it is identified as the endpoint of the right hepatic vein. The final set of vascular key points contains all verified bifurcation points and endpoints, categorized and labeled by vascular type, such as "Bifurcation point-PV-001 (portal vein bifurcation), Endpoint-HV-002 (left hepatic vein terminal)". The size of the set depends on the complexity of the vascularity and usually contains 50-200 key points.
[0058] Calculate the geometric features of each key point, including its coordinates in three-dimensional space, the estimated radius of the blood vessel it is located in, and the local blood vessel orientation vector at that point, and generate a key point feature descriptor;
[0059] This step quantifies the geometric properties of key points to endow the nodes of the topological graph with rich feature information, supporting subsequent graph convolutional coding and feature fusion. The specific implementation method is as follows:
[0060] The 3D spatial coordinates of each keypoint are directly taken from the pixel coordinates of the vascular skeleton line. Since the skeleton line is one pixel wide, the coordinate accuracy is consistent with the normalized voxel size (1 mm), and the format is (x, y, z) in mm. For example, the coordinates of the bifurcation point of the portal vein are (125.3, 98.7, 105.2), accurately reflecting its spatial position in the 3D CT image. The coordinate values are retained to one decimal place to ensure the accuracy of subsequent distance calculations and topology construction.
[0061] The radius of the blood vessel is estimated based on a binary mask, using a local region statistical method: a spherical neighborhood (corresponding to 5 voxels) with a radius of 5 mm is constructed centered on the key point. The number of pixels in the blood vessel mask (voxel value = 1) within the neighborhood is counted. The radius is estimated by the relationship between the volume of the sphere and the cross-sectional area of the blood vessel, using the following formula: Where N is the number of blood vessel voxels in the neighborhood, and h is the length of the neighborhood along the blood vessel axis (taken as 5 mm). For example, if there are 78 blood vessel voxels in the spherical neighborhood of a key point, substituting into the formula yields... This means the vessel radius at that point is approximately 2.2 mm, consistent with the anatomical dimensions of portal vein branches (portal vein branch radius is typically 1-3 mm). The radius estimate is rounded to one decimal place to distinguish between main trunk vessels (radius ≥ 3 mm) and branch vessels (radius < 3 mm).
[0062] Local vascular orientation vectors are used to characterize the extension direction of blood vessels at key points. For bifurcation points, the orientation vector for each branch needs to be calculated, while for endpoints, only the orientation vector for a single branch is calculated. The calculation method is as follows: For an endpoint, select its five adjacent skeletal points and obtain the vascular axis through linear fitting. The direction furthest from the endpoint is taken as the orientation vector, with a unit vector format of (dx, dy, dz) and a magnitude of 1. For bifurcation points, for each branch, select five skeletal points on that branch furthest from the bifurcation center and obtain the orientation vector of that branch through linear fitting. For example, the orientation vector of a branch at a portal vein bifurcation point might be (0.3, 0.4, -0.2), indicating that the branch extends along the positive x-axis, positive y-axis, and negative z-axis. Each component of the orientation vector is retained to three decimal places to ensure accurate directional representation.
[0063] The keypoint feature descriptor is a high-dimensional vector that integrates the above three types of features, with 10 dimensions: 3D coordinates + 1D radius estimate + 6D orientation vector (for bifurcation points, the orientation vectors of the first two main branches are taken, each with 3 dimensions; for endpoints, a single orientation vector is taken, and the remaining 3 dimensions are padded with 0). For example, the feature descriptor of a certain bifurcation point is (125.3,98.7,105.2,2.2,0.3,0.4,-0.2,-0.1,0.5,0.3), which fully quantifies the spatial location, blood vessel size, and extension direction of the node, providing a foundation for subsequent topological relationship construction and feature encoding.
[0064] Based on the Euclidean distance between key points, the K-nearest neighbor algorithm is applied to establish the connection relationship between each key point and its K nearest neighbors, thereby constructing an initial adjacency matrix representing the connectivity of blood vessels.
[0065] This step transforms discrete key points into a network structure with preliminary connectivity by establishing initial connections between them, providing an initial framework for subsequent correction. The specific implementation method is as follows:
[0066] Euclidean distance is used to quantify the spatial straight-line distance between two key points. The calculation formula is as follows: The unit is mm. This distance directly reflects the spatial proximity of key points and is the core basis for establishing connection relationships. For example, the Euclidean distance between key point A (125.3, 98.7, 105.2) and key point B (128.5, 100.2, 103.1). This indicates that the two points are spatially adjacent.
[0067] In the K-Nearest Neighbors (KNN) algorithm, the K value is determined to be 5 through cross-validation. This means that each keypoint establishes initial connections only with its five spatially nearest keypoints. This K value ensures that each node has enough connections to reflect the topological structure while avoiding excessive redundant connections that would complicate subsequent corrections. The algorithm execution process is as follows: iterates through each keypoint, calculates its Euclidean distance to all other keypoints in the set, sorts them by distance in ascending order, selects the top five keypoints as its neighbors, and establishes bidirectional connections (i.e., if A is a neighbor of B, then B is also a neighbor of A).
[0068] The initial adjacency matrix is an N×N square matrix (N is the size of the keypoint set, for example, 100 keypoints correspond to a 100×100 matrix). The matrix element A[i][j] takes the value 1 or 0. A[i][j]=1 indicates that keypoint i and keypoint j have an initial connection relationship, and A[i][j]=0 indicates that there is no connection relationship. The matrix is a symmetric matrix (because the connection relationship is bidirectional), and the diagonal element A[i][i] is set to 0 (nodes themselves have no connection). For example, in the adjacency matrix of 100 keypoints, A[0][1]=1 indicates that keypoint 0 is adjacent to keypoint 1, and A[0][6]=0 indicates that keypoint 0 is not adjacent to keypoint 6. This matrix intuitively represents the initial connectivity between keypoints. For example, the neighboring points of the bifurcation point (keypoint 0) of the portal vein include keypoints 1, 2, and 3 on its branch. In the adjacency matrix, A[0][1], A[0][2], and A[0][3] are all 1, which accurately reflects the branch connection relationship.
[0069] By combining prior knowledge of vascular anatomy, the initial adjacency matrix is corrected to ensure that the connection relationship conforms to the tree-like hierarchical structure of portal veins converging from coarse to fine and hepatic veins converging from fine to coarse, and finally a complete vascular topology map containing node features and adjacency relationships is generated.
[0070] This step corrects for deviations in the initial connections using anatomical principles, ensuring that the topological relationships conform to clinical anatomy. It is crucial for constructing an effective vascular topology map, and the specific implementation method is as follows:
[0071] The core of prior anatomical knowledge of blood vessels lies in the hierarchical tree-like structure of two types of vessels: the portal venous system follows a branching pattern from "thick to thin," meaning that starting from the main portal vein (radius ≥ 5 mm), it branches successively into secondary branches (radius 3-5 mm) and tertiary branches (radius 1-3 mm), with no direct cross-level connections between branches; the hepatic venous system follows a convergence pattern from "thin to thick," meaning that terminal branches (radius 1-2 mm) gradually converge into secondary branches (radius 2-4 mm) and the main vein (radius ≥ 4 mm), with no retrograde connections during the convergence process. Furthermore, there is no direct connection between the portal vein and the hepatic venous system (anatomically, they are indirectly connected through hepatic sinusoids, and are considered two independent subgraphs in the topological diagram).
[0072] The correction process consists of three steps: First, the key point sets of the portal vein and hepatic vein are separated. By clustering based on radius features and direction vectors (the portal vein branches are dispersed, while the hepatic vein converges), the key points are divided into two subsets, ensuring that all connections between the two types of nodes in the adjacency matrix are set to 0 (removing cross-system connections). Second, the key points in each subset are classified according to their radius (portal vein: trunk level, secondary branch level, tertiary branch level; hepatic vein: terminal level, secondary branch level, trunk level), and connections across two or more levels are deleted. For example, the initial connection (A[i][j]=1) between a trunk level node (radius 5mm) and a tertiary branch level node (radius 1.5mm) of the portal vein needs to be set to 0, and only connections between adjacent levels are retained. Finally, reasonable missing connections are added. For key points on the same branch that are close to each other (≤5mm) but not selected by the K nearest neighbors, connections are manually added (A[i][j]=1). For example, key point 4 and key point 5 on a branch of the portal vein are 4.8mm apart. Since K=5 was not selected, a connection is added during correction.
[0073] The corrected adjacency matrix strictly follows the anatomical hierarchy. For example, in the portal vein subgraph, trunk-level nodes are only connected to second-level branch-level nodes, and second-level branch-level nodes are only connected to trunk-level and third-level branch-level nodes, with no cross-level or reverse connections. In the hepatic vein subgraph, terminal-level nodes are only connected to second-level branch-level nodes, and second-level branch-level nodes are only connected to terminal-level and trunk-level nodes, which conforms to the convergence rule.
[0074] The final generated vascular topology map is a binary structure of "node features + adjacency relationships": node features are 10-dimensional feature descriptors for each key point, including spatial coordinates, radius, and direction vector; adjacency relationships are corrected N×N adjacency matrices, accurately reflecting the tree-like topology and hierarchical distribution of blood vessels. For example, the portal vein topology submap contains 30 key points, and the adjacency matrix retains only the connections between adjacent levels, forming a complete tree structure from the trunk to the third-level branches; the hepatic vein topology submap contains 25 key points, forming a convergence structure from the terminal to the trunk. Together, they constitute a complete vascular topology map, providing accurate structured input for subsequent graph convolutional coding.
[0075] S203, the blood vessel topology map is input into a graph convolutional neural network for high-dimensional structured encoding, the topological features of the blood vessels are extracted, and the topological features are fused with the image features extracted by the liver segmentation backbone network using a cross-attention mechanism;
[0076] Specifically, the vascular topology map can be input into a multi-layer graph convolutional neural network. Each layer updates the node representation by aggregating the features of the node itself and its neighboring nodes. After multiple layers of propagation, the output is a high-dimensional topological feature vector that can characterize the global vascular hierarchy and spatial relationship.
[0077] This step uses a graph convolutional neural network (GCN) to deeply encode the structured information of the vascular topology map, transforming discrete node features and adjacency relationships into globally unified high-dimensional feature vectors. This provides structured vascular information support for subsequent fusion. The specific implementation method is as follows:
[0078] The vascular topology map contains N nodes (N being the number of key points in the blood vessels, e.g., 100) and an N×N adjacency matrix. The initial features of each node are 10-dimensional key point feature descriptors (3D coordinates + radius + 6D orientation vector). The multilayer graph convolutional neural network adopts a 3-layer stacked structure, with each layer following the core process of "adjacency matrix aggregation + linear transformation + activation function". By aggregating neighborhood information layer by layer, it achieves dimensionality enhancement and structured encoding of node features.
[0079] The input to the first layer of the GCN is the initial node feature matrix (dimension N×10) and the normalized adjacency matrix (dimension N×N). The adjacency matrix is normalized using a symmetric normalization method (A_norm=D^(-1 / 2)×A×D^(-1 / 2), where D is the degree matrix of the adjacency matrix A), in order to avoid feature aggregation bias caused by differences in node degree. The aggregation process is as follows: the normalized adjacency matrix is multiplied by the node feature matrix to obtain the neighborhood feature aggregation result (dimension N×10), and then added to the node's own feature matrix to achieve the fusion of "own features + neighborhood features"; then, the fused features are upscaled to 64 dimensions through a linear transformation layer (weight matrix dimension 10×64), and finally, nonlinearity is introduced through the ReLU activation function to output the first layer node feature matrix (dimension N×64). For example, if the initial features of a node are [125.3,98.7,105.2,2.2,0.3,0.4,-0.2,-0.1,0.5,0.3], after neighborhood aggregation (merging the features of 3 neighboring nodes) and linear transformation, a 64-dimensional feature vector is output, where the first 5 elements are [0.82,0.35,-0.12,0.56,0.71].
[0080] The second-layer GCN takes the N×64 feature matrix output from the first layer and the same normalized adjacency matrix as input. The aggregation logic is the same as the first layer, but the linear transformation layer weight matrix has a dimension of 64×128, and the activation function remains ReLU. The output is an N×128 node feature matrix. This layer further aggregates second-order neighborhood information (the neighborhood of a node), strengthening the representation of hierarchical relationships within blood vessels. For example, after encoding the portal vein trunk node in the second layer, its feature vector contains indirect information about second-level branch nodes, reflecting the hierarchical relationship between the trunk and branches.
[0081] The third layer, GCN, is a global aggregation layer. The input is the N×128 feature matrix output from the second layer. First, global average pooling aggregates the features of N nodes into a 1×128 global feature vector (formula: G=(1 / N)×ΣX_i, where X_i is the 128-dimensional feature of the i-th node). Then, a linear transformation layer (128×256) is used to increase the dimension to 256. Finally, the vector is normalized using the Sigmoid activation function, outputting a high-dimensional topological feature vector of 256. This vector integrates the spatial location, radius, orientation, and hierarchical relationships of all vascular nodes. For example, the output vector is [0.62, 0.45, ..., 0.38] (256 elements, values ranging from 0 to 1), accurately representing the tree-like topological structure of the portal vein ("coarse to fine") and the hepatic vein ("fine to coarse").
[0082] Standardized 3D CT image volume data is input into a Transformer-based liver segmentation backbone network. This network extracts global contextual features of the image through a multi-head self-attention mechanism and outputs multi-scale deep feature maps of the image.
[0083] This step leverages the global modeling capabilities of Transformer to extract global contextual features such as liver parenchyma and perivascular texture from CT images, generating multi-scale feature maps. This provides image-level semantic support for fusion with vascular topological features. The specific implementation is as follows:
[0084] The standardized 3D CT image volume data has a size of 512×512×100 (H×W×D) and a voxel resolution of 1mm×1mm×1mm. Before being input into the backbone network, it needs to be reshaped into a two-dimensional sequence (the 3D image is sliced along the Z-axis, each slice is a 512×512 two-dimensional image, the sequence length is 100, and each element is a 512×512×1 two-dimensional feature). Then, through the embedding layer (convolution kernel size 7×7, stride 2, number of output channels 64), each two-dimensional slice is mapped to a 512×512×64 feature map, finally forming an input feature sequence with a sequence length of 100 and each element being 512×512×64.
[0085] The Transformer-based backbone network consists of six encoder layers, each containing two core modules: a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism has eight heads, each with a feature dimension of 64, for a total feature dimension of 8 × 64 = 512. Its core function is to capture global dependencies between different slices and between different regions within the same slice. The calculation process is as follows: the input features are mapped to three matrices: query (Q), key (K), and value (V) (each with a dimension of 100 × (512 × 512) × 512). The similarity between Q and K is calculated using dot product operations through the eight heads, and the attention weights are obtained after Softmax normalization. These are then weighted and summed with V to output the globally associated features. For example, for the vascular orientation region across liver slices, the multi-head self-attention mechanism assigns a high weight (0.8-0.9), strengthening the feature representation of this region and capturing the continuous distribution information of blood vessels in three-dimensional space.
[0086] The feedforward neural network comprises two layers of linear transformation (intermediate dimension 2048, output dimension 512) and a ReLU activation function, used to perform non-linear transformations on the features of the multi-head self-attention output, enhancing feature representation capabilities. The output of each encoder layer is stabilized and trained through residual connections and layer normalization (LN). The parameters of layer normalization are 512 feature dimensions, and the normalized mean and variance are dynamically calculated based on each sample.
[0087] After processing by a 6-layer encoder, a multi-scale feature extraction module generates deep feature maps at three scales: Scale 1 (fine-grained) is 256×256×128 (downsampled by convolution with a stride of 2, increasing the number of channels from 64 to 128), corresponding to local liver details (such as blood vessel wall texture and liver segment boundary details); Scale 2 (medium-grained) is 128×128×256 (further downsampled, doubling the number of channels), corresponding to medium-scale liver structural features (such as the overall morphology of liver segments); Scale 3 (coarse-grained) is 64×64×512 (downsampled to 1 / 8 size, further doubling the number of channels), corresponding to global liver context features (such as the relative position of the liver to surrounding organs). All three scales of feature maps retain three-dimensional spatial information, providing multi-dimensional image semantic support for subsequent cross-attention fusion.
[0088] The cross-attention fusion module is designed, which uses the topological feature vector output by the graph convolutional neural network as the query vector, linearly projects the image feature map output by the segmentation backbone network into key vector and value vector, obtains the attention weight by calculating the similarity between the query and the key, and performs a weighted summation of the value vector to generate enhanced image features guided by vascular topology information.
[0089] This step is the core of fusing vascular topology information with image semantic features. By using a cross-attention mechanism, image features are focused on key regions related to vascular topology, improving the boundary accuracy of liver segment segmentation. The specific implementation method is as follows:
[0090] The input to the cross-attention fusion module consists of two parts: a 256-dimensional high-dimensional topological feature vector (query vector Q) output by graph convolution, and three scale image feature maps output by the backbone network (which need to be fused separately; here, we will explain in detail the scale 2 feature map as an example). First, the 128×128×256 image feature map at scale 2 is linearly projected, and the number of channels is reduced from 256 to 256 (consistent with the dimension of the query vector) by using 1×1×1 convolution kernels (number 256), resulting in a key vector K (dimension 128×128×256) and a value vector V (dimension 128×128×256). The role of the 1×1×1 convolution is to adjust the channel dimension without changing the spatial size, ensuring that it matches the dimension of the query vector.
[0091] The query vector Q (256 dimensions) needs to be dimensionally expanded, reshaped from 1×256 to 1×1×256, and then expanded to 128×128×256 through a broadcast mechanism to maintain consistency with the spatial size of the key vector K. Similarity calculation uses a dot product approach; for each spatial location (x,y), the dot product of Q(x,y,:) and K(x,y,:) is calculated (the formula is Sim(x,y)=Q(x,y,:)·K(x,y,:)^T / ),in This is a scaling factor to prevent the dot product from becoming too large, which could cause the Softmax gradient to vanish. For example, in a spatial location corresponding to the bifurcation region of the portal vein, the dot product of Q and K is 180, which becomes 180 / 16 = 11.25 after scaling. However, the dot product of the background region may only be 30, which becomes 1.875 after scaling, significantly distinguishing the key region from the background.
[0092] Attention weights are obtained by normalizing the similarity of each spatial location using the Softmax function (formula: Weight(x,y)=exp(Sim(x,y)) / Σexp(Sim(x,y))). The weight values range from 0 to 1, with high weights (0.7-1.0) corresponding to key regions related to vascular topology (such as vascular bifurcation and liver segment boundaries), and low weights (0-0.3) corresponding to irrelevant background regions. For example, the spatial location corresponding to the main portal vein has a weight of 0.85, the internal region of the liver segment has a weight of 0.4, and the background region has a weight of 0.05.
[0093] Finally, the value vector V is weighted and summed using the formula Enhanced_Feat(x,y,:)=ΣWeight(x,y)×V(x,y,:), generating an enhanced image feature of 128×128×256. This feature incorporates structured information from the vascular topology, enhancing image features in vascular-related regions and suppressing them in background regions. The same process is used to fuse image feature maps at scales 1 and 3, resulting in three scales of enhanced feature maps (Scale 1: 256×256×128, Scale 2: 128×128×256, Scale 3: 64×64×512), providing multi-scale, topology-guided enhanced feature support for the subsequent decoder.
[0094] The enhanced image features are input into the segmentation decoder, and the spatial resolution is restored through layer-by-layer upsampling and skip connections. Finally, a preliminary probability map of liver segments incorporating vascular topology information is output.
[0095] This step restores the enhanced feature map to the original image resolution using a decoder, supplements detailed information with skip connections, and outputs the probability distribution of each liver segment, laying the foundation for subsequent loss calculation and segmentation result output. The specific implementation method is as follows:
[0096] The segmentation decoder employs a symmetrical encoder-decoder structure of "upsampling + skip connections + convolutional fusion," containing three upsampling modules corresponding to enhanced feature maps at three different scales. The final output is a liver segment probability map with the same size as the original image. The decoder's input is a 64×64×512 enhanced feature map at scale 3. The first upsampling module uses transposed convolution (kernel size 3×3×3, stride 2, output channels 256) to upsample the 64×64×512 feature map to 128×128×256. This is then concatenated with the 128×128×256 enhanced feature map at scale 2 through skip connections (the number of channels becomes 256+256=512). Finally, a 1×1×1 convolution (output channels 256) is used to fuse the features, resulting in a 128×128×256 fused feature map. The transposed convolution restores the spatial dimensions through inverse convolution, while the skip connections supplement the details lost during upsampling (such as liver segment boundary texture).
[0097] The second upsampling module also uses a 3×3×3 transposed convolution (stride 2, output channels 128) to upsample the 128×128×256 fused feature map to 256×256×128. This is then concatenated with the 256×256×128 enhanced feature map at scale 1 (channels 128+128=256), and fused using a 1×1×1 convolution (output channels 128) to obtain a 256×256×128 feature map. The fused feature map at this layer retains both global topological guidance information and restores the detailed features of liver segments at medium scales, such as clearly distinguishing the boundary regions between liver segments and blood vessels.
[0098] The third upsampling module is the final resolution restoration layer. It uses a 3×3×3 transposed convolution (stride 2, output channels 8) to upsample the 256×256×128 feature map to 512×512×8, at which point the number of channels is 8, corresponding to the probability distribution of the eight Couinaud segments. After the transposed convolution, the Softmax function is used to normalize the 8 channels at each spatial location, so that the sum of the 8 probability values at each location is 1. The higher the probability value, the higher the probability that the location belongs to the corresponding liver segment. For example, if a spatial location corresponds to liver segment III, the probability value of its 3rd channel is 0.92, and the probability values of the other channels are all less than 0.1, indicating that the location is highly likely to belong to liver segment III. At the boundary of a liver segment, two channels may have relatively high probability values (such as the probabilities of liver segments II and III being 0.45 and 0.48, respectively), reflecting the transition characteristics of the boundary.
[0099] The final output of the preliminary liver segment probability map is 512×512×8 (consistent with the planar size of the original standardized CT image, and all slices need to be processed synchronously in the Z-axis direction). The probability map of each slice integrates vascular topology information and image semantic features, providing the core output for subsequent training and optimization by combining the structural contrast loss function, ensuring that the liver segment segmentation results can follow the anatomical rules guided by vascular topology.
[0100] S204 combines the structural contrast loss function and the global memory to perform multi-center training and optimization on the fused network, so as to improve the model's generalization ability and boundary segmentation accuracy, and finally output multi-class liver segmentation results that conform to the Couinaud segmentation anatomy rules.
[0101] Specifically, a multi-center training dataset can be constructed, which contains labeled CT images from different hospitals and different scanning devices. A global memory can be initialized to store representative liver segment feature prototypes from each data center.
[0102] This step is fundamental to improving the model's multi-center generalization ability. By integrating diverse data and an initial feature prototype library, it provides a unified feature reference benchmark for subsequent comparative training. The specific implementation method is as follows:
[0103] The construction of the multi-center training dataset focuses on data diversity and annotation standardization, covering 5 hospitals of different levels (3 tertiary hospitals and 2 secondary hospitals). Each hospital provided 200 clinical liver-enhanced CT images, totaling 1000 data points. The dataset covers different scanning equipment (e.g., spiral CT, multi-slice CT), scanning parameters (tube voltage 100-140kV, tube current 200-300mA), and enhancement protocols (differences in arterial and portal venous phase scan times), ensuring that the data distribution reflects the diversity of real clinical scenarios. All images were jointly annotated by two senior radiologists, with annotations including Couinaud eight-segment semantic segmentation labels and portal and hepatic vein vascular masks. Annotation consistency was verified using the Dice similarity coefficient, ensuring that the Dice coefficient for all samples was ≥0.95, avoiding annotation errors from affecting training results. The dataset was divided into a 7:2:1 ratio: training set (700 cases), validation set (200 cases), and test set (100 cases). The training set was used for model parameter optimization, the validation set for hyperparameter tuning, and the test set for evaluating generalization ability.
[0104] The core of the global memory initialization is to store representative feature prototypes for each liver segment category, providing a reference for subsequent comparison loss calculation. The memory adopts a key-value pair storage structure, where the key is the liver segment category identifier (1-8 correspond to the eight Couinaud segments), and the value is the feature prototype vector of that category. The vector dimension is consistent with the region feature dimension output by the liver segmentation backbone network (256 dimensions). The initialization process is as follows: 20 samples are randomly selected from each liver segment category in the training set, and the liver segment region features of each sample are extracted (obtained from the preliminary probability map through global average pooling). The mean of the 20 feature vectors for each category is calculated as the initial feature prototype for that category. For example, the initial feature prototype for liver segment III is the mean vector of the region features of the 20 samples, with a dimension of 256, and some elements are [0.62, 0.35, -0.18, 0.51, ..., 0.43], with a value range of [-1, 1]. The storage capacity of the memory bank is set to 8 (corresponding to 8 liver segment categories). Only one feature prototype is retained for each category. The quality of the prototype is then optimized through a dynamic update mechanism to ensure that it can represent the common features of the data from each center.
[0105] During the forward propagation of the liver segmentation backbone network training, the basic segmentation loss between the preliminary liver segment probability map and the true label is calculated. At the same time, the regional features of each liver segment category are extracted from the preliminary probability map and compared with the corresponding feature prototypes in the global memory.
[0106] This step achieves "segmentation accuracy optimization" and "feature consistency constraint" through dual-task parallelism, ensuring both basic segmentation performance and providing high-quality features for contrastive loss. The specific implementation method is as follows:
[0107] The basic segmentation loss is calculated using a weighted combination of "Dice loss + cross-entropy loss," balancing the overlap of segmented regions with pixel-level classification accuracy, with a weight ratio of 1:1. Dice loss measures the overlap between the initial liver segment probability map and the true label; the core formula is 2 × (predicted ∩ true) / (predicted + true), with a value range of [0,1]. A loss value closer to 0 indicates a higher overlap. Cross-entropy loss penalizes pixel-level classification errors, calculating the average logarithmic loss between the predicted probability and the true label for each pixel. For example, in the initial liver segment probability map of a sample, the overlap between the predicted region of liver segment V and the true label is 0.85, and the Dice loss is 0.15; the pixel-level classification error rate is 5%, the cross-entropy loss is 0.08, and the basic segmentation loss is (0.15 + 0.08) / 2 = 0.115. During loss calculation, only pixels within the liver region are considered, ignoring the background region to avoid interference from background pixels.
[0108] Region feature extraction is performed independently for each liver segment category. Pixel regions with a probability value ≥ 0.5 for that category are selected from the preliminary liver segment probability map (considered as the predicted region for that liver segment). Global average pooling is then performed on the enhanced image features (256-dimensional feature map output by the decoder) corresponding to this region to obtain the 256-dimensional region feature vector for that liver segment. For example, the predicted region for liver segment IV contains 12,000 pixels. The 256-dimensional features of these pixels are averaged to obtain the region feature vector [0.58, 0.29, -0.21, ..., 0.39]. After extraction, the eight region feature vectors are compared with the corresponding feature prototypes in the global memory according to the liver segment category. The similarity is calculated using cosine similarity, with the formula Sim(a,b)=(a·b) / (||a||×||b||), ranging from [-1, 1]. The closer the value is to 1, the more similar the features are. For example, the cosine similarity between the regional features of liver segment IV and the prototype in the memory bank is 0.82, indicating that the features of this sample have a high degree of consistency with the prototype; if the similarity is less than 0.5, it suggests that the features of this sample may deviate from the common features and need to be constrained by contrastive loss.
[0109] By using the structural contrast loss function, we maximize the feature similarity of the same liver segment category among different samples and minimize the feature similarity between different liver segment categories, thereby narrowing the intra-class distance and widening the inter-class distance, thus generating the structural contrast loss.
[0110] This step is crucial for improving feature consistency and inter-class discriminative power. It involves learning stable liver segment features by comparing loss-constrained models. The specific implementation method is as follows:
[0111] The core logic of the Structural Contrast Loss (SCL) function is to construct "positive sample pairs" and "negative sample pairs" to enhance intra-class clustering and inter-class separation through loss calculation. A positive sample pair is defined as "a combination of regional features and prototypes from the memory database for different samples within the same liver segment category." For example, the liver segment III feature of sample A and its prototype from the memory database, as well as the liver segment III feature of sample B and its prototype from the memory database, are both positive sample pairs. A negative sample pair is defined as "a combination of regional features and prototypes from the memory database for different liver segment categories." For example, the liver segment III feature of sample A and its prototype from the memory database, as well as the liver segment V feature of sample B and its prototype from the memory database, are both negative sample pairs. Each liver segment feature of each sample corresponds to one positive sample pair and seven negative sample pairs (covering the other seven liver segment categories), ensuring comprehensive comparison.
[0112] In the loss calculation process, a temperature coefficient τ (with a value of 0.1, used to adjust the smoothness of the similarity distribution) is introduced. First, the similarity score of the positive sample pair is calculated and divided by τ to obtain the positive sample score; then, the similarity score of all negative sample pairs is calculated and divided by τ to obtain the negative sample score; finally, all scores are normalized by the Softmax function, and the log loss of the positive sample score is calculated as the structural contrast loss. For example, the similarity between the regional features of liver segment III of a certain sample and the positive sample of the prototype is 0.82, which, after dividing by τ=0.1, yields 8.2; the similarities of the 7 negative sample pairs are 0.35, 0.28, 0.41, 0.32, 0.29, 0.37, and 0.33, respectively, which, after dividing by τ, yield 3.5, 2.8, 4.1, 3.2, 2.9, 3.7, and 3.3; after Softmax normalization, the proportion of positive sample scores is exp(8.2) / (exp(8.2)+sum(exp(negative sample scores)))=2678 / (2678+33+16+60+24+18+41+28)≈0.93, and the structural contrast loss is -log(0.93)≈0.073. The smaller this loss value, the higher the intra-class similarity, the lower the inter-class similarity, and the better the feature discrimination.
[0113] To avoid individual outliers affecting loss calculation, a hard negative sample mining mechanism is introduced. Only the three most similar samples in a negative sample pair are selected for loss calculation, enhancing the constraint of the loss on difficult-to-distinguish samples. For example, in the seven negative sample pairs mentioned above, the three with the highest similarity are 4.1, 3.7, and 3.5. Only these three are retained for calculation, making the loss more focused on inter-class boundary samples and improving the model's discriminative ability.
[0114] The basic segmentation loss and structural contrast loss are weighted and combined to form the total training loss. The network parameters are optimized through the backpropagation algorithm, and the feature prototypes in the global memory are dynamically updated. After multiple rounds of iterative training, the liver segmentation backbone network model obtains strong generalization ability, performs direct inference on unseen test data, and outputs high-precision multi-class liver segmentation results that conform to the Couinaud eight-segment anatomical rules.
[0115] This step achieves simultaneous improvement in model accuracy and generalization ability through multi-loss joint optimization and dynamic updating of the memory bank, ultimately outputting clinically usable segmentation results. The specific implementation method is as follows:
[0116] The weighted combination of the total training loss adopts an adaptive weighting strategy. The weight α of the basic segmentation loss is set to 0.7, and the weight β of the structural contrastive loss is set to 0.3. The total loss L = α × L_base + β × L_SCL. This weight ratio is determined through validation set tuning to ensure both the accuracy of basic segmentation and to fully leverage the constraint effect of the contrastive loss on feature consistency. For example, in a certain iteration, L_base = 0.115, L_SCL = 0.073, and the total loss L = 0.7 × 0.115 + 0.3 × 0.073 = 0.0805 + 0.0219 = 0.1024.
[0117] The backpropagation algorithm employs the AdamW optimizer, with an initial learning rate of 0.001 and a weight decay coefficient of 0.0001. Every 100 training epochs, the learning rate decays to 0.9 times its original value to avoid oscillations in the later stages of training. During optimization, the total loss is backpropagated to each layer of the network using a chain rule, updating the parameters of the graph convolutional neural network, the Transformer backbone network, and the cross-attention fusion module, allowing the network to gradually learn the dual characteristics of "high segmentation accuracy + strong feature consistency." The total number of training iterations is set to 2000 epochs, employing an early stopping strategy: if the total loss of the validation set does not decrease for 50 consecutive epochs, training is stopped, and the current optimal model parameters are saved to avoid overfitting.
[0118] The global feature memory's dynamic update strategy is "periodic update + quality screening." It is updated every 20 training epochs. During the update, regional features for each liver segment category are extracted from the current training batch, and their similarity to existing prototypes in the feature memory is calculated. If the similarity between the new feature and the prototype is ≥0.85, the prototype is retained. If multiple new features have a similarity <0.7, the mean of these new features is calculated, and the original prototype is replaced, ensuring that the prototypes in the feature memory always represent the common features of the latest training data. For example, after 20 training epochs, if the average similarity between the prototype of liver segment II in the feature memory and the features of 10 samples in the current batch is 0.68 <0.7, the mean vector of these 10 features is calculated, and the original prototype is replaced, making the prototype more consistent with the current training trend.
[0119] After 2000 rounds of iterative training, the model's performance on the test set reached clinically usable standards: the Dice coefficients for all eight liver segments were ≥0.92, with the Dice coefficients for segments adjacent to the portal vein (such as segments IV and V) ≥0.94; the segmentation boundary accuracy for vascular bifurcation regions was ≥95%; and the generalization error for multi-center data was ≤5%, significantly outperforming traditional models. When inferring on unseen test data, the model directly inputs standardized 3D CT image volume data, and through topological feature encoding, cross-attention fusion, and decoder output, generates a 512×512×100 multi-class liver segmentation result. Each pixel is labeled with the corresponding liver segment category (1-8), and background pixels are labeled 0. The segmentation results strictly follow the Couinaud eight-segment anatomical rules, the boundaries of liver segments corresponding to the portal vein branches from thick to thin are clear, and there are no cross-segment misclassifications in the liver segmentation of the hepatic vein confluence area, which can be directly used for clinical surgical planning.
[0120] As can be seen, the process involves acquiring and preprocessing 3D CT image data, segmenting liver vessels in the images, constructing a vascular topology map to characterize the spatial distribution and hierarchical relationship of blood vessels based on the segmentation results, inputting the vascular topology map into a graph convolutional neural network for high-dimensional structured encoding, extracting the topological features of blood vessels, and fusing these topological features with the image features extracted by the liver segmentation backbone network using a cross-attention mechanism. The fused network is then optimized through multi-center training using a structure contrast loss function and a global memory, ultimately outputting multi-class liver segmentation results that conform to the Couinaud segmentation anatomical rules. This enables the direct derivation of results conforming to the Couinaud segmentation standard from the spatial distribution of blood vessels, improving the anatomical accuracy of segmentation and the model's generalization ability.
[0121] Another embodiment of the present invention provides a Couinaud liver segmentation system based on vascular topology guidance, see [link to relevant documentation]. Figure 3 The system may include:
[0122] The acquisition module 301 is used to acquire and preprocess three-dimensional CT image data and segment the liver blood vessels in the images;
[0123] The extraction module 302 is used to extract the topological structure, key points and geometric features of the liver blood vessels based on the segmentation results of the liver blood vessels, so as to construct a blood vessel topology map to characterize the spatial distribution and hierarchical relationship of blood vessels.
[0124] The fusion module 303 is used to input the blood vessel topology map into the graph convolutional neural network for high-dimensional structured encoding, extract the topological features of the blood vessels, and use a cross-attention mechanism to fuse the topological features with the image features extracted by the liver segmentation backbone network.
[0125] The segmentation module 304 is used to perform multi-center training and optimization of the fused network by combining the structural contrast loss function and the global memory library, so as to improve the generalization ability and boundary segmentation accuracy of the model, and finally output multi-class liver segmentation results that conform to the Couinaud segmentation anatomy rules.
[0126] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.
[0127] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0128] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.
[0129] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.
Claims
1. A Couinaud liver segmentation method based on vascular topology guidance, characterized in that, The method includes: Acquire and preprocess 3D CT image data, segment the liver vessels in the images, and generate the final liver vessel segmentation result and centerline skeleton; Based on the segmentation results of the liver vessels, the topological structure, key points, and geometric features of the liver vessels are extracted to construct a vascular topology map representing the spatial distribution and hierarchical relationship of the vessels. Specifically, based on the liver vessel segmentation results and the central axis skeleton, a 3D corner detection algorithm is used to identify the bifurcation points and endpoints of the central axis skeleton, defining these points as nodes in the topology map and generating a set of vascular key points. The geometric features of each key point are calculated, including its coordinates in 3D space, the estimated radius of the vessel it belongs to, and the local vessel orientation vector at that key point, generating a key point feature descriptor. Based on the Euclidean distance between key points, the K-nearest neighbor algorithm is applied to establish connections between each key point and its K nearest neighbors, thus constructing an initial adjacency matrix representing vascular connectivity. The initial adjacency matrix is corrected using prior anatomical knowledge of the vessels to ensure that the connections conform to a tree-like hierarchical structure where portal veins converge from thick to thin and hepatic veins converge from thin to thick, ultimately generating a complete vascular topology map containing node features and adjacency relationships. The vascular topology map is input into a graph convolutional neural network for high-dimensional structured encoding to extract the topological features of the blood vessels. The topological features are then fused with the image features extracted by the liver segmentation backbone network using a cross-attention mechanism. Finally, a preliminary liver segment probability map that incorporates vascular topology information is output through a segmentation decoder. By combining a structural contrast loss function with a global memory, the fused network is optimized through multi-center training to improve model generalization ability and boundary segmentation accuracy, ultimately outputting multi-class liver segmentation results that conform to Couinaud's segmented anatomy rules. Specifically, a multi-center training dataset is constructed, containing labeled CT images from different hospitals and scanning devices. A global memory is initialized to store representative liver segment feature prototypes from each data center. During the forward propagation of the liver segmentation backbone network training, the basic segmentation loss between the preliminary liver segment probability map and the ground truth labels is calculated. Simultaneously, regional features for each liver segment category are extracted from the preliminary probability map and compared with the global memory. The system compares the corresponding feature prototypes in the memory; it uses the structural contrast loss function to maximize the feature similarity of the same liver segment category among different samples and minimize the feature similarity between different liver segment categories, thereby narrowing the intra-class distance and widening the inter-class distance, thus generating the structural contrast loss; it combines the basic segmentation loss and the structural contrast loss in a weighted manner to form the total training loss, optimizes the network parameters through the backpropagation algorithm, and dynamically updates the feature prototypes in the global memory. After multiple rounds of iterative training, the liver segmentation backbone network model obtains strong generalization ability, performs direct inference on unseen test data, and outputs high-precision multi-class liver segmentation results that conform to the Couinaud eight-segment anatomical rules.
2. The method according to claim 1, characterized in that, The acquisition and preprocessing of 3D CT image data, and the segmentation of liver blood vessels in the images, include: Acquire 3D CT image data containing the liver and generate raw DICOM format image sequences; The original DICOM image sequence is preprocessed, including resampling the anisotropic original voxels to isotropic resolution, standardizing the image intensity using the Z-Score method, and applying a nonlocal mean filtering algorithm to reduce image noise, thereby generating preprocessed standardized 3D CT image volume data. Standardized 3D CT image volume data is input into a pre-trained liver and blood vessel segmentation network. This network adopts the 3DU-Net++ architecture and generates liver region masks and blood vessel system segmentation masks by segmenting the liver parenchyma on arterial phase images and segmenting the portal vein and hepatic vein system in the liver on portal venous phase images. The vascular system segmentation mask is post-processed, and the main vascular structure is preserved by three-dimensional connected domain analysis. The centerline skeleton of the blood vessels is extracted by morphological thinning algorithm to generate the final liver blood vessel segmentation result and centerline skeleton.
3. The method according to claim 2, characterized in that, The process of inputting the vascular topology map into a convolutional neural network for high-dimensional structured encoding, extracting the topological features of the blood vessels, and fusing these topological features with image features extracted by the liver segmentation backbone network using a cross-attention mechanism includes: The vascular topology map is input into a multi-layer graph convolutional neural network. Each layer updates the node representation by aggregating the features of the node itself and its neighboring nodes. After multiple layers of propagation, the output is a high-dimensional topological feature vector that can characterize the global vascular hierarchy and spatial relationship. Standardized 3D CT image volume data is input into a Transformer-based liver segmentation backbone network. This network extracts global contextual features of the image through a multi-head self-attention mechanism and outputs multi-scale deep feature maps of the image. The cross-attention fusion module is designed, which uses the topological feature vector output by the graph convolutional neural network as the query vector, linearly projects the image feature map output by the segmentation backbone network into key vector and value vector, obtains the attention weight by calculating the similarity between the query and the key, and performs a weighted summation of the value vector to generate enhanced image features guided by vascular topology information. The enhanced image features are input into the segmentation decoder, and the spatial resolution is restored through layer-by-layer upsampling and skip connections. Finally, a preliminary probability map of liver segments incorporating vascular topology information is output.
4. A Couinaud liver segmentation system based on vascular topology guidance, characterized in that, The system includes: The acquisition module is used to acquire and preprocess 3D CT image data, segment the liver vessels in the images, and generate the final liver vessel segmentation result and centerline skeleton. An extraction module is used to extract the topological structure, key points, and geometric features of the liver vessels based on the segmentation results, in order to construct a vascular topology map representing the spatial distribution and hierarchical relationship of the vessels. Specifically, based on the liver vessel segmentation results and the central axis skeleton, a 3D corner detection algorithm is used to identify the bifurcation points and endpoints of the central axis skeleton, defining these points as nodes in the topology map and generating a set of vascular key points. The geometric features of each key point are calculated, including its coordinates in 3D space, the estimated radius of the vessel it belongs to, and the local vessel orientation vector at that key point, generating a key point feature descriptor. Based on the Euclidean distance between key points, the K-nearest neighbor algorithm is applied to establish connections between each key point and its K nearest neighbors, thus constructing an initial adjacency matrix representing vascular connectivity. The initial adjacency matrix is corrected using prior anatomical knowledge of the vessels to ensure that the connections conform to a tree-like hierarchical structure where portal veins converge from thick to thin and hepatic veins converge from thin to thick, ultimately generating a complete vascular topology map containing node features and adjacency relationships. The fusion module is used to input the vascular topology map into the graph convolutional neural network for high-dimensional structured encoding, extract the topological features of the blood vessels, and use the cross-attention mechanism to fuse the topological features with the image features extracted by the liver segmentation backbone network. Finally, the segmentation decoder outputs a preliminary liver segment probability map that incorporates the vascular topology information. The segmentation module combines a structural contrast loss function with a global memory to perform multi-center training and optimization of the fused network, thereby improving the model's generalization ability and boundary segmentation accuracy. The final output is a multi-class liver segmentation result conforming to Couinaud's segmented anatomical rules. Specifically, a multi-center training dataset is constructed, containing labeled CT images from different hospitals and scanning devices. A global memory is initialized to store representative liver segment feature prototypes from each data center. During the forward propagation of the liver segmentation backbone network training, the basic segmentation loss between the preliminary liver segment probability map and the ground truth labels is calculated. Simultaneously, regional features for each liver segment category are extracted from the preliminary probability map. The model compares the feature prototypes with the corresponding features in the global memory. Using the structural contrast loss function, it maximizes the feature similarity of the same liver segment category across different samples and minimizes the feature similarity between different liver segment categories, thereby narrowing the intra-class distance and widening the inter-class distance, generating the structural contrast loss. The basic segmentation loss and the structural contrast loss are weighted and combined to form the total training loss. The network parameters are optimized using the backpropagation algorithm, and the feature prototypes in the global memory are dynamically updated. After multiple rounds of iterative training, the liver segmentation backbone network model acquires strong generalization ability, directly inferring from unseen test data and outputting high-precision, multi-class liver segmentation results that conform to the Couinaud eight-segment anatomical rules.
5. The system according to claim 4, characterized in that, The acquisition module is specifically used for: Acquire 3D CT image data containing the liver and generate raw DICOM format image sequences; The original DICOM image sequence is preprocessed, including resampling the anisotropic original voxels to isotropic resolution, standardizing the image intensity using the Z-Score method, and applying a nonlocal mean filtering algorithm to reduce image noise, thereby generating preprocessed standardized 3D CT image volume data. Standardized 3D CT image volume data is input into a pre-trained liver and blood vessel segmentation network. This network adopts the 3DU-Net++ architecture and generates liver region masks and blood vessel system segmentation masks by segmenting the liver parenchyma on arterial phase images and segmenting the portal vein and hepatic vein system in the liver on portal venous phase images. The vascular system segmentation mask is post-processed, and the main vascular structure is preserved by three-dimensional connected domain analysis. The centerline skeleton of the blood vessels is extracted by morphological thinning algorithm to generate the final liver blood vessel segmentation result and centerline skeleton.
6. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-3 when it is run.
7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-3.
Citation Information
Patent Citations
Modeling method for blood vessel segmentation in medical image based on topological knowledge
CN115908297A
Liver enhanced CT feature extraction and fusion method based on vascular topological structure
CN119904724A