A liver cancer thermal ablation path planning system based on deep reinforcement learning
By using a path planning system based on deep reinforcement learning, combined with multimodal imaging and expert reward functions, the thermal ablation path for liver cancer is optimized, solving the problem of inaccurate path planning in existing technologies and achieving safe and efficient ablation under individualized and dynamic constraints.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- THE THIRD MEDICAL CENT OF THE CHINESE PEOPLES LIBERATION ARMY GENERAL HOSPITAL
- Filing Date
- 2026-04-01
- Publication Date
- 2026-06-26
AI Technical Summary
Existing liver cancer thermal ablation pathway planning technology suffers from several problems: reliance on clinician experience, lack of multimodal image fusion, difficulty in adapting to individualized cases, and inability to accurately predict the extent of thermal damage and respiratory motion interference. These issues lead to inaccurate pathway planning and a higher risk of complications.
A path planning system based on deep reinforcement learning is adopted. By generating target and hazard area masks, constructing pipeline distance fields, predicting thermal fields, predicting breathing deformation, and fusing multi-scale features, combined with an expert reward function, the system optimizes path decision-making and achieves individualized and dynamically constrained path planning.
It improves the precision and safety of thermal ablation surgery for liver cancer, reduces damage to normal tissues and complications, and provides intelligent interventional treatment decision support.
Smart Images

Figure CN121962543B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ablation path planning technology, specifically to a liver cancer thermal ablation path planning system based on deep reinforcement learning. Background Technology
[0002] Hepatocellular carcinoma (HCC) thermal ablation, a core minimally invasive interventional technique for treating HCC, uses an ablation needle to release energy into the lesion area to induce coagulative necrosis. Its therapeutic efficacy highly depends on the optimized design of the ablation needle puncture path, requiring simultaneous fulfillment of three core requirements: avoidance of dangerous areas (such as blood vessels and bile ducts), complete lesion coverage, and minimization of damage to normal liver tissue. Existing path planning technologies have significant limitations: traditional methods heavily rely on the clinical physician's experience and judgment, exhibiting strong subjectivity and individual differences, making it difficult to accurately compensate for dynamic deformation of liver tissue caused by respiratory movements, and easily leading to serious complications such as hemorrhage and bile leakage due to path deviations; some intelligent planning methods are mostly based on static constraint models constructed from single-modal images, failing to effectively integrate the complementary features of multi-modal images, and lacking the ability to accurately predict the extent of ablation thermal damage, making it difficult to achieve a dynamic balance between path safety and ablation effectiveness; furthermore, existing algorithms generally use fixed weight mechanisms for path optimization, making it difficult to adapt to individualized cases with significant differences in lesion size, location, and distribution of dangerous areas, thus limiting generalization performance. In summary, there is an urgent need to develop an intelligent path planning algorithm that can integrate multi-source dynamic constraints, incorporate clinical expert experience, and adapt to individualized case characteristics in order to break through existing technical bottlenecks and meet the needs of precision clinical treatment. Summary of the Invention
[0003] To address the shortcomings of existing methods and the needs of practical applications, and in order to solve the aforementioned problems, this invention provides a deep reinforcement learning-based system for planning the thermal ablation path for liver cancer, comprising the following modules:
[0004] The system comprises the following modules: a target and hazard area mask generation module for generating target and hazard area masks; a pipeline distance field construction module for constructing the pipeline distance field based on parallel refinement and graph structure optimization; a thermal field prediction module for obtaining the three-dimensional thermal field distribution through a fast thermal conduction approximation solution using a physical information neural network; a respiratory deformation prediction module for extracting respiratory motion features using a spatiotemporal sequence encoding and decoding network; a multi-scale feature fusion module for fusing target masks, pipeline distance fields, thermal fields, and deformation fields to construct a reinforcement learning state space; a behavior preference reward module for extracting an expert reward function through expert behavior preference distillation using maximum entropy inverse reinforcement learning; a composite reward function module for obtaining a composite reward function by combining target and hazard area masks, pipeline distance fields, and expert reward functions through multi-objective adaptive weighting and sparse reward shaping; and a decision output module for solving the optimal path using a deep deterministic policy gradient algorithm by combining the reinforcement learning state space, the composite reward function, and the spatiotemporal joint taboo field.
[0005] Optionally, generating the target and danger zone mask includes the following steps:
[0006] By adaptively adjusting parameters and enhancing the weight of lesion regions, a spatially perfectly aligned multimodal fusion image is generated. Clinical priors for liver cancer are incorporated into the segmentation network. Using the multimodal fusion image as input, initial masks for target and danger regions are generated through prior knowledge guidance and loss function optimization. The initial masks are then post-processed to obtain high-precision target and danger region masks.
[0007] Optionally, the construction of the pipeline distance field based on parallel refinement and graph structure optimization includes the following steps:
[0008] Based on the clinical semantic labeling system, dangerous area masks for pipelines are screened to obtain a pure subset of pipeline masks. Based on the three-dimensional voxels of the pipeline masks, the initial centerline is extracted through parallel refinement. The initial centerline is optimized using a graph structure to obtain the optimized complete three-dimensional centerline. Using the optimized pipeline centerline as a reference, the shortest distance from the spatial voxel to the pipeline is calculated using the fast travel method. Combined with clinical pipeline wall thickness parameters, a quantified pipeline distance field is constructed.
[0009] Optionally, obtaining the three-dimensional thermal field distribution through the rapid heat conduction approximation solution of the physical information neural network includes the following steps:
[0010] A physical constraint equation for heat conduction is constructed, and the loss function of the physical constraint is deeply embedded in the physical information neural network. The network is trained using thermal ablation data under multiple operating conditions, and the three-dimensional thermal field distribution is obtained through the trained physical information neural network.
[0011] Optionally, the extraction of respiratory motion features using a spatiotemporal sequence coding and decoding network includes the following steps:
[0012] A spatiotemporal sequence encoding and decoding network is constructed. Based on the dynamic CT sequence of the respiratory cycle, the spatiotemporal correlation features of respiratory motion in liver cancer are captured by the spatiotemporal sequence encoding and decoding network to obtain an initial deformation field. The initial deformation field is regularized and constrained. The optimized deformation field is fused with the static duct distance field to generate a spatiotemporal joint taboo field that is dynamically updated with the respiratory phase.
[0013] Optionally, the fusion of the target mask, pipe distance field, thermal field, and deformation field to construct the reinforcement learning state space includes the following steps:
[0014] Based on the target mask, pipe distance field, thermal field, and deformation field, a multi-scale fusion feature map is obtained through hierarchical feature extraction and attention-weighted fusion. The multi-scale fusion feature map is then subjected to spatial sparsity filtering and dimensionality compression encoding to generate a reinforcement learning state space.
[0015] Optionally, the extraction of the expert reward function through expert behavior preference distillation using maximum entropy inverse reinforcement learning includes the following steps:
[0016] Based on the optimality probability distribution of expert paths, a maximum entropy inverse reinforcement learning model framework is constructed. The model parameters of the maximum entropy inverse reinforcement learning model framework are optimized by gradient descent, and an expert reward function containing clinical decision-making logic is distilled from the expert path demonstration data.
[0017] Optionally, the step of combining the target and hazard area mask, pipeline distance field, and expert reward function to obtain a composite reward function through multi-objective adaptive weighting and sparse reward shaping includes the following steps:
[0018] By combining the target and hazard area mask and pipeline distance field, an objective function for path planning is constructed. The objective function is adaptively weighted according to the expert reward function, and intermediate reward or penalty signals are set at key nodes of path planning to construct a sparse reward mechanism. The adaptively weighted multi-objective reward and sparse reward are integrated to obtain a composite reward function.
[0019] Optionally, the step of constructing a sparse reward mechanism by setting intermediate reward or penalty signals at key nodes in path planning includes the following steps:
[0020] A multi-node hierarchical intermediate reward system was constructed. Based on the clinical process and constraints of liver cancer pathway planning, the reward or penalty rules of each node were clarified. The gradient distribution of reward signals was optimized to ensure that the endpoint reward remains the core incentive. The reward shaping effect was verified through training to obtain a sparse reward mechanism.
[0021] Optionally, the step of combining the reinforcement learning state space, the composite reward function, and the spatiotemporal joint taboo field to solve for the optimal path using the deep deterministic policy gradient algorithm includes the following steps:
[0022] A deep deterministic policy gradient network structure is constructed based on the reinforcement learning state space; reinforcement learning training is carried out by using the co-design of experience playback and exploration policies based on the composite reward function and spatiotemporal joint taboo field; the optimal path is obtained by optimizing the initial path generated by the policy network after training convergence through post-processing.
[0023] This invention focuses on the challenge of planning ablation pathways for liver cancer. Through adaptive registration and segmentation techniques, it accurately extracts lesions and dangerous areas such as blood vessels and bile ducts. By extracting the centerline of the pathways and constructing a distance field, using the PINN model combined with the heterogeneity of liver cancer tissue to predict the three-dimensional thermal field, and generating a dynamic taboo field through spatiotemporal network modeling of respiratory deformation, it achieves a deep fusion of static and dynamic constraints. Based on multi-source feature standardization and multi-scale fusion, a reinforcement learning state space is constructed. Using MaxEnt IRL to distill clinical expert path preferences, a composite reward function with adaptive weights and sparse reward shaping is designed. Through DDPG network learning and optimization, a safe, effective, and efficient individualized ablation pathway is derived. Clinical validation ensures that the pathway adapts to surgical procedures and thermal damage control requirements. This algorithm addresses clinical pain points such as respiratory interference, danger zone avoidance, and individualized adaptation, improving the accuracy and safety of liver cancer thermal ablation surgery, reducing damage to normal tissues and complications, providing intelligent decision support for interventional therapy, and promoting the upgrade of minimally invasive liver cancer treatment towards precision and intelligence. Attached Figure Description
[0024] Figure 1 This is a framework diagram of a liver cancer thermal ablation path planning system based on deep reinforcement learning, provided for an embodiment of the present invention. Detailed Implementation
[0025] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.
[0026] Throughout this specification, references to an embodiment, example, or illustration mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, phrases appearing in various places throughout the specification, such as "in one embodiment," "in an embodiment," "an example," or "an illustration," do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in any suitable combination and / or sub-combination in one or more embodiments or examples. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0027] Please see Figure 1 To address the aforementioned problems, this invention provides a liver cancer thermal ablation path planning system based on deep reinforcement learning, such as... Figure 1 As shown, it includes: target and danger area mask generation module 1, pipeline distance field construction module 2, thermal field prediction module 3, breathing deformation prediction module 4, multi-scale feature fusion module 5, behavior preference reward module 6, composite reward function module 7, and decision output module 8.
[0028] The target and danger zone mask generation module 1 is used to generate target and danger zone masks.
[0029] In this embodiment, spatial and grayscale differences between multi-source images are eliminated through adaptive modal registration. Combined with prior clinical knowledge of liver cancer, the segmentation network is guided to focus on key regions, achieving high-precision mask extraction of liver cancer lesions (target regions) and blood vessels, bile ducts, gallbladder, etc. (dangerous regions). The generation of target and danger region masks includes the following steps:
[0030] S11. By adjusting adaptive parameters and enhancing the weight of the lesion area, a spatially fully aligned multimodal fusion image is generated.
[0031] First, grayscale normalization is performed. For CT images, HU value standardization is used (mapping the HU value range corresponding to liver tissue to the [0,255] interval), and for T2-weighted MRI images, Z-score normalization is used (eliminating grayscale shifts caused by different equipment and scanning parameters) to ensure that the grayscale dynamic range of the two modalities is consistent.
[0032] Subsequently, hybrid noise filtering was performed. First, a Gaussian filter with a 3×3 convolution kernel was used to smooth the high-density salt-and-pepper noise in the CT images (such noise is easily generated in CT images of liver cancer patients due to uneven distribution of contrast agent). Then, a median filter with a 5×5 convolution kernel was used to remove motion artifacts and magnetic susceptibility artifacts in the MRI images (respiratory motion can easily cause artifacts around the liver cancer area). During the filtering process, the signal-to-noise ratio (SNR) of the images was monitored in real time to ensure that the edge features of the liver cancer lesion area were not overly smoothed after filtering.
[0033] Finally, the resolution of the two modalities was unified, and the MRI images were resampled to the spatial resolution of the CT images to avoid registration errors caused by resolution differences.
[0034] Furthermore, by dynamically adjusting the registration parameters and focusing on the lesion area, the registration accuracy of the liver cancer lesion area is improved, ensuring the spatial alignment accuracy of multimodal images.
[0035] Specifically, we first construct an adaptive registration optimizer based on normalized mutual information (NMI), select MRI images as the reference modality (MRI provides better soft tissue contrast between liver cancer lesions and normal liver tissue), and CT images as the floating modality (CT provides clearer display of dangerous areas such as blood vessels). We use the NMI value as the evaluation index of registration accuracy (the higher the NMI value, the better the registration effect).
[0036] The dynamic parameter adjustment process is then initiated, with the initial iteration step size set to 0.1. Every 5 iterations, the NMI value change rate is calculated. If the change rate is greater than 5%, the current step size is maintained and iteration continues. If the change rate is between 1% and 5%, the step size is adjusted to 0.05 to improve registration accuracy. If the change rate is less than 1%, the step size is adjusted to 0.01 to avoid excessive iteration. The registration transformation first achieves global alignment of the image through affine transformation (solving global differences such as translation, rotation, and scaling), and then achieves fine local alignment through thin-plate spline transformation (adapting to the irregular deformation of tissues surrounding liver cancer lesions).
[0037] Meanwhile, a registration weight mask for liver cancer lesion areas is generated based on a prior clinical knowledge base, assigning a weight of 1.5 times to lesion areas and a weight of 0.8 times to non-lesion areas, making the registration optimization process more focused on key areas.
[0038] After registration, a resampling algorithm is used to map the floating modal CT images to the spatial coordinate system of the reference modal MRI, outputting a spatially aligned multimodal fusion image.
[0039] S12. Integrate the clinical priors of liver cancer into the segmentation network. Using the multimodal fused image as input, generate the initial masks of the target region and the danger region through prior knowledge guidance and loss function optimization.
[0040] By incorporating prior clinical knowledge of liver cancer into the segmentation network, the problems of ambiguous lesion boundaries and missed segmentation of dangerous areas are solved, achieving precise pixel-level annotation of target and dangerous areas.
[0041] Specifically, a basic segmentation network was first built based on U-Net++, and the decoder structure of the network was optimized. The original single-scale upsampling was changed to multi-scale upsampling (fusing feature maps of three scales: 1×1, 2×2, and 4×4) to improve the adaptability to liver cancer lesions of different sizes.
[0042] Subsequently, a prior knowledge guidance submodule was designed to extract key information from the clinical prior knowledge base, including the common topological morphology of liver cancer lesions (mostly circular with irregular boundaries) and the relative positional relationship between dangerous areas and lesions (blood vessels often surround lesions, and bile ducts are often distributed parallel to the liver lobes). This information was transformed into a structured attention weight map, assigning 2.0 times the attention weight to the boundary pixels of dangerous areas such as blood vessels and bile ducts, 1.8 times the attention weight to the edge pixels of liver cancer lesions, and 0.5 times the weight to the pixels of normal liver tissue. This weight map was then input into the fourth feature fusion layer of the network, where the feature response of key areas was enhanced through element-wise multiplication.
[0043] In terms of loss function design, a hybrid loss function of Dice loss and boundary enhancement loss is adopted. Dice loss is used to optimize the overall segmentation accuracy, while boundary enhancement loss is used to specifically penalize boundary segmentation error by calculating the Hausdorff distance between the predicted boundary and the true boundary, thus solving the problem of blurred boundaries between liver cancer lesions and surrounding tissues.
[0044] During network training, clinically labeled multimodal fusion images of liver cancer were used as the training set, of which 60% were primary liver cancer cases and 40% were metastatic liver cancer cases, covering lesions of different sizes (1-10cm) and different locations (left lobe of liver, right lobe of liver, and porta hepatis) to improve the network's generalization ability.
[0045] After training, the registered fused images are input into the network, and the network outputs the initial masks of the target region (liver cancer lesion) and the danger region (blood vessels, bile ducts, gallbladder). The mask is a single-channel binary image, where the foreground pixels (value 1) represent the target / danger region and the background pixels (value 0) represent normal tissue.
[0046] S13. Post-process the initial mask to obtain a high-precision target and danger area mask.
[0047] Post-processing optimizations are performed to address issues such as holes and artifacts in the initial mask, improving the integrity and accuracy of the mask and providing a reliable spatial constraint benchmark for subsequent path planning.
[0048] First, perform a morphological closing operation, select a circular structuring element (with a radius of 2 pixels to fit the size of the tiny holes in liver cancer lesions), dilate the initial mask and then erode it. The dilation operation can fill the tiny holes inside the lesion (caused by the segmentation network's insufficient response to low-contrast areas), and the erosion operation can eliminate the tiny protrusions at the boundary, keeping the overall shape of the lesion unchanged.
[0049] Connectivity analysis was then performed to calculate the area of all connected components in the mask. An area threshold was set based on clinical priors: for liver cancer lesion masks, areas smaller than 5 mm were removed. 2 Connected components (mostly artifacts) with a preserved area greater than or equal to 5 mm 2 Connected components (meeting the minimum clinically ablationable lesion size); for danger zone masks, areas smaller than 2mm are removed. 2 Connected regions (mostly noise interference) with a reserved area greater than or equal to 2 mm 2 The connected domain (meets the minimum diameter of blood vessels and bile ducts).
[0050] Finally, boundary smoothing is performed. A Gaussian smoothing algorithm is used to filter the boundary pixels of the mask, with a smoothing coefficient of 0.8. This eliminates jagged artifacts at the boundaries while maintaining accurate boundary positions.
[0051] After post-processing, the mask is quality verified by calculating the Dice similarity coefficient between the mask and the results manually annotated by clinical experts. The Dice coefficient for the target area is required to be ≥0.9 and the Dice coefficient for the danger area is required to be ≥0.85. If the requirements are not met, the morphological operation parameters are readjusted and the processing procedure is repeated until the requirements are met, and finally a high-precision target and danger area mask is obtained.
[0052] The pipeline distance field construction module 2 is used to construct the pipeline distance field based on parallel refinement and graph structure optimization.
[0053] In this embodiment, for pipe-like structures (blood vessels, bile ducts) in hazardous areas, parallel refinement improves the efficiency of centerline extraction, and graph structure optimization addresses the centerline breakage problem, ultimately constructing a three-dimensional pipe distance field to quantify the shortest distance from spatial points to pipes. The construction of the pipe distance field based on parallel refinement and graph structure optimization includes the following steps:
[0054] S21. Based on the clinical semantic labeling system, filter the pipeline-type hazardous area masks to obtain a pure pipeline mask subset.
[0055] First, the danger zone mask is classified and defined based on the clinical semantic labeling system, and the semantic features of duct-type regions (blood vessels and bile ducts) are clarified. Blood vessel regions show high-density enhancement features in enhanced CT registration images, and bile duct regions show high signal features in T2-weighted MRI registration images. These features are transformed into screening judgment conditions.
[0056] Then, a semantic filtering operation is performed. Through pixel-level feature matching, regions that simultaneously satisfy continuous tubular topology and corresponding modal signal features are extracted from the danger area mask and marked as pipeline candidate regions.
[0057] Next, non-pipe areas are eliminated. For non-pipe danger areas such as the gallbladder, whose semantic features are a circular closed structure with no connection to the pipeline, the connectivity relationship between the candidate area and the pipeline structure is determined by connected component analysis, and the circular closed areas with no connection are eliminated.
[0058] Finally, the screening results were verified by calculating the Dice similarity coefficient between the duct mask subset and the duct regions manually annotated by clinical experts. A coefficient ≥ 0.85 was required to ensure that no duct regions were missed and no non-duct regions remained after screening, resulting in a pure duct mask subset. This process accurately separates duct-like structures (blood vessels, bile ducts) from the mixed hazardous area mask, eliminating non-duct hazardous areas such as the gallbladder. This provides pure target data for subsequent centerline extraction, meeting the clinical need to prioritize the avoidance of duct-like hazardous structures during liver cancer thermal ablation.
[0059] S22. Based on pipe masking, three-dimensional voxels are used to extract the initial centerline through parallel refinement.
[0060] To address the issues of low efficiency and easy breakage of the centerline due to the complex branching of liver cancer ducts in traditional serial refinement algorithms, a three-dimensional voxel block parallel processing method is used to improve efficiency while ensuring the connectivity of the duct topology, thus achieving rapid and accurate extraction of the initial centerline.
[0061] Specifically, three-dimensional voxel preprocessing is first performed. Based on the clinical topological characteristics of liver cancer duct structure (mostly tree-like branches with a branch diameter of 2-10 mm), the three-dimensional voxel space of the duct mask is divided into 16×16×16 mm cubic sub-blocks, with a 2 mm overlap area set between adjacent sub-blocks (to avoid breakage caused by branches crossing sub-blocks).
[0062] Subsequently, a parallel thinning operation pool was constructed, employing a multi-threaded parallel mechanism. An independent computation thread was allocated to each sub-block, and thinning operations were performed simultaneously. For each sub-block, edge voxels (voxels that are only connected to 3 or fewer adjacent voxels) were first identified. It was then determined whether stripping these voxels would disrupt the connectivity of the pipeline. If not, the stripping was performed, and this process was repeated until the voxel width of the pipeline structure within the sub-block was reduced to 1 pixel. A unified thinning stop condition was set, and the pipeline voxel width of each sub-block was monitored in real time using a voxel width detection algorithm. Thinning was stopped when the pipeline voxel width of all sub-blocks was 1 pixel and the overall connectivity was not disrupted.
[0063] Finally, sub-block fusion is performed. Based on voxel matching of overlapping areas, the refined results of each sub-block are stitched together to form a complete three-dimensional initial centerline, eliminating the boundary gaps caused by the block processing.
[0064] S23. Optimize the initial centerline using a graph structure to obtain the optimized complete three-dimensional centerline.
[0065] To address defects such as branch breaks and edge burrs in the initial centerline, and combining the clinical topological priors of liver cancer ducts (such as portal vein and hepatic vein), graph structure modeling is used to complete and smooth the centerline, ensuring the consistency between the centerline and the actual duct structure.
[0066] The specific operation is as follows: First, perform graph structure modeling of the center line, taking each voxel of the initial center line as a node of the undirected graph. If two voxels are adjacent in three-dimensional space, then establish an edge connection between the corresponding nodes to form an undirected graph representing the topological relationship of the center line.
[0067] Subsequently, clinical topological constraint rules were constructed, and constraint conditions were extracted based on the prior clinical knowledge base of liver cancer channels: such as the angle between the main branch and branches of the portal vein ranging from 30° to 60°, and the branches of the hepatic veins being mostly perpendicular to the liver surface. These conditions were then transformed into connection constraints for graph nodes. Branch break completion was performed, and a graph node traversal algorithm was used to identify break regions (regions with isolated nodes or nodes with a degree of 1 and no adjacent nodes). The spatial distance between the nodes at both ends of the break was calculated (a distance threshold of ≤5mm was set to meet the reasonable interval of clinical channel branches) and the directional angle (which must meet the topological angle constraint of the corresponding channel). If both distance and directional conditions were met, Dijkstra's shortest path algorithm was used to generate a completion path between the nodes at both ends of the break, integrating the isolated nodes into the overall graph structure.
[0068] Next, a smoothing process to eliminate burrs is performed. The 3-neighbor mean filtering algorithm is used to smooth the coordinates of each node in the graph. The average coordinates of the current node and its three neighboring nodes are used as the new coordinates of the node. This process is repeated three times to eliminate the tiny burrs on the surface of the center line (caused by edge voxel recognition errors during the thinning process).
[0069] Finally, the optimized centerline was verified by calculating the degree of overlap between the centerline and the clinical anatomical structure of the channel. The degree of overlap was required to be ≥0.9. At the same time, it was ensured that the centerline had no intersections or isolated branches, thus obtaining the optimized complete three-dimensional centerline.
[0070] S24. Using the optimized pipeline centerline as a reference, calculate the shortest distance from the spatial voxel to the pipeline using the rapid travel method, and construct a quantitative pipeline distance field by combining clinical pipeline wall thickness parameters.
[0071] First, the parameters of the fast travel method are adapted. Based on the spatial range of the liver cancer surgery plan, the voxel resolution of the computation grid is set to be consistent with the image resolution to ensure the accuracy of distance calculation. The optimized centerline voxel is set as the seed point (distance value is 0), and the voxels outside the liver are set as unreachable regions (distance value is set to infinity) to clarify the computational boundary of the fast travel method.
[0072] Subsequently, three-dimensional distance calculations are performed. Based on the partial differential equation solution logic of the fast travel method, the calculation is performed by gradually expanding outward from the seed point to calculate the shortest Euclidean distance from each reachable voxel in three-dimensional space to the center line, thereby generating the initial distance field.
[0073] Next, clinical tube wall thickness parameters were incorporated. Based on clinical pathological data and surgical guidelines, safe distance thresholds for different tubes were determined. For example, the wall thickness of large blood vessels such as the portal vein and hepatic vein is about 1-2 mm, and the safe distance threshold is set to 3 mm (wall thickness + 1 mm safety redundancy). The wall thickness of small intrahepatic bile ducts is about 0.5 mm, and the safe distance threshold is set to 1.5 mm. These thresholds were embedded into the initial distance field to form a mapping relationship between distance value and risk level (distance < threshold is high risk, distance ≥ threshold is low risk).
[0074] Finally, a complete pipeline distance field is constructed, which associates and stores distance values, hazard levels, and pipeline topology information to form a pipeline distance field containing multi-dimensional information such as spatial location, distance value, hazard level, and pipeline branches. The pipeline distance field can quickly provide quantitative constraints on the pipeline hazard area from any spatial location for the path planning algorithm, avoiding complications such as bleeding and bile leakage caused by the ablation needle invading the pipeline.
[0075] The thermal field prediction module 3 is used to obtain the three-dimensional thermal field distribution by rapidly solving the thermal conduction approximation through a physical information neural network.
[0076] In this embodiment, obtaining the three-dimensional thermal field distribution through the rapid heat conduction approximation solution of a physical information neural network includes the following steps:
[0077] S31. Construct the physical constraint equation for heat conduction, and embed the loss function of the physical constraint deep into the physical information neural network.
[0078] The Pennes equation, which considers the effect of blood perfusion, is used as the basis for biological heat conduction. This equation can accurately describe the coupling effect of heat conduction and blood convection heat transfer in biological tissues, which fits the in vivo heat transfer characteristics of liver cancer tissue. Subsequently, the heterogeneity parameters of liver cancer tissue are modeled. Based on the clinical pathology report, the physical parameters of two key tissues are extracted: thermal conductivity, specific heat capacity and blood perfusion rate of tumor tissue and normal liver tissue. A tissue type-physical parameter mapping table is established.
[0079] Next, the physical constraints for the partitioning are defined. The target region mask is used as the basis for partitioning. By matching the mask with image voxels pixel by pixel, corresponding physical parameters are assigned to the tumor region and the normal liver tissue region, respectively, and the physical constraint equations for spatial partitioning are constructed. In the tumor region, due to the low blood perfusion rate, the weight of the convection heat transfer term is weakened; in the normal liver tissue region, the weight of the convection heat transfer term is strengthened to ensure that the equations adapt to the heat transfer laws of different regions. Finally, the effectiveness of the constraint equations is verified. The thermal field distribution under simple working conditions is calculated by the finite element method and compared with the theoretical derivation results of the partition constraint equations. The consistency of the temperature distribution trend is required to be ≥90% to ensure the physical rationality of the constraint equations.
[0080] Furthermore, a multi-branch PINN (Physical Information Neural Network) structure was designed to deeply embed physical constraints into the loss function. First, a multi-branch input layer was designed, constructing three parallel input branches: ① Spatial path branch, which inputs the three-dimensional coordinate sequence of the ablation needle, and uses an embedding layer to map the coordinates into a 64-dimensional feature vector; ② Parameter branch, which inputs ablation needle parameters (power, diameter, ablation time) and tissue physical parameters (obtained from a mapping table), and transforms them into a 64-dimensional feature vector through a fully connected layer; ③ Region identification branch, which inputs the target region mask, labels the tissue type of the currently predicted voxel, and outputs a 1-dimensional binary feature (1 for tumor, 0 for normal liver tissue).
[0081] Subsequently, a hidden layer structure was designed, employing four fully connected hidden layers (with 256, 128, 64, and 32 neurons per layer, respectively). Residual connections were added between layers to directly transmit features from the previous layer via shortcut paths, thus solving the gradient vanishing problem in deep networks. Additionally, a LeakyReLU activation function (with a slope of 0.01) was added after each layer to enhance the network's ability to fit the nonlinear features of the thermal field.
[0082] Next, the output layer was designed, adopting a single-output node structure to output the absolute temperature value of the corresponding voxel. The core loss function was constructed using a hybrid mode of data loss and physical constraint loss: the data loss used mean squared error (MSE) to calculate the error between the network's predicted temperature and a small number of clinically measured temperatures; the physical constraint loss calculated the residual term of the heat conduction equation (substituting the temperature value output by the network into the partitioned physical constraint equation, the smaller the residual, the more the prediction result conforms to physical laws), and also added boundary condition constraint terms—ablation needle surface temperature constraint (set to 100-120℃ to match the working temperature of clinical radiofrequency ablation needles) and normal human body temperature constraint (the boundary temperature of the liver region was set to 37℃). Finally, the weight allocation of the hybrid loss function was: data loss accounted for 40%, physical constraint loss accounted for 50%, and boundary condition constraint loss accounted for 10%.
[0083] Finally, network initialization is performed, using a He normal distribution to initialize weights and setting the bias term to 0 to ensure network training stability. Through deep fusion of multi-branch input and physical constraint loss, the PINN network can accurately capture the thermal conductivity characteristics of liver cancer tissue while improving gradient propagation efficiency.
[0084] S32. Use multi-condition thermal ablation data for training, and obtain the three-dimensional thermal field distribution through the trained physical information neural network.
[0085] First, the training dataset was constructed. For the clinical data portion, intraoperative measured data from liver cancer thermal ablation surgery were collected, including ablation parameters, lesion location and size, and temperature measurements at key intraoperative points. Preoperative and postoperative ablation range images of the corresponding cases were also collected simultaneously. For the numerical simulation data portion, based on the partitioned physical constraint equations, multi-condition thermal field data were generated using finite element software (such as ANSYS), covering common clinical scenarios: ablation power of 50 / 80 / 100 / 120 / 150W, ablation needle diameter of 1.6 / 1.8 / 2.0 / 2.2 / 2.4mm, lesion size of 1-10cm (divided in 2cm intervals), and lesion location covering the left / right lobe / hepatic hilum of the liver. Three-dimensional thermal field data of 300×300×300mm were generated for each condition, ultimately producing multiple sets of numerical simulation data.
[0086] Subsequently, data preprocessing was performed, mapping clinical data and numerical simulation data to the same spatial coordinate system (based on preoperative CT images). Temperature data was normalized (mapped to the [0,1] interval, with a normalization baseline of 37℃ (normal body temperature) to 120℃ (ablation needle temperature)). The dataset was divided into training set (80%), validation set (10%), and test set (10%). Next, training weights for lesion regions were set. Based on the target region mask, voxels of lesion regions were assigned a training weight of 1.5 times, and voxels of normal liver tissue were assigned a weight of 0.8 times, making the network training process more focused on the thermal field fitting of the lesion region.
[0087] The training process uses the Adam optimizer with a learning rate decay strategy. Finally, the validation results are summarized to ensure that all indicators meet clinical requirements, forming the final thermal field prediction model. The model output is three-dimensional thermal field distribution data containing spatial coordinates, temperature values, and ablation range markers.
[0088] The respiratory deformation prediction module 4 is used to extract respiratory motion features using a spatiotemporal sequence encoding and decoding network.
[0089] In this embodiment, to address the dynamic deformation of liver cancer cells during respiration, a spatiotemporal sequence encoder-decoder network is used to extract respiratory motion features. Combined with deformation field regularization constraints, the smoothness and topological consistency of deformation prediction are ensured, enabling tissue deformation prediction at any time point within the respiratory cycle. The extraction of respiratory motion features using the spatiotemporal sequence encoder-decoder network includes the following steps:
[0090] S41. Construct a spatiotemporal sequence encoding and decoding network. Based on the dynamic CT sequence of the respiratory cycle, use the spatiotemporal sequence encoding and decoding network to capture the spatiotemporal correlation features of respiratory motion in liver cancer and obtain the initial deformation field.
[0091] First, respiratory phase calibration is performed. By collecting respiratory waveform signals from the patient's intraoperative respiratory sensors, key phases such as end of inspiration (peak value of respiratory waveform) and end of expiration (trough value of respiratory waveform) are initially marked. Then, based on the spatial location characteristics of the liver contour (the liver volume is largest at the end of inspiration and smallest at the end of expiration), the phase is finely corrected, and the entire respiratory cycle is divided into 10 uniform phase frames (to ensure temporal continuity).
[0092] Subsequently, feature enhancement of the target and danger areas is performed. Based on the mask of the target (hepatocellular carcinoma lesion) and the danger area (blood vessels and bile ducts), contrast-limited adaptive histogram equalization (CLAHE) is performed on the mask-covered area to improve the edge clarity of the lesion and danger area. At the same time, mild Gaussian smoothing (3×3 convolution kernel) is applied to the normal liver tissue outside the area to reduce background noise while preserving key features.
[0093] Finally, invalid frames were removed using a dual set of criteria: ① Signal-to-noise ratio (SNR) threshold: the SNR of each frame was calculated, and blurry frames with an SNR < 20dB (artifact frames caused by shortness of breath) were removed; ② Deformation consistency threshold: deformation consistency was calculated by the overlap of the liver contour in adjacent frames, and aberration frames with an overlap < 85% (invalid frames caused by patient movement) were removed.
[0094] After processing, a standardized dynamic CT sequence of the respiratory cycle is output (containing 10 uniform phase frames, with continuous temporal sequence between frames and clear features of key regions), along with respiratory phase labels for each frame. Using the original dynamic CT sequence as input, targeted phase calibration and artifact removal provide a reliable dynamic image foundation for subsequent spatiotemporal feature extraction.
[0095] Furthermore, targeting the dual characteristics of spatial deformation and temporal evolution of respiratory motion in liver cancer, a 3D spatiotemporal convolutional coding-decoding network is constructed to achieve deep fusion of respiratory temporal features and spatial deformation features, simultaneously capturing the spatiotemporal correlation of respiratory motion and accurately outputting the initial three-dimensional deformation field.
[0096] First, the overall network architecture is designed, adopting a three-stage structure of encoding-attention-decoding:
[0097] The input is a processed 10-frame continuous phase dynamic CT sequence (size 256×256×64, voxel resolution 1×1×1mm), and the output is a three-dimensional deformation field of each phase relative to the reference frame (end of expiration) (size is the same as the input, each voxel contains displacement components in the x, y, and z directions).
[0098] Encoding layer design: A 4-layer 3D spatiotemporal convolution (Conv3d) is adopted, with a kernel size of 3×3×3 and a stride of 2×2×2. The number of channels is 64→128→256→512 respectively. BatchNorm normalization and LeakyReLU activation function (slope 0.01) are added after each layer. The spatiotemporal convolution simultaneously extracts the spatial structural features of respiratory motion (such as liver contour morphology) and temporal evolution features (such as the change of contour with phase). The feature map downsampling is achieved by increasing the stride to improve the feature abstraction ability.
[0099] Attention gating mechanism design: Four parallel attention gating modules are embedded between the encoding and decoding layers (each corresponding to an output feature of the encoding layer). Figure 1 (One-to-one correspondence) The target and danger area mask are transformed into attention weight maps. The output feature map of the coding layer is weighted by element-wise multiplication to enhance the spatiotemporal feature response of lesions and danger areas and suppress invalid features of background areas.
[0100] Decoding layer design: A 4-layer 3D deconvolution (Deconv3d) is adopted, with a kernel size of 3×3×3, a stride of 2×2×2, and the number of channels successively 512→256→128→64. BatchNorm and LeakyReLU are also added after each layer. The spatial resolution of the feature map is gradually restored through deconvolution. Finally, the number of channels is mapped to 3 (corresponding to the displacement in the x, y, and z directions) through 1×1×1 convolution, and the initial three-dimensional deformation field is output.
[0101] The network initialization uses a Xavier normal distribution to initialize the convolutional layer weights, with the bias term initialized to 0 to ensure gradient stability during training. Using dynamic image sequences as input, deep fusion of spatiotemporal convolution and attention-guided methods accurately captures the spatiotemporal correlation features of respiratory motion in liver cancer patients. The output initial deformation field can preliminarily reflect the dynamic displacement patterns of tissues within the respiratory cycle.
[0102] S42. The initial deformation field is regularized and constrained. The optimized deformation field and the static pipeline distance field are fused to generate a spatiotemporal joint taboo field that is dynamically updated with the breathing phase.
[0103] To address potential issues such as local distortion, abrupt displacement, and topological disruption in the initial deformation field, a dual regularization constraint is designed based on the physiological characteristics of liver cancer tissue (non-rigid, continuous deformation) to optimize the smoothness and topological consistency of the deformation field, ensuring that the deformation prediction results conform to clinical physiological logic. The specific operation is as follows:
[0104] First, we define the mathematical expression for the double regularization constraint:
[0105] ① Smoothness constraint: The L2 gradient norm of the deformation field is used to construct the constraint term. The first-order partial derivatives of the three-dimensional deformation field in the x, y, and z directions are calculated. The average of the squares of the partial derivatives is used as the smoothness penalty term. The larger the penalty term value, the more drastic the local change in the deformation field. The constraint objective is to minimize this penalty term to ensure the spatial continuity of the deformation field (which conforms to the physiological characteristics of non-rigid continuous deformation of liver tissue). The weight of the smoothness constraint term is set to 0.15.
[0106] ② Topological consistency constraint: Based on the dangerous area mask, a topological invariance verification rule is constructed. The topological structure parameters (such as the number of blood vessel branches and bile duct connectivity) of the dangerous area (blood vessels, bile ducts) under the action of the initial deformation field are calculated and compared with the topological parameters of the reference frame (end-expiration). The parameter difference value is used as a topological consistency penalty term to constrain the topological structure of the dangerous area from being destroyed (avoiding results that do not conform to clinical anatomical logic, such as blood vessel rupture and bile duct blockage due to deformation prediction errors). The weight of the topological consistency constraint term is set to 0.2.
[0107] Subsequently, a network loss function including a regularization term was constructed. The total loss = deformation loss (MSE error between the initial deformation field and the true deformation field, weighted at 0.65) + smoothness constraint term + topology consistency constraint term. The Adam optimizer was used to optimize the total loss. During the optimization process, the quality of the deformation field was monitored in real time, and the displacement distribution of the deformation field was observed using visualization tools. If local distortion regions existed, the weight of the smoothness constraint term was appropriately increased. If the topology of the danger zone was abnormal, deformation samples of such regions were added, and the model was fine-tuned. Based on the initial deformation field, through precise control of the dual regularization constraints, an optimized deformation field with both smoothness and topology consistency meeting clinical requirements was obtained, providing a reliable foundation for subsequent dynamic danger zone prediction.
[0108] Furthermore, the optimized deformation field is deeply integrated with the static pipeline distance field to generate a spatiotemporal joint taboo field that is dynamically updated with the breathing phase. This provides dynamic hazard constraints that link spatial location and temporal phase for subsequent path planning, preventing the static taboo field from being unable to adapt to the dynamic deformation of breathing.
[0109] First, dynamic location prediction of the danger zone is performed. Based on the optimized three-dimensional deformation field, the danger zone mask of the reference frame (end-expiration) is mapped to each phase frame within the respiratory cycle. By calculating the displacement component of the deformation field, the new coordinates of all voxels in the danger zone under each phase are calculated, generating the dynamic danger zone mask corresponding to each phase (including the real-time spatial position of blood vessels and bile ducts in that phase). Then, the dynamic danger zone and the tube distance field are fused. Based on the tube distance field, the voxel distance values corresponding to the dynamic danger zone mask of each phase are extracted, and a phase-distance association table is constructed. The distance values are dynamically classified into danger levels: ① High danger level (distance < safety threshold, such as 3mm for large blood vessels and 1.5mm for small bile ducts), marked as the forbidden core area; ② Medium danger level (distance ≥ safety threshold and < safety threshold + 1mm), marked as the forbidden transition area; ③ Low danger level (distance ≥ safety threshold + 1mm), marked as the safe area.
[0110] Next, a spatiotemporal joint taboo field data structure is constructed, which uses a four-dimensional tensor to store information with three-dimensional spatial coordinates and one-dimensional time phase. Each tensor element contains four core contents: spatial location (x, y, z), breathing phase (0-9), hazard level (high / medium / low), and corresponding pipeline branch information, so as to realize the rapid query of the hazard level at any spatiotemporal point.
[0111] Finally, spatiotemporal joint taboo field verification was performed. Respiratory motion measurement data from multiple patients (including the actual danger zone locations in each phase) were selected, and the overlap between the predicted danger zone locations and the actual locations was calculated (Dice coefficient ≥ 0.85). At the same time, the real-time performance of taboo field updates under different phases was verified (single phase update time ≤ 0.1s) to ensure that the taboo field can accurately and quickly reflect the dynamic changes of danger zones within the respiratory cycle.
[0112] The multi-scale feature fusion module 5 is used to fuse target mask, pipeline distance field, thermal field and deformation field to construct reinforcement learning state space.
[0113] In this embodiment, the fusion of the target mask, pipe distance field, thermal field, and deformation field to construct the reinforcement learning state space includes the following steps:
[0114] S51. Based on the target mask, pipeline distance field, thermal field and deformation field, a multi-scale fused feature map is obtained through hierarchical feature extraction and attention-weighted fusion.
[0115] First, differentiated standardization strategies are designed for different types of features to ensure that they are appropriate for the clinical physical significance of each feature:
[0116] ① Mask feature standardization: The target (liver cancer lesion) and the danger area mask itself are single-channel binary images (foreground 1, background 0), directly retaining the original binary features. At the same time, the mask edges are refined through morphological erosion operation (1×1 pixel structuring element) to avoid feature redundancy caused by excessive expansion of edge pixels.
[0117] ②Standardization of the pipe distance field: Based on the pipe distance field, the distribution of distance values of all voxels in the liver region is first statistically analyzed to determine the effective range of distance values (0-20mm, covering common safe pipe distance thresholds in clinical practice). The min-max normalization method is used to linearly map the distance values to the [0,1] interval, where the minimum distance value is 0 (pipe centerline) and the maximum distance value is set to 20mm (distances outside this range have no significant significance for risk assessment and are uniformly mapped to 1).
[0118] ③Standardization of thermal field characteristics: Based on the absolute temperature value output by the thermal field prediction model, the relative temperature value (relative to normal body temperature 37℃) is calculated, that is, relative temperature = absolute temperature - 37℃. Then, the relative temperature value is normalized to the [0,1] interval. The effective temperature range for clinical ablation is 45-120℃, corresponding to a relative temperature of 7-83℃. Temperature values outside this range are truncated to 7℃ (lower limit) and 83℃ (upper limit) respectively, to ensure that the standardized value can accurately reflect the degree of thermal damage.
[0119] ④ Deformation field feature standardization: Extract the displacement components in the x, y, and z directions of the deformation field, and count the maximum displacement value of the liver region during the respiratory cycle (the liver displacement caused by breathing in clinical liver cancer patients is mostly 5-30 mm). Divide the displacement component in each direction by the maximum displacement value in the corresponding direction to obtain the relative displacement value in the interval [-1,1]. Then, map it to the interval [0,1] through translation transformation (add 1 and divide by 2) to retain the displacement direction information while achieving scale uniformity.
[0120] After standardization, feature distribution consistency verification is performed: calculate the mean (required to be between 0.4 and 0.6) and standard deviation (required to be between 0.2 and 0.3) of each feature after standardization. If the standard is not met, readjust the normalization range to ensure that all features are in similar numerical distribution ranges and completely eliminate the interference of dimensional differences on subsequent fusion.
[0121] Furthermore, to address the dual requirements of fine-grained spatial localization (such as lesion margins and duct locations) and macroscopic physical constraints (such as thermal field distribution and respiratory deformation) in liver cancer pathway planning, a pyramid-shaped multi-scale fusion module is constructed. Through hierarchical feature extraction and attention weighting, efficient fusion of multi-source features is achieved, enhancing the key feature responses of targets and dangerous areas.
[0122] Specifically, a three-tiered pyramid architecture (shallow-middle-deep) is constructed, with each level's functions and operations precisely adapted to the characteristics of liver cancer.
[0123] ① Shallow Feature Fusion (Fine-Grained Spatial Feature Layer): Focusing on the accuracy of spatial location, the standardized mask features (target + danger zone) and the pipeline distance field features are input. Both are fine-grained spatial features. The channel splicing method is used to fuse them, splicing the two channels of the mask (target and danger zone) and the one channel of the distance field into a three-channel feature map, preserving the accurate spatial attributes of each voxel, and adapting to the fine positioning requirements of the ablation needle path.
[0124] ② Mid-level feature fusion (mesoscale physical feature layer): Focusing on the effectiveness of physical constraints, standardized thermal field features and deformation field features are input. Thermal field features reflect energy distribution, and deformation field features reflect dynamic displacement. Both are mesoscale physical features. An element-weighted product + channel fusion strategy is adopted: Based on clinical priors, weights are set (thermal field feature weight 0.6, deformation field feature weight 0.4, because thermal damage risk has a higher priority for path planning constraints). The corresponding voxel values of the two are weighted and multiplied, and then stitched with the original thermal field and deformation field features to form a 3-channel physical feature map, which strengthens the thermal-deformation coupling constraint information.
[0125] ③ Deep Feature Fusion (Cross-Scale Attention Fusion Layer): A spatial attention mechanism is constructed to achieve weighted fusion of shallow and mid-level features. First, an initial attention weight map is generated based on the target and danger region masks (target region weight 1.8, danger region weight 2.0, normal liver tissue weight 0.5). Then, the weight map is smoothed by a 3×3×3 three-dimensional convolution kernel to make the weight transition more consistent with the spatial continuity of the tissue. Subsequently, the shallow 3-channel spatial features and the mid-level 3-channel physical features are concatenated (6 channels in total) and multiplied element-wise with the smoothed attention weight map to enhance the key region features. Finally, the 6-channel fused features are reduced to 3 channels by a 1×1×1 convolution kernel to obtain the final multi-scale fused feature map (resolution 1×1×1mm, containing spatial-physical coupling constraint information).
[0126] After fusion, the feature response intensity is verified: the average feature response value of the target and the danger zone is calculated, which is required to be more than 3 times higher than that of normal liver tissue, to ensure that the fused features can accurately focus on the key areas.
[0127] S52. Perform spatial sparse filtering and dimensional compression encoding on the multi-scale fused feature map to generate a reinforcement learning state space.
[0128] First, spatial sparsity screening is performed. Based on preoperative CT images, the reachable space of the ablation needle is defined—with the liver outline as the boundary. Combined with clinical surgical approach guidelines (such as percutaneous approach, which must avoid ribs, intestines, etc.), a reachable space mask is generated: only voxels of connected regions from the liver interior and skin needle entry point to the liver surface are retained, while voxels of inaccessible regions such as the outside, bones, and intestines are removed (these voxels are meaningless for path planning). The integrity of the reachable space is verified through connected component analysis, requiring the reachable space to cover all possible needle entry paths and target areas to ensure that no critical areas are missed. After screening, the number of voxels of multi-scale fusion features can be reduced from millions to less than hundreds of thousands, achieving preliminary dimensional simplification.
[0129] Next, voxelized sparse coding and dimensionality reduction are performed. The selected voxel features are structured and encoded by concatenating the 3-channel fusion features of each voxel with the 3D spatial coordinates (x, y, z) to form a 6-dimensional original feature vector for each voxel. Principal component analysis (PCA) is used for dimensionality reduction, with core parameters set based on the importance of clinical features: the cumulative contribution rate of the principal components is retained to be ≥95% (ensuring that key decision information is not lost). The feature vector is calculated through the covariance matrix, and the top K principal components with the largest variance are selected (K is usually 12-16 dimensions in clinical testing, and can be dynamically adjusted according to the case characteristics). The 6-dimensional original feature vector of each voxel is projected onto the K-dimensional principal component space to obtain the simplified voxel feature vector.
[0130] Then, a sparse state space data structure is constructed, using a dictionary format to store the mapping relationship between voxel space coordinates and simplified feature vectors, storing only voxel information within the reachable space, further reducing storage and computational overhead. After encoding, the dimensionality reduction effect is verified: the cosine similarity of features before and after dimensionality reduction is calculated, requiring ≥0.98, to ensure that key constraint information is not lost during the dimensionality reduction process, while the state space dimension is reduced to the range that reinforcement learning can efficiently process (12-16 dimensions).
[0131] Finally, we verify whether the encoded reinforcement learning state space fully contains the key decision information required for path planning (spatial location, hazard constraints, thermal damage constraints, and dynamic deformation constraints). By filtering redundant information through feature selection, we ensure the effectiveness and efficiency of the state space, and finally obtain a reinforcement learning state space that is complete in information, has reduced dimensions, and is efficient in decision-making.
[0132] The behavior preference reward module 6 is used to extract the expert reward function through expert behavior preference distillation by maximum entropy inverse reinforcement learning.
[0133] In this embodiment, maximum entropy inverse reinforcement learning (MaxEnt IRL) is used to distill implicit expert behavioral preferences from the ablation path data of clinical experts, and to construct an expert reward function that conforms to clinical diagnosis and treatment logic. The extraction of the expert reward function through expert behavioral preference distillation using maximum entropy inverse reinforcement learning includes the following steps:
[0134] S61. Based on the optimality probability distribution of expert paths, construct a maximum entropy inverse reinforcement learning model framework.
[0135] First, the clinical expert path data is standardized, cleaned, and spatially aligned to remove invalid and noisy data. The expert paths are then accurately mapped to the constructed reinforcement learning state space to generate high-quality state-action sequence pairs. This provides reliable demonstration data for the subsequent MaxEntIRL model and meets the specific needs of extracting expert experience for liver cancer thermal ablation paths.
[0136] Specifically, a dual screening standard of effectiveness and safety was constructed: ① Effectiveness screening: invalid paths were removed if they did not reach the target lesion area (distance between the path endpoint and the lesion center > 5 mm, exceeding the effective range of clinical ablation) or if the path covered less than 80% of the lesion volume; ② Safety screening: high-risk paths were removed if they invaded dangerous areas (distance from large blood vessels < 3 mm, distance from small bile ducts < 1.5 mm) or caused postoperative complications such as bleeding / bile leakage; at the same time, data integrity verification was added to ensure that each retained path contained a complete three-dimensional coordinate sequence (sampling interval 1 mm, covering the entire distance from the needle entry point to the lesion endpoint), corresponding ablation parameters (power, time, needle diameter) and postoperative efficacy evaluation results (whether the lesion was completely necrotic) 3 months after the operation.
[0137] Subsequently, path coordinate space alignment was performed. Using the state space coordinate system (preoperative CT image reference coordinate system, resolution 1×1×1mm) as the target coordinate system, affine transformation was used to map the original coordinates of the expert path (from the intraoperative navigation system, which may have coordinate system offset) to the target coordinate system. This was done by selecting more than three liver anatomical landmarks (such as the bifurcation point of the portal vein and the confluence point of the hepatic vein) to establish coordinate mapping relationships, calculating the affine transformation matrix, and transforming the coordinates of each sampling point on the path. After mapping, the alignment accuracy was verified, requiring that the coordinate mapping error of the landmarks be ≤0.5mm and the overlap between the overall path contour and the liver anatomical structure in the preoperative image be ≥0.95.
[0138] Finally, state-action sequence pairs are generated. The mapped path coordinates are decomposed into continuous state nodes according to the sampling order. Each state node corresponds to a voxel feature vector in the state space. The action space is defined as the needle insertion direction of the ablation needle (the angle between the x, y, and z directions) and the step size (a fixed step size of 1 mm to adapt to the precision of clinical operation). The transformation relationship between adjacent state nodes is an action. Finally, continuous state-action-next state sequence pairs are formed for each expert path. They are divided into a distillation dataset and a preliminary validation dataset in an 8:2 ratio for subsequent model training and intermediate validation.
[0139] To address the implicit and generalization requirements of expert decision-making in liver cancer thermal ablation pathways, a maximum entropy inverse reinforcement learning (IRL) model adapted to medical scenarios is constructed. This model uses a feature extractor to focus on key clinical decision features and combines entropy maximization constraints to ensure the adaptability of the reward function to different liver cancer cases, thereby improving the generalization ability of the IRL model. The specific operation is as follows:
[0140] First, the core framework of the model is built. Based on the probabilistic model framework of maximum entropy inverse reinforcement learning, the optimality probability distribution of the expert path is defined. The probability of the expert path is proportional to exp(β·R(s,a)), where β is a temperature parameter that controls the entropy value. The smaller β is, the larger the entropy value and the stronger the generalization. Considering the conservative requirements of the medical scenario, β is set to 0.8, and R(s,a) is the reward function to be learned. The model objective is to maximize the log-likelihood probability of all expert paths, while avoiding overfitting of the reward function to a single case through the entropy maximization constraint.
[0141] Subsequently, a clinically adapted feature extractor was designed, employing a shallow convolutional + fully connected structure: the input is a 12-16 dimensional state feature vector from the reinforcement learning state space. First, a 1×1×1 three-dimensional convolutional kernel (with 32 channels) is used to enhance the spatial feature correlation. Then, two fully connected layers (with 64 and 32 neurons respectively) are used for feature dimensionality reduction and nonlinear mapping. Combining prior clinical screening of key feature dimensions for liver cancer, four core decision features are focused on:
[0142] ① Target area distance characteristics, Euclidean distance from the current state to the lesion center;
[0143] ② Danger zone distance characteristics: the shortest distance from the current state to the nearest large blood vessel / small bile duct;
[0144] ③ Path smoothness feature: the angle between the direction of the current action and the previous action;
[0145] ④ Thermal damage prediction characteristics: the proportion of thermal damage range corresponding to the current path predicted by the thermal field model.
[0146] By using a feature selection matrix, the weights of these four feature types are increased by 1.5 times, thereby enhancing the responsiveness of key clinical decision-making information.
[0147] Finally, the model parameters are initialized. The policy function adopts a Gaussian policy (mean is the predicted action value, variance is set to 0.1 to ensure the rationality of action exploration), the value function is initialized as a 0 vector, and the feature weights of the reward function are initialized using a normal distribution (mean 0, variance 0.01). The model iteration termination condition is set: when the rate of change of the log-likelihood probability of the expert path is <1% for 10 consecutive iterations, the initialization and optimization of the model structure parameters are stopped, and a stable MaxEnt IRL model framework is obtained.
[0148] S62. Optimize the model parameters of the maximum entropy inverse reinforcement learning model framework by gradient descent, and distill the expert reward function containing clinical decision-making logic from the expert path demonstration data.
[0149] First, we construct the mathematical expression of the reward function, using a linear reward function structure:
[0150] ,in The four core features selected are: target area distance, danger area distance, path smoothness, and thermal damage prediction. The weights are the feature weights to be learned (i.e., expert preference weights).
[0151] Gradient descent optimization is then performed, with the objective of maximizing the log-likelihood probability of the expert path, and the weights of the objective function with respect to the reward function are calculated. The gradient is calculated; the Adam optimizer is used, and an L2 regularization term (regularization coefficient set to 0.001) is added during the optimization process to avoid overfitting of the weights; the expert path matching degree (average distance between the predicted path and the real expert path) of the validation dataset is monitored in real time. When the matching degree is ≤2mm and there is no decrease for 15 consecutive batches, the optimization is stopped, and the optimal feature weights are output. .
[0152] Next, the reward function was visualized and preference weights were extracted. A heatmap was used to visualize the distribution of reward values under different states—reward values decreased significantly near the danger zone (negative reward) and increased significantly in the center of the lesion (positive reward), visually verifying the clinical rationality of the reward function. Expert preference was quantified based on the optimized feature weights: the proportion of each weight was calculated, where the distance from the danger zone to the feature weights was... The percentage represents the weight of experts' preference for security, and the weight of the target area distance feature. Weighting of thermal damage prediction features The sum represents the preference weights for efficiency and the path smoothness feature weights. The weighting of surgical efficiency is used as a reference. Clinical tests show that the typical weighting distribution of liver cancer thermal ablation experts is: safety 45%-55%, effectiveness 35%-45%, and efficiency 5%-10%. If the extracted weights deviate from this range, it is necessary to go back to the data preprocessing steps to check the representativeness of the expert path (such as supplementing the expert path for high-risk lesion cases) and re-optimize.
[0153] Finally, the effectiveness and generalization of the expert reward function can be verified through multi-dimensional clinical validation to ensure that it can accurately distinguish between the optimal and suboptimal paths, adapt to the decision-making needs of different types of liver cancer cases, and finally output the expert reward function, which includes a linear mathematical expression and the corresponding expert preference weight matrix.
[0154] The composite reward function module 7 is used to combine the target and danger zone mask, pipeline distance field and expert reward function to obtain the composite reward function through multi-objective adaptive weighting and sparse reward shaping.
[0155] In this embodiment, the step of combining the target and hazard area mask, pipeline distance field, and expert reward function to obtain a composite reward function through multi-objective adaptive weighting and sparse reward shaping includes the following steps:
[0156] S71. Combine the target and hazard area mask and pipeline distance field to construct the objective function for path planning.
[0157] Based on clinical diagnosis and treatment guidelines and surgical practice requirements for liver cancer, three core quantitative objectives and evaluation criteria for pathway planning are defined:
[0158] ① Safety Objective: With the pipeline distance field as the core constraint, the quantitative indicator is the shortest distance from all voxels on the path to the nearest danger zone (large blood vessels, small bile ducts). A graded reward rule is set: when the distance is ≥ the safety threshold of 3mm for large blood vessels / 1.5mm for small bile ducts, the reward value increases linearly with the distance (increase slope 0.2 / mm); when the distance is < the safety threshold, a negative penalty is triggered (the penalty value increases exponentially with the distance, and reaches the maximum value of -50 when the distance is 0), ensuring that the path completely avoids danger zones.
[0159] ②Effectiveness Target: Based on the target area mask, the quantitative indicators include the Euclidean distance between the path endpoint and the lesion center and the proportion of lesion volume covered by the path. Double guarantee of ablation effectiveness: When the distance between the endpoint and the lesion center is ≤5mm (clinically effective ablation range), a full positive reward is given. When the distance is >5mm, the reward value decreases linearly with the increase of distance (attenuation slope 0.3 / mm). When the proportion of lesion volume covered by the path is ≥80%, an additional reward is added. For every 10% decrease in the proportion, the reward value is halved to ensure that the ablation range can completely cover the lesion.
[0160] ③ Efficiency target: The quantitative indicator is the path length from the ablation needle insertion point to the lesion endpoint. A reasonable range is set in combination with the convenience of clinical surgical operation (path length of 10-50mm is the optimal range). A basic reward is given within this range. When the path length is <10mm (which may increase the difficulty of operation due to the steep angle), the reward value decreases linearly. When it is >50mm (which may increase the risk of damage to normal tissue), the reward value decreases linearly with the length (attenuation slope 0.1 / mm), thus balancing surgical efficiency and operational safety.
[0161] S72. Adaptively assign weights to the objective function based on the expert reward function, and construct a sparse reward mechanism by setting intermediate reward or penalty signals at key nodes in the path planning.
[0162] First, the core framework of the adaptive weight allocator is constructed, adopting a three-stage structure of case feature extraction, weight mapping, and dynamic adjustment. The input is the specific feature parameters of liver cancer cases, and the output is the real-time weight ratio of the three objectives.
[0163] ① Case-specific feature extraction: Key feature parameters are extracted from the target and danger area mask and the tube distance field, including: lesion size (divided into small / medium / large categories according to 1-3cm, 3-5cm, and 5-10cm), the shortest distance between the danger area and the lesion (divided into high-risk / medium-risk / low-risk categories according to <5mm, 5-10mm, and >10mm), lesion location (hepatic hilum / non-hepatic hilum), and accessibility of the needle insertion path (whether it is necessary to avoid obstacles such as ribs / intestines), forming a 4-dimensional case feature vector;
[0164] ② Initial weighting benchmark setting: The extracted expert preference weights are used as the initial benchmark. A typical benchmark weight distribution is as follows: safety weight (45%-55% range), effectiveness weight (35%-45% range), efficiency weight (5%-10% range) to ensure that the initial direction of weight allocation conforms to the clinical decision-making logic of experts;
[0165] ③ Dynamic weight adjustment rules: Construct a weight adjustment mapping table based on case feature vectors and set differentiated adjustment strategies:
[0166] a) Adjustment of the association between dangerous areas: If the shortest distance between the dangerous area and the lesion is <5mm (high-risk scenario), the safety weight will be increased to 0.6-0.65, the effectiveness weight will be decreased to 0.3-0.35, and the efficiency weight will remain unchanged at 0.1, prioritizing the safety of the path; if the distance is >10mm (low-risk scenario), the safety weight will be decreased to 0.4-0.45, and the effectiveness weight will be increased to 0.45-0.5, appropriately increasing the priority of effectiveness;
[0167] b) Adjustment based on lesion size: If the lesion volume is >5cm (large lesion), the effectiveness weight is increased to 0.5-0.55, the safety weight is adjusted to 0.4-0.45, and the efficiency weight is 0.05-0.1 to ensure that the ablation range completely covers the large lesion; if the lesion is 1-3cm, the efficiency weight is increased to 0.15-0.2, and the safety and effectiveness weights are reduced to 0.45-0.5 respectively to improve the convenience of surgical operation;
[0168] c) Special location association adjustment: If the lesion is located in the porta hepatis (a densely populated high-risk area), the safety weight will be increased to 0.65-0.7, the effectiveness weight to 0.25-0.3, and the efficiency weight to 0.05, to ensure the safety of the operation to the fullest extent.
[0169] ④ Weight Normalization and Verification: The dynamically adjusted weights need to be normalized (to ensure...) Then, the process is validated through expert paths—the adjusted weights are applied to the corresponding type of expert paths, and the cumulative reward value of the path is calculated. The reward value ranking is required to be consistent with the experts' rating of the path's merits. If they are inconsistent, the parameters of the adjustment rules are fine-tuned until the validation requirements are met.
[0170] Furthermore, by setting intermediate reward or penalty signals at key nodes in path planning, the agent is guided to quickly focus on effective exploration directions, accelerating model convergence. Simultaneously, it is ensured that the setting of intermediate rewards conforms to the clinical operational logic of liver cancer thermal ablation, avoiding misleading the agent into generating non-compliant paths. The construction of a sparse reward mechanism by setting intermediate reward or penalty signals at key nodes in path planning includes the following steps:
[0171] S721. Construct a multi-node hierarchical intermediate reward system, based on the clinical process and constraints of liver cancer pathway planning, and clarify the reward or punishment rules for each node.
[0172] First, compliance of needle insertion points: When the needle insertion point selected by the agent is located in a safe area of the skin (avoiding ribs, blood vessels, and intestines) and can directly reach the surface of the liver, a positive basic reward is given (the reward value is 20% of the effective reward at the end point); if the needle insertion point is located in a dangerous area or cannot directly reach the liver, a negative penalty is given (the penalty value is 30% of the maximum penalty at the end point), thus avoiding invalid needle insertion paths from the source.
[0173] Second, path safety exploration nodes: During the path from the needle insertion point to the lesion, a safety verification node is set every 5mm. If the distance from the voxel at the node to the danger zone is greater than or equal to the safety threshold (3mm for large blood vessels and 1.5mm for small bile ducts), a positive cumulative reward is given (the reward value of each node is 5% of the effective reward at the end point); if the voxel at the node is close to the danger zone (distance < safety threshold but not invaded), a slight negative penalty is given (the penalty value is 10% of the maximum penalty at the end point); if it invades the danger zone, a severe negative penalty is given (the penalty value is 80% of the maximum penalty at the end point) and the current path exploration is terminated, forcibly guiding the agent to avoid the danger zone.
[0174] Third, target area approach node: when the path enters the transition area 5mm outside the lesion, a positive guidance reward is given (the reward value is 30% of the effective reward at the end point); when the path reaches the edge of the lesion, a positive reinforcement reward is given (the reward value is 50% of the effective reward at the end point), guiding the agent to accurately approach the target area.
[0175] S722. Optimize the gradient distribution of reward signals to ensure that the endpoint reward remains the core incentive.
[0176] Optimize the gradient distribution of reward signals, ensuring that the cumulative value of intermediate rewards does not exceed 80% of the effective reward at the endpoint, thus preventing the agent from stagnating at intermediate nodes. At the same time, set a gradient increment rule for penalty signals, where the closer to the danger zone and the more severe the violation, the larger the penalty gradient, thereby enhancing the agent's sensitivity to dangerous constraints.
[0177] S723. By training and verifying the effect of reward shaping, a sparse reward mechanism is obtained.
[0178] Finally, the training process was used to verify the effect of reward shaping. The training convergence speed of the model before and after shaping was compared. It was required that the number of iterations required for the model to reach stable convergence after shaping be reduced by more than 40%, and the proportion of compliant paths generated in the initial exploration phase be increased by more than 50%, so as to ensure the effectiveness of sparse reward shaping.
[0179] S73. Integrate the adaptively weighted multi-objective reward and sparse reward to obtain a composite reward function.
[0180] By systematically integrating adaptive weighted multi-objective rewards with hierarchical intermediate rewards, a composite reward function with both dynamic adaptability and efficient guidance is constructed. The mathematical relationships and weight ratios of each reward component are clearly defined to ensure that the reward function can accurately reflect the clinical decision-making logic, provide clear and reasonable reward signals for the reinforcement learning model, and drive the model to learn the optimal path that meets clinical requirements.
[0181] Specifically, an integrated logic of adaptive weighted multi-objective reward + tiered intermediate reward - violation penalty is adopted to satisfy:
[0182]
[0183] in, These are the quantitative reward values for the safety, effectiveness, and efficiency objectives, respectively (the calculation method is described in the multi-objective definition sub-step). The dynamic weights output by the adaptive weight allocator. The multi-objective reward weighting coefficient is set to 0.6 based on expert experience to ensure the dominant position of the core objective reward. This is the cumulative value of intermediate rewards (including the sum of rewards for compliance at the needle insertion point, rewards for safe path exploration, and rewards for approaching the target area). The intermediate reward weighting coefficient is set to 0.4 to balance the proportion of intermediate rewards and core objective rewards, preventing intermediate rewards from becoming overly dominant. This is the cumulative penalty value for violations (including penalties for approaching dangerous areas, intrusion, and violations at the needle insertion point). The penalty coefficient is set to 1.2 to strengthen the constraint on violations and ensure that the model prioritizes avoiding dangerous operations.
[0184] Two checks need to be performed during the integration process: ① Reward range constraint: the composite reward value Normalize to the [-100, 100] range to avoid model training oscillations caused by extreme reward values -- when >100 is truncated to 100, <-100 is truncated to -100; ② Clinical logical consistency check: Select different types of liver cancer cases (high-risk / low-risk, large / small lesions), substitute the optimal and suboptimal paths marked by experts into the composite reward function, and require the optimal path to be... The value should be at least 30% higher than the second-best path to ensure the reward function accurately distinguishes between superior and inferior paths. If the validation fails, backtrack to the weight allocation or reward shaping step to adjust parameters (e.g., adjust...). The coefficients or reward values of intermediate reward nodes are used until the requirements are met. The final output composite reward function has both dynamic adaptability (through adaptive weights) and efficient guidance (through intermediate rewards). It can be directly used for training subsequent reinforcement learning models, providing accurate reward guidance for learning the optimal clinical path.
[0185] The decision output module 8 is used to combine the reinforcement learning state space, the composite reward function, and the spatiotemporal joint taboo field to solve for the optimal path according to the deep deterministic policy gradient algorithm.
[0186] In this embodiment, the Deep Deterministic Policy Gradient (DDPG) algorithm is employed to achieve continuous spatial optimal solution for the ablation needle path within the reinforcement learning state space, based on the constraints of a composite reward function and a spatiotemporal joint taboo field. The step of combining the reinforcement learning state space, the composite reward function, and the spatiotemporal joint taboo field to solve for the optimal path using the Deep Deterministic Policy Gradient algorithm includes the following steps:
[0187] S81. Construct a deep deterministic policy gradient network structure based on the reinforcement learning state space.
[0188] Specifically, a collaborative architecture of policy network (Actor) and value network (Critic) is adopted, both of which are built on fully connected networks. The input layers are adapted to the feature vectors of the reinforcement learning state space to ensure the consistency and integrity of the input features. The policy network is positioned as an action generator, with the core function of outputting continuous and precise ablation needle operation actions. The network structure is designed as follows: input layer (12-16 dimensional) → hidden layer 1 (128 neurons) → hidden layer 2 (64 neurons) → output layer (4 dimensional). The hidden layers use the LeakyReLU activation function (with a slope of 0.01 to improve gradient propagation stability), and the output layer uses the Tanh activation function to map the action values to a reasonable range. The output 4-dimensional action vector corresponds to the angle between the x, y, and z axes of the needle insertion direction and the step size, respectively. The value network is positioned as an action evaluator, with the core function of quantifying the value of the current state-action pair and providing gradient guidance for policy optimization. The network structure is as follows: joint input layer (state 12-16 dimensional + action 4 dimensional, totaling 16-20 dimensional) → hidden layer 1 (128 neurons) → hidden layer 2 (64 neurons) → output layer (1 dimensional). The hidden layers also use LeakyReLU activation, and the output layer uses a linear activation function to directly output the action value evaluation value.
[0189] Then, a spatiotemporal taboo field constraint module is embedded. To enhance the clinical safety constraints of the network, a dynamic constraint embedding layer is added to the dual network to call the spatiotemporal joint taboo field data in real time. The constraint mechanism of the strategy network is as follows: the spatial coordinates and respiratory phase corresponding to the current action are input into the taboo field to query the risk level. If the action points to a high-risk area (distance from large blood vessels <3mm, small bile ducts <1.5mm), a gradient penalty (penalty coefficient set to 1.5) is applied to the output action through the constraint layer to forcibly correct the action direction to the safe area. The constraint mechanism of the value network is as follows: when calculating the value of the action, the risk level weight of the taboo field is introduced. The value assessment value of the action in the high-risk area decreases linearly according to the risk level (60% reduction in high-risk area, 30% reduction in medium-risk area). The value-guided strategy network avoids dangerous actions.
[0190] Furthermore, network initialization and regularization optimization were performed: To improve the stability and generalization of network training, the weights of both the policy network and the value network were initialized using a He normal distribution, and the bias term was initialized to 0. Dropout layers (with a dropout rate of 0.1 to avoid overfitting) were added to the hidden layers of both networks, and L2 regularization (regularization coefficient 0.001) was introduced into the value network to suppress excessive fluctuations in value evaluation. Finally, the rationality of the network output actions was verified using state feature samples from typical liver cancer cases (high-risk lesions in the porta hepatis, large-volume lesions, and small-volume superficial lesions)—the output needle insertion angle and step length were required to be within the range allowed by clinical operation, and the danger zone constraint module could accurately identify high-risk actions and complete corrections, ensuring that the network architecture conforms to the clinical operation requirements of liver cancer thermal ablation from a design perspective.
[0191] S82. Based on the composite reward function and the spatiotemporal joint taboo field, reinforcement learning training is carried out by using the collaborative design of experience playback and exploration strategy.
[0192] First, an experience replay pool is constructed to store five-dimensional samples generated during training: state s-action a-reward r-next state s'-termination signal done. The reward r is calculated in real-time by a composite reward function. The termination signal done is triggered when the action reaches the lesion center (distance between endpoint and lesion center ≤ 2mm) or when the action invades a high-risk area (triggers severe penalty in the taboo field). To improve the clinical representativeness of the samples, a priority experience replay strategy is adopted—high-reward samples (meeting the characteristics of the optimal clinical path) and high-penalty samples (borderline samples approaching the danger zone) are given higher sampling weights (weight coefficient 1.2-1.5) to ensure the network focuses on learning the decision-making logic of key clinical scenarios. The sample update strategy for the experience replay pool is first-in, first-out, and low-value samples (samples with reward values < -20) are cleaned up every 1000 iterations to ensure the quality of the sample pool.
[0193] To balance the exploratory nature of the initial training phase with the convergence stability of the later phase, a decaying ε-greedy exploration strategy is adopted. The initial ε value is set to 0.3 (to ensure sufficient action diversity to explore safe paths in the early stage). Every 500 iterations, the ε value decays by 0.05 until it decays to 0.01 and is maintained (to retain a small amount of exploratory nature to cope with case differences). The training optimizer is Adam.
[0194] During training, the spatiotemporal joint taboo field data is updated in real time every 100 iterations, and the dynamic displacement of the danger zone caused by the breathing phase change is synchronized to ensure that the network always learns under dynamic constraints. The cumulative reward value of the composite reward function is the core optimization objective, and an alternating update strategy is used to optimize the parameters of the two networks. First, the policy network is fixed, and the mean squared error loss of the value network (the error between the predicted value and the target value) is minimized through gradient descent. Then, based on the value evaluation result of the value network, the cumulative reward expectation of the policy network is maximized through policy gradient ascent.
[0195] To improve training stability, a target network mechanism is introduced. Both the policy network and the value network are set with corresponding target networks. The parameters of the target network adopt a soft update strategy (update rate τ=0.001) to avoid parameter update oscillations.
[0196] Finally, a dual convergence criterion was set: a) the cumulative reward value fluctuation range of 2000 consecutive iterations is <5%; b) the clinical compliance rate of the generated path in the test set (10 liver cancer cases that did not participate in training) is ≥95% (compliance criteria are no dangerous area invasion + lesion coverage ratio ≥85% + path length in the range of 10-50mm).
[0197] During training, three key indicators are monitored in real time: the cumulative reward value change curve, the path compliance rate, and the number of times dangerous areas are invaded. If the reward value decreases or the compliance rate decreases, the exploration strategy ε value or the experience replay sampling weight is adjusted in a timely manner to ensure stable convergence of training.
[0198] This invention is based on a composite reward function and dynamic spatiotemporal constraints. Through the collaborative design of experience playback and exploration strategies, it achieves stable convergence of the DDPG network, allowing the policy network to gradually learn path planning strategies that meet the requirements of clinical safety, effectiveness, and efficiency, while adapting to the individualized constraint differences of different liver cancer cases.
[0199] S83. The optimal path is obtained by optimizing the initial path generated by the policy network after training convergence through post-processing.
[0200] The initial path output by the policy network after training convergence is post-processed and optimized to eliminate defects such as sawtooth fluctuations and endpoint offsets, improve the clinical operability and accuracy of the path, and ensure that the optimized path is fully adapted to the physical operation characteristics of the ablation needle and the requirements of the clinical surgical approach.
[0201] Specifically, the path smoothing optimization begins: To address potential jagged fluctuations in the initial path (caused by subtle differences in motion exploration during training), a 3rd-order Bézier curve fitting smoothing strategy is employed. First, the coordinate sequence of sampling points from the initial path is extracted (sampled at 1mm intervals) and used as control points for the Bézier curve. The curve parameters are optimized using the least squares method, ensuring that the average deviation between the fitted curve and the initial path is ≤0.5mm (ensuring the core path direction remains unchanged). Simultaneously, curvature constraints are added, requiring the maximum radius of curvature of the smoothed path to be ≥5mm (adapting to the physical rigidity of the ablation needle and avoiding operational difficulties caused by excessive bending). If the curvature of a certain segment of the path does not meet the standard, control points are added and the path is refitted until the curvature requirement is met.
[0202] Next, precise endpoint calibration is performed: using the lesion center coordinates as a reference, the accuracy of the initial path endpoints is verified—the Euclidean distance between the endpoints and the lesion center is calculated. If the distance is >2mm (exceeding the clinical ablation precision range), a fine-tuning mechanism is initiated: based on the overall direction vector of the path, the gradient descent method is used to adjust the motion angles of the three sampling points at the end, with each adjustment step being 0.01rad, until the distance between the endpoints and the lesion center is ≤2mm. During the fine-tuning process, it is simultaneously checked whether the path has invaded the danger zone. If the fine-tuning causes the danger zone to be invaded, the path state before the adjustment is traced back, and a combination strategy of segmented fine-tuning + overall smoothing is adopted to balance the endpoint accuracy and path safety.
[0203] Finally, the path length and approach were optimized: the path length was further optimized based on the convenience of clinical surgical operation. If the path length was <10mm (the needle angle was too steep and the operating field of view was limited), the coordinates of the needle entry point were adjusted appropriately (translated within the safe area of the skin, with a translation range of ≤10mm) to extend the path to the 10-15mm range. If the path length was >50mm (increasing the risk of damage to normal tissue), the path direction was optimized and ineffective path segments were shortened (such as avoiding unnecessary detours around normal liver tissue) while ensuring safety and endpoint accuracy.
[0204] The optimized path is validated in three-dimensional space, with indicators including: path smoothness (angle between tangents of adjacent sampling points ≤ 5°), endpoint accuracy (distance from lesion center ≤ 2mm), danger zone avoidance rate (100% avoidance of high-risk areas), and path length (10-50mm). After all indicators meet the standards, the final optimized path is output, including complete needle insertion point coordinates, path sampling point sequence (1mm interval), endpoint coordinates, and needle insertion angle parameters.
[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the present invention.
Claims
1. A liver cancer thermal ablation path planning system based on deep reinforcement learning, characterized in that, Includes the following modules: The target and danger zone mask generation module is used to generate target and danger zone masks; The pipeline distance field construction module is used to construct the pipeline distance field based on parallel refinement and graph structure optimization. The thermal field prediction module is used to obtain the three-dimensional thermal field distribution by rapidly solving the thermal conduction approximation through a physical information neural network. The respiratory deformation prediction module is used to extract respiratory motion features using a spatiotemporal sequence encoding and decoding network; A multi-scale feature fusion module is used to fuse target masks, pipe distance fields, thermal fields, and deformation fields to construct a reinforcement learning state space; The behavior preference reward module is used to extract the expert reward function through expert behavior preference distillation by maximum entropy inverse reinforcement learning. The composite reward function module is used to combine the target and hazard area mask, pipeline distance field and expert reward function to obtain the composite reward function through multi-objective adaptive weighting and sparse reward shaping. The decision output module is used to combine the reinforcement learning state space, the composite reward function, and the spatiotemporal joint taboo field to solve for the optimal path according to the deep deterministic policy gradient algorithm. The generation of the target and danger zone mask includes the following steps: By adjusting adaptive parameters and enhancing the weight of the lesion area, a spatially fully aligned multimodal fusion image is generated. Clinical priors about liver cancer are incorporated into the segmentation network. Using the multimodal fused image as input, initial masks for the target region and danger region are generated through prior knowledge guidance and loss function optimization. The initial mask is post-processed to obtain a high-precision target and danger zone mask; The construction of the pipeline distance field based on parallel refinement and graph structure optimization includes the following steps: Based on the clinical semantic labeling system, the dangerous area masks for pipelines are screened to obtain a pure subset of pipeline masks; Three-dimensional voxels based on pipe masks are used to extract the initial centerline through parallel refinement. The initial centerline is optimized using a graph structure to obtain the optimized complete three-dimensional centerline. Using the optimized pipe centerline as a reference, the shortest distance from the spatial voxel to the pipe is calculated by the rapid travel method, and a quantitative pipe distance field is constructed by combining clinical pipe wall thickness parameters. The method of obtaining the three-dimensional thermal field distribution through the rapid heat conduction approximation solution using a physical information neural network includes the following steps: Construct a physical constraint equation for heat conduction, and embed the loss function of the physical constraint deep into a physical information neural network; The three-dimensional thermal field distribution is obtained by training a neural network based on multi-condition thermal ablation data. The method of extracting respiratory motion features using a spatiotemporal sequence encoding and decoding network includes the following steps: A spatiotemporal sequence coding and decoding network is constructed. Based on the dynamic CT sequence of the respiratory cycle, the spatiotemporal correlation features of respiratory motion in liver cancer are captured by the spatiotemporal sequence coding and decoding network to obtain the initial deformation field. The initial deformation field is regularized and constrained. The optimized deformation field and the static pipe distance field are fused together to generate a spatiotemporal joint taboo field that is dynamically updated with the breathing phase. The fusion of the target mask, pipe distance field, thermal field, and deformation field constructs a reinforcement learning state space, including the following steps: Based on the target mask, pipe distance field, thermal field and deformation field, a multi-scale fused feature map is obtained through hierarchical feature extraction and attention-weighted fusion. Spatial sparse filtering and dimensionality compression encoding are performed on the multi-scale fused feature map to generate a reinforcement learning state space; The process of extracting the expert reward function through expert behavior preference distillation using maximum entropy inverse reinforcement learning includes the following steps: Based on the optimality probability distribution of expert paths, a maximum entropy inverse reinforcement learning model framework is constructed. By optimizing the model parameters of the maximum entropy inverse reinforcement learning model framework through gradient descent, an expert reward function containing clinical decision-making logic is distilled from expert path demonstration data. The method of combining the target and hazard area mask, pipeline distance field, and expert reward function to obtain a composite reward function through multi-objective adaptive weighting and sparse reward shaping includes the following steps: By combining the target and hazard area mask and pipeline distance field, an objective function for path planning is constructed; The objective function is adaptively weighted according to the expert reward function, and a sparse reward mechanism is constructed by setting intermediate reward or penalty signals at key nodes of path planning. By integrating adaptively weighted multi-objective rewards and sparse rewards, a composite reward function is obtained. The method of combining reinforcement learning state space, composite reward function, and spatiotemporal joint taboo field to solve for the optimal path using deep deterministic policy gradient algorithm includes the following steps: Construct a deep deterministic policy gradient network structure based on the reinforcement learning state space; Based on the composite reward function and the spatiotemporal joint taboo field, reinforcement learning training is carried out by the collaborative design of experience replay and exploration strategy; The optimal path is obtained by optimizing the initial path generated by the policy network after training convergence through post-processing.
2. The liver cancer thermal ablation path planning system based on deep reinforcement learning according to claim 1, characterized in that, The method of constructing a sparse reward mechanism by setting intermediate reward or penalty signals at key nodes in path planning includes the following steps: Construct a multi-node hierarchical intermediate reward system, based on the clinical process and constraints of liver cancer pathway planning, and clarify the reward or punishment rules for each node; Optimize the gradient distribution of reward signals to ensure that the endpoint reward remains the core incentive; The sparse reward mechanism is obtained by verifying the effect of reward shaping through training.
Citation Information
Patent Citations
Preoperative ablation area simulation method and device for tumor thermal ablation
CN111973271A
Intraoperative dynamic path optimization method and device for tumor thermal ablation
CN114451990A
Liver tumor ablation path planning method and device
CN115375621A
Reinforced learning training method and system based on multi-target structured reward function
CN121257636A