Intelligent thoracic surgery planning system based on multi-modal image fusion
Through a multi-level feature fusion registration algorithm and a multi-task deep network segmentation architecture, combined with an improved multi-objective optimization strategy, the problems of insufficient registration accuracy and inaccurate anatomical structure segmentation of multimodal images in thoracic surgery are solved, high-precision and safe thoracic surgery planning is achieved, real-time risk assessment and path optimization are provided, and the reliability and efficiency of the surgery are improved.
Patent Information
- Application Number
- CN202510747190.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing multimodal imaging technology in thoracic surgery has problems such as insufficient registration accuracy, low anatomical structure sensitivity, and lack of dynamic risk assessment and real-time interactive adjustment capabilities in surgical path planning. This leads to significant deviations between preoperative planning plans and actual intraoperative needs, affecting the accuracy and safety of the surgery.
It adopts a multi-level feature fusion registration algorithm, a multi-task deep network segmentation architecture and an improved multi-objective optimization strategy. It performs cross-modal registration through the image processing module, segmentes key anatomical structures through the 3D reconstruction module, and generates candidate path plans for risk assessment through the surgical planning module. It also provides real-time risk feedback and visual adjustment through the interactive adjustment module.
It achieves high-precision and high-safety thoracic surgery planning, solves the problems of insufficient accuracy of multimodal image registration and inaccurate anatomical structure segmentation, provides real-time risk assessment and path optimization capabilities, and improves the reliability and efficiency of surgical planning.
Smart Images

Figure CN120656639A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of surgical planning, and in particular to an intelligent planning system for thoracic surgery based on multimodal image fusion. Background Art
[0002] Thoracic surgery is a highly challenging area of clinical surgery due to its complex anatomical environment (such as the dense distribution of lung lobes, blood vessels, bronchi, and nerve plexuses) and stringent requirements for operational precision. Traditional surgical planning relies heavily on the physician's experience, resulting in long preoperative planning times and a high risk of intraoperative accidents. With the development of minimally invasive technology, precise surgical procedures such as thoracoscopy have placed higher demands on lesion localization, instrument path planning, and risk prediction, necessitating an urgent need for intelligent auxiliary systems to improve planning efficiency and safety.
[0003] Currently, multimodal imaging technologies such as CT, MRI, and PET have been gradually applied to thoracic surgery planning, assisting doctors in lesion identification and spatial localization by providing complementary information (such as high-resolution anatomical structure from CT and metabolic activity from PET). However, the heterogeneous nature of multimodal imaging leads to insufficient cross-modality registration accuracy. This is especially true in scenarios with soft tissue deformation and respiratory motion interference. Image fusion errors can easily be transmitted to the 3D reconstruction and path planning stages, affecting the reliability of surgical plans.
[0004] Existing systems generally have problems such as low sensitivity of multimodal registration algorithms to local anatomical features, single surgical path optimization goals that make it difficult to balance safety and efficiency, and lack of real-time risk dynamic feedback mechanism in the interactive interface. Ultimately, there is a significant deviation between preoperative planning plans and actual intraoperative needs, which restricts the realization of precision medicine. Summary of the Invention
[0005] In view of this, the present invention proposes an intelligent planning system for thoracic surgery based on multimodal image fusion. By designing a multi-level feature fusion registration algorithm, a multi-task deep network segmentation architecture and an improved multi-objective optimization strategy, it solves the problems of insufficient accuracy of existing multimodal image registration, low sensitivity of three-dimensional reconstruction to tiny anatomical structures, and lack of dynamic risk assessment and real-time interactive adjustment capabilities in surgical path planning. Ultimately, it achieves the generation and optimization of high-precision and high-safety thoracic surgery planning plans.
[0006] The technical solution of the present invention is achieved as follows:
[0007] The present invention provides an intelligent planning system for thoracic surgery based on multimodal image fusion, comprising:
[0008] An image processing module is used to receive multimodal medical image data, perform preprocessing and cross-modal registration on the medical image data, and generate fused image data;
[0009] A 3D reconstruction module is used to segment key anatomical structures within the chest cavity based on fused image data and construct a 3D model including lesion location, blood vessels, and bronchial structures;
[0010] The surgical planning module is used to generate candidate surgical path plans with risk assessment based on the three-dimensional model and a multi-objective optimization algorithm;
[0011] The interactive adjustment module provides a 3D visual interface for doctors to evaluate, select and modify candidate surgical pathways, determine the final surgical pathway and update risk assessment results in real time;
[0012] The report generation module is used to integrate the final surgical path plan, three-dimensional model and risk assessment results to generate a standardized surgical planning report.
[0013] Preferably, the cross-modal registration of the image processing module adopts a multi-level feature fusion algorithm, including:
[0014] Step 1: Based on the imaging principles and anatomical structure correspondences of the two input modal images, their correlation levels are determined using a predefined modality association rule library. A high-correlation modality combination is used for registration between anatomical imaging modalities; a low-correlation modality combination is used for registration between anatomical and functional imaging modalities.
[0015] Step 2: For high-correlation modal combinations, non-rigid registration is performed using a first metric function, which includes a gradient direction consistency constraint. For low-correlation modal combinations, non-rigid registration is performed using a second metric function, which disables the gradient direction consistency constraint. The deformation field is calculated based on the registration results.
[0016] Step 3: Use B-spline interpolation to constrain the deformation field to smooth and generate the registered fusion image data.
[0017] Preferably, the first metric function is:
[0018] M1=∑ r∈Ω [W(r)·(H(A|r)+H(B|r)-H(A,B|r))]+λ·Φ(A,B)
[0019] Where Ω represents the set of key chest regions defined by a pre-segmentation template or an independent anatomical atlas, including the anatomical divisions of the lung lobes, blood vessels, and bronchi; W(r) is the clinical weight of region r, which is preset according to the anatomical structure type to which the region belongs; H(A|r) is the conditional entropy value of image A within region r; H(B|r) is the conditional entropy value of image B within region r; H(A,B|r) is the joint conditional entropy value of images A and B within region r; Φ(A,B) is the gradient direction consistency function; λ is the dynamic weight coefficient;
[0020] The second metric function is the first metric function with the gradient direction consistency constraint term λ·Φ(A,B) removed.
[0021] Preferably, the gradient direction consistency function Φ(A,B) is:
[0022]
[0023] Where θ A (i,j) and θ B (i, j) is the gradient direction angle of image A and image B at position (i, j), which is obtained by calculating the derivatives of the image in the x and y directions and then taking the inverse tangent; and is the gradient amplitude of image A and image B at position (i, j); cos(θ A (i,j)-θ B (i,j)) is the cosine value between the gradient direction angles of the two images, and its value range is [-1,1].
[0024] Preferably, the dynamic weight coefficient λ is updated in the following manner:
[0025] λ n+1 =λ0·e -k·n +c·ε
[0026] Where λ n+1 represents the dynamic weight coefficient of the next iteration; λ0 is the initial weight coefficient; k is the decay rate coefficient, k0 is the static adjustment factor, σ 2 is the statistical variance of the historical registration error; c is the error sensitivity coefficient; ε is the registration error of the current iteration, calculated by the root mean square error; n represents the number of current iterations.
[0027] Preferably, the 3D reconstruction module uses a multi-task deep network to segment key anatomical structures, including:
[0028] A multi-task deep network was constructed, consisting of a backbone network based on a 3D residual network, a lesion detection subnetwork, and an anatomical boundary localization subnetwork. The backbone network enhanced its feature extraction capabilities for lung lobes, blood vessels, and bronchi by embedding a channel attention mechanism. The lesion detection subnetwork and the anatomical boundary localization subnetwork were deployed in parallel, interacting with each other through a spatial attention mechanism.
[0029] The fused image data is preprocessed to obtain its voxel data, which is then fed into a multi-task deep network. Features are extracted through the backbone network, and candidate lesion regions are extracted using the lesion detection subnetwork. Lesion masks and classification confidence scores are then output. The anatomical boundary localization subnetwork captures vascular and bronchial boundary features and outputs an anatomical boundary probability map.
[0030] The output of the multi-task deep network is processed by voxel-level fusion and three-dimensional reconstruction to generate a three-dimensional model including the lesion location, blood vessel direction and bronchial branching structure.
[0031] Preferably, the voxel-level fusion and 3D reconstruction process includes:
[0032] Perform weighted fusion of the lesion mask, anatomical boundary probability map and the original fused image data to obtain fused voxel data;
[0033] Perform morphological closing operations on the fused voxel data to fill holes and breaks in the segmentation results and obtain optimized voxel data.
[0034] Binarize the optimized voxel data to generate binary voxel labels;
[0035] The Marching Cubes algorithm is used to traverse the binary voxel labels to generate a three-dimensional model including the lesion location, blood vessels and bronchial structures.
[0036] Preferably, the multi-objective optimization algorithm in the surgical planning module includes:
[0037] Three optimization objectives of the surgical path are defined in the 3D model, including: operation distance cost, whose value is determined by the total length of the surgical path; tissue damage cost, whose value is quantified by the volume of sensitive tissue passed by the path; and operation time cost, whose value is calculated based on the path complexity and the preset operation speed model.
[0038] A risk assessment index is introduced as a constraint condition. The risk assessment index is calculated by combining the closest distance between the path and the blood vessels and neural structures and the path curvature parameters, and a maximum risk threshold is set.
[0039] The improved NSGA-II algorithm is used for iterative search to search for non-dominated solution sets in the solution space that meets the maximum risk threshold and generate the Pareto front of candidate surgical paths.
[0040] Preferably, the improved NSGA-II algorithm includes:
[0041] S1. Encode candidate surgical paths as genetically manipulated individuals and generate an initial population that satisfies basic anatomical constraints; calculate three types of optimization objective values for each candidate surgical path; and simultaneously calculate the risk assessment index for each candidate surgical path.
[0042] S2. In each iteration:
[0043] The parent population P t and the offspring population Q t Merge into A t , traverse At Among all candidate solutions, eliminate all those with risk assessment index R(X)>R max The solution, R max is the maximum risk threshold, and the safe solution set S is obtained t ;
[0044] To S t Perform hierarchical sorting to generate non-dominated levels Z1, Z2, ..., Z m , where m is the number of levels and Z1 is the optimal frontier. For the solutions at the same level, the original congestion degree D(X) is calculated, and the risk penalty term is introduced to correct the congestion degree:
[0045]
[0046] Where D′(X) is the modified congestion degree, and α is the risk penalty coefficient;
[0047] From S t The next generation of population is selected by level and corrected crowding;
[0048] S3. Repeat step S2 until the iteration stop condition is met, and output the Pareto optimal solution set from the optimal frontier of the final population as the candidate surgical path plan.
[0049] Preferably, the interactive adjustment module includes:
[0050] The multi-view linkage operation interface simultaneously displays the axial, sagittal, coronal and 3D perspectives of the 3D model, and supports multi-view synchronous zooming, rotation and point selection operations;
[0051] Real-time collision detection mechanism, based on octree spatial hierarchical algorithm and adaptive radius search strategy, calculates the distance between candidate surgical paths and vascular and bronchial structures in real time, and issues warnings when approaching danger thresholds;
[0052] The risk dynamic calculation unit calls the risk assessment function in the surgical planning module to perform immediate risk reassessment of the user-adjusted path and present it numerically.
[0053] The present invention has the following beneficial effects compared to the prior art:
[0054] (1) The present invention constructs a complete intelligent planning technology chain for thoracic surgery through the collaborative design of image processing module, 3D reconstruction module, surgical planning module and interactive adjustment module. The image processing module adopts a multi-level feature fusion registration algorithm to effectively solve the problem of local anatomical structure misalignment caused by differences in imaging principles in multimodal images; the 3D reconstruction module improves the segmentation accuracy of micro-anatomical structures such as lung lobes and blood vessels through a multi-task deep network; the surgical planning module generates candidate path solutions that take into account operation distance, tissue damage and time cost based on an improved multi-objective optimization algorithm; the interactive adjustment module provides real-time risk feedback and visual adjustment functions, forming a closed-loop process from image registration to path optimization;
[0055] (2) To address the problem of insufficient sensitivity of existing registration algorithms to local anatomical features, the present invention defines regional correlations through pre-segmentation templates, combines gradient direction consistency constraints with a dynamic weight adaptation mechanism, and enhances local structure alignment capabilities while retaining global statistical information. The gradient direction consistency function is weighted by the cosine of the gradient direction angle to suppress edge matching errors caused by contrast differences in multimodal images; the dynamic weight coefficient is iteratively updated according to the registration error to balance the contribution ratio of global mutual information and local structure constraints.
[0056] (3) The 3D reconstruction module of the present invention uses a multi-task deep network to enhance the feature response of the backbone network to tubular structures such as blood vessels and bronchi through a channel attention mechanism. At the same time, a spatial attention mechanism is used to achieve feature interaction between the lesion detection and boundary localization sub-networks. This design solves the feature conflict problem caused by target scale differences in traditional single-task models.
[0057] (4) The surgical planning module of the present invention introduces risk constraint pre-screening and risk-aware congestion correction, uses the risk assessment index as a hard constraint, and superimposes a risk penalty term in the congestion calculation of the NSGA-II algorithm. This improvement enables the algorithm to prioritize the elimination of high-risk candidate solutions while preserving the diversity of the Pareto front;
[0058] (5) The interactive adjustment module of the present invention supports doctors to intuitively adjust path nodes on the three-dimensional model through a multi-view linkage operation interface and a real-time collision detection mechanism, and calculates the distance between the path and sensitive structures in real time based on the octree spatial hierarchical algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0060] Figure 1 It is a system framework diagram of the present invention;
[0061] Figure 2 It is a technical implementation diagram of the present invention;
[0062] Figure 3 Schematic diagram of the multi-task deep network framework of the present invention. DETAILED DESCRIPTION
[0063] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0064] like Figure 1 As shown, the present invention provides an intelligent planning system for thoracic surgery based on multimodal image fusion, comprising:
[0065] An image processing module is used to receive multimodal medical image data, perform preprocessing and cross-modal registration on the medical image data, and generate fused image data;
[0066] A 3D reconstruction module is used to segment key anatomical structures within the chest cavity based on fused image data and construct a 3D model including lesion location, blood vessels, and bronchial structures;
[0067] The surgical planning module is used to calculate the safe margin for lesion resection based on the anatomical constraints in the 3D model and generate candidate surgical path plans with risk assessment using a multi-objective optimization algorithm;
[0068] The interactive adjustment module provides a 3D visual interface for doctors to evaluate, select and modify candidate surgical pathways, determine the final surgical pathway and update risk assessment results in real time;
[0069] The report generation module is used to integrate the final surgical path plan, three-dimensional model and risk assessment results to generate a standardized surgical planning report.
[0070] See also Figure 2The technical implementation process of the present invention is as follows: the present invention receives multimodal medical imaging data (such as CT, MRI, PET) through an image processing module, and adopts a multi-level feature fusion algorithm for preprocessing and cross-modal registration, wherein a registration metric function (high correlation modality combination) or a pure mutual information metric function (low correlation modality combination) containing a gradient direction consistency constraint is dynamically selected based on modality correlation, and smooth fused image data is generated through B-spline interpolation; the three-dimensional reconstruction module uses a multi-task deep network to perform voxel-level segmentation on the fused image, and extracts the global features of the lung lobes, blood vessels and bronchi through a 3D residual network embedded with a channel attention mechanism, and adopts a parallel sub-network combined with a spatial attention mechanism to realize the coordinated positioning of lesion detection and anatomical boundaries, and through morphological closing operation and Marching The Cubes algorithm constructs a high-precision three-dimensional model; the surgical planning module defines three optimization objectives on this model: operation distance, tissue damage, and operation time. It uses the improved NSGA-II algorithm to search for Pareto optimal path solutions under risk constraints, and corrects the congestion calculation through a risk penalty factor to prioritize and eliminate high-risk solutions; the interactive adjustment module uses the octree spatial hierarchical algorithm to detect the distance between the path and sensitive structures in real time, providing a multi-view linkage interface for doctors to adjust path nodes and dynamically update risk assessment parameters. Finally, the report generation module integrates the planning plan, model, and risk data to generate a standardized surgical report, forming a full-process closed-loop system from image registration to plan output.
[0071] Specifically, in one embodiment of the present invention, the image processing module is responsible for receiving multimodal medical imaging data (such as CT, MRI, PET, etc.) from different imaging devices, eliminating the resolution and signal differences between modalities through standardized preprocessing, and completing cross-modal registration based on a multi-level feature fusion algorithm to generate high-precision fused image data.
[0072] Specifically, multimodal imaging data is acquired from medical imaging devices (such as CT scanners and MRI devices) through the DICOM protocol. The original data format is the DICOM 3.0 standard, which includes patient position information, imaging parameters, and voxel matrix.
[0073] In this embodiment, the image preprocessing process includes: resampling and spatial alignment: resampling the different modal images to a uniform resolution (default 1×1×1mm 3 Grayscale normalization: Z-score normalization was used to map voxel values across modalities to the range [-1, 1] to eliminate inter-modality signal intensity differences. Background segmentation: For CT images, air background was removed using a thresholding method (threshold range: -1000 to 0 HU); for MRI images, noise regions were removed based on signal intensity (threshold <10%).
[0074] In this embodiment, the cross-modal registration of multimodal images adopts a multi-level feature fusion algorithm, including:
[0075] Step 1: Based on the imaging principles and anatomical structure correspondence of the two input modal images, their correlation level is determined through a predefined modality association rule library, where a high-correlation modality combination is the registration between anatomical imaging modalities; a low-correlation modality combination is the registration between anatomical imaging and functional imaging modalities.
[0076] Specifically, the modal association rule base is constructed based on the following three dimensions:
[0077] Imaging principle similarity: similarity scores are calculated based on the physical imaging mechanisms of different modalities; anatomical structure correspondence: spatial mapping relationships are calculated based on predefined anatomical landmarks; image feature correlation: quantitative evaluation is performed through mutual information coefficient and structural similarity index.
[0078] Specific judgment criteria:
[0079] High correlation: similarity score > 0.7 and mutual information coefficient > 0.5; low correlation: similarity score ≤ 0.7 or mutual information coefficient ≤ 0.5.
[0080] Step 2: For high-correlation modal combinations, non-rigid registration is performed using a first metric function, which includes a gradient direction consistency constraint. For low-correlation modal combinations, non-rigid registration is performed using a second metric function, which disables the gradient direction consistency constraint. The deformation field is calculated based on the registration results.
[0081] The first metric function is used for high-correlation mode combinations, and its mathematical expression is:
[0082] M1=∑ r∈Ω [W(r)·(H(A|r)+H(B|r)-H(A,B|r))]+λ·Φ(A,B)
[0083] where Ω represents the set of key thoracic regions defined by a pre-segmentation template (e.g., lung lobe, vascular ROI) or an independent anatomical atlas (e.g., the MNI standard brain atlas), including the anatomical divisions of lung lobes, blood vessels, and bronchi; W(r) is the clinical weight of region r, which is preset according to the anatomical structure type to which the region belongs. Specifically, the weight of the lung lobe region is set to 1.0, the weight of the vascular dense region is increased to 1.5, and the weight of the bronchial branch is 0.8; H(A|r) is the conditional entropy value of image A within region r; H(B|r) is the conditional entropy value of image B within region r; H(A,B|r) is the joint conditional entropy value of images A and B within region r; Φ(A,B) is the gradient direction consistency function; and λ is the dynamic weight coefficient.
[0084] The gradient direction consistency function Φ(A,B) is:
[0085]
[0086] Where θ A (i,j) and θ B (i, j) is the gradient direction angle of image A and image B at position (i, j), which is obtained by calculating the derivatives of the image in the x and y directions and then taking the inverse tangent; and is the gradient amplitude of image A and image B at position (i, j); cos(θ A (i,j)-θ B (i, j) is the cosine value between the gradient angles of the two images, ranging from -1 to 1. Specifically, the gradient angle can be obtained by calculating the x and y derivatives using the Sobel operator and then taking the inverse tangent.
[0087] The second metric function is used for low correlation combinations (such as CT-PET), disabling the gradient direction consistency constraint, and is expressed as:
[0088] M2=∑ r∈Ω [W(r)·(H(A|r)+H(B|r)-H(A,B|r))]
[0089] In order to make the registration process more stable and accurate, the dynamic weight coefficient λ is updated as follows:
[0090] λ n+1 =λ0·e -k·n +c·ε
[0091] Where λ n+1 represents the dynamic weight coefficient of the next iteration; λ0 is the initial weight coefficient; k is the decay rate coefficient, k0 is the static adjustment factor, σ 2 is the statistical variance of the historical registration error; c is the error sensitivity coefficient; ε is the registration error of the current iteration, calculated by the root mean square error; n represents the number of current iterations.
[0092] Specifically, in the non-rigid registration algorithm, a complete iteration represents a closed-loop process that completes the following operations:
[0093] 1. Parameter update: based on the current deformation field T n and weight λ n , calculate the registration objective function value M.
[0094] 2. Optimization direction search: Through strategies such as gradient descent and genetic algorithms, find the deformation field increment ΔT that minimizes the objective function.
[0095] 3. Deformation field update: Apply the increment to get the new deformation field T n+1 =T n +ΔT.
[0096] 4. Error calculation: Calculate the updated registration error ε.
[0097] 5. Weight coefficient adjustment: Calculate the next round of weight λ based on the error ε and the update formula n+1 .
[0098] 6. Convergence judgment: Check the termination conditions (such as ε < threshold or number of iterations exceeds the limit) and decide whether to continue iterating.
[0099] The registration in this embodiment is divided into an initial stage and a later stage, and a dynamic weight coefficient λ is introduced to adapt to the accuracy requirements of different registration stages.
[0100] Step 3: Use B-spline interpolation to constrain the deformation field to smooth and generate the registered fusion image data.
[0101] Specifically, uniform cubic B-spline interpolation is used to smooth the deformation field. The control point grid is initially set to a uniform distribution with a control point spacing of 10 mm. As the registration iterations proceed, the control point grid is gradually refined, eventually reaching a control point spacing of 5 mm to achieve more precise local deformation control. During the deformation field optimization process, anatomical structure preservation constraints are introduced to ensure that key anatomical structures such as lung lobe boundaries, major blood vessel directions, and bronchial branches maintain morphological integrity during the registration process.
[0102] Based on the optimized deformation field, the images of different modalities are transformed into the same coordinate system, and then fused image data is generated. The specific fusion strategy is as follows: For highly correlated modal combinations, a weighted average fusion strategy is adopted, with weights dynamically assigned based on the modality's ability to display different anatomical structures. For low-correlation modal combinations, a feature-level fusion strategy is adopted, first extracting the salient features of each modality and then fusing them in the feature domain. The registered fused image is output in NIfTI format, a dual-channel matrix containing the original CT / MRI voxel values and PET metabolic activity values.
[0103] Specifically, in one embodiment of the present invention, the three-dimensional reconstruction module is responsible for accurately segmenting the key anatomical structures in the chest cavity based on the fused image data generated by the image processing module, and constructing a detailed three-dimensional model that includes the lesion location, blood vessels, and bronchial structures. This module adopts a multi-task deep learning network architecture and achieves high-precision recognition and segmentation of complex anatomical structures in the chest cavity by processing different anatomical tasks in parallel. Compared with traditional single-task segmentation methods, this module can simultaneously focus on two interrelated tasks: lesion detection and anatomical boundary positioning, and enhance the segmentation effect through feature sharing and interaction.
[0104] like Figure 3 As shown in the figure, the multi-task deep network used in this embodiment is an end-to-end 3D medical image segmentation framework consisting of three main components: a backbone network based on a 3D residual network, a lesion detection subnetwork, and an anatomical boundary localization subnetwork. The backbone network embeds a channel-wise attention mechanism to enhance feature extraction for lung lobes, blood vessels, and bronchi. The lesion detection subnetwork and the anatomical boundary localization subnetwork are deployed in parallel, interacting with each other through a spatial attention mechanism.
[0105] The backbone network is based on a modified 3D residual network (ResNet) architecture and consists of five encoding stages, each consisting of two residual blocks. Each residual block adopts a "convolution-batch normalization-ReLU-convolution-batch normalization-ReLU" structure, and introduces skip connections to alleviate the vanishing gradient problem of deep networks. The network initially uses 32 3×3×3 convolution kernels. As the network depth increases, the number of feature channels increases to 32, 64, 128, 256, and 512, respectively. To enhance feature extraction of different anatomical structures such as lung lobes, blood vessels, and bronchi, a channel attention mechanism is embedded after each encoding stage. This mechanism generates channel descriptors through global average pooling and maximum pooling. After processing by a shared multi-layer perceptron, channel weights are generated using a sigmoid function, and the feature maps are adaptively recalibrated to highlight channel features related to the target anatomical structure. The mathematical expression of the channel attention mechanism is:
[0106] M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))
[0107] Where F is the input feature map, AvgPool and MaxPool represent global average pooling and maximum pooling operations, respectively, MLP is a shared multi-layer perceptron, and σ is the sigmoid activation function. Through this mechanism, the network can adaptively enhance the feature representation of different anatomical structures, improving segmentation accuracy.
[0108] The lesion detection subnetwork adopts a U-Net-like decoder architecture and consists of four decoding stages. Each decoding stage consists of an upsampling operation and two consecutive 3D convolutional blocks. A skip connection is introduced to the feature map from the corresponding encoding stage in the backbone network to restore spatial detail lost during the encoding process. Finally, a 1×1×1 convolutional layer projects the feature map into a lesion probability map, representing the probability that each voxel belongs to a lesion. To improve the detection capability of lesions of varying sizes and shapes, the lesion detection subnetwork also incorporates a multi-scale feature fusion module. This module upsamples feature maps from different decoder layers to the same resolution, performs channel alignment via 1×1×1 convolutions, and then performs weighted fusion to generate a multi-scale-aware lesion feature map.
[0109] The anatomical boundary localization subnetwork also uses a U-Net-like decoder structure but is specifically optimized for slender structures such as blood vessels and bronchi. This subnetwork introduces a boundary enhancement module at each decoding stage. By combining 3D dilated convolutions with different dilation rates, it expands the receptive field while maintaining computational efficiency. Furthermore, this subnetwork employs a deep supervision strategy, adding auxiliary loss functions at different decoding stages to guide the network in learning multi-scale boundary features. Finally, a multi-channel 1×1×1 convolutional layer projects the feature map into a multi-category anatomical boundary probability map, including segmentation results for key structures such as arteries, veins, and bronchi.
[0110] To achieve feature interaction and complementary enhancement between the two sub-networks, the system introduces a spatial attention mechanism as a bridge. This mechanism first generates a spatial attention map from the intermediate feature maps of the lesion detection and anatomical boundary localization sub-networks:
[0111] M s (F1,F2)=σ(Conv([AvgPool(F1,F2);MaxPool(F1,F2)]))
[0112] Here, F1 and F2 are the feature maps of the two sub-networks, AvgPool and MaxPool perform average pooling and maximum pooling, respectively, in the channel dimension. [;] represents channel concatenation, Conv is the convolution operation, and σ is the sigmoid activation function. This mechanism allows the two sub-networks to focus on each other's spatial regions of interest, enhancing feature complementarity. For example, the lesion detection sub-network can use the location information of blood vessels and bronchi to assist in determining the extent of lesion invasion, while the anatomical boundary localization sub-network also benefits from lesion area information to improve the accuracy of vessel and bronchi segmentation around the lesion.
[0113] The training of multi-task deep network adopts end-to-end joint optimization strategy and designs a combined loss function to balance the learning objectives of different tasks. The total loss function consists of three parts: lesion detection loss L lesion, anatomical boundary localization loss L boundary and inter-task association loss L relation , expressed as:
[0114] L total =α·L lesion +β·L boundary +γ·L relation
[0115] Among them, α, β, and γ are weight coefficients, and their initial values are set to 0.4, 0.4, and 0.2 respectively. They are dynamically adjusted according to the difficulty of the task during training. lesion A combination of Focal Loss and Dice Loss is used to deal with the serious category imbalance problem; anatomical boundary positioning loss L boundary A combination of weighted cross entropy and boundary-sensitive Dice Loss is used to enhance the learning of the boundaries of slender structures such as blood vessels and bronchi; the inter-task association loss L relation It is designed as the mutual information loss of the attention maps of the two sub-networks to promote collaborative learning between sub-networks.
[0116] The network training uses the Adam optimizer, and the initial learning rate is set to 10 -3 , and dynamically adjusted using a cosine annealing strategy. To prevent overfitting, in addition to standard L2 regularization, data augmentation was implemented, including random rotation, random scaling, random brightness and contrast adjustments, and random Gaussian noise. The training batch size was set to 4, and a total of 300 iterations were performed.
[0117] To improve the generalization ability of the model, the system adopts a two-stage transfer learning strategy: first, pre-training is performed on a large-scale chest CT dataset to master common chest anatomical features; then fine-tuning is performed on multimodal fusion data of specific tasks to adapt to specific clinical scenario requirements.
[0118] The 3D reconstruction data processing process is as follows:
[0119] First, the fused image data was preprocessed, including intensity normalization, voxel spacing adjustment, and image cropping. Intensity normalization employed a percentile-based linear transformation to adjust the image intensity distribution to the range [-1, 1]. Voxel spacing was uniformly adjusted to 1 mm × 1 mm × 1 mm. Image cropping involved threshold segmentation and morphological operations to identify the thoracic region of interest, minimizing computational resource consumption.
[0120] The preprocessed fused image data is fed into a multi-task deep network using a sliding window configuration with a window size of 96×96×96 voxels and a stride of 48 voxels. Adjacent windows are guaranteed to have 50% overlap to eliminate boundary effects. The predictions for each window are fused using a weighted average, with the weight inversely proportional to the distance from the voxel to the window center, to improve prediction reliability for the central region. Features extracted from the backbone network are fed into two subnetworks: the lesion detection subnetwork outputs a lesion mask and classification confidence score, while the anatomical boundary localization subnetwork outputs an anatomical boundary probability map, including segmentation results for arteries, veins, and bronchi.
[0121] The output of the multi-task deep network is processed through voxel-level fusion and 3D reconstruction to generate a 3D model that includes the lesion location, blood vessel direction, and bronchial branching structure. The specific process is as follows:
[0122] The lesion mask, anatomical boundary probability map and original fused image data are weightedly fused to obtain fused voxel data. The weighted fusion process assigns different weights to each anatomical structure category to dynamically balance the original image information and segmentation results. The fusion formula is expressed as:
[0123]
[0124] Among them, V fusion (x, y, z) is the voxel value after fusion, V0(x, y, z) is the voxel value of the original fused image, V lesion (x,y,z) is the lesion mask voxel value, is the voxel value of the probability map of the anatomical structure of type i, α, β, γ i is the weight coefficient, and α+β+∑γ i =1.
[0125] Morphological closing is performed on the fused voxel data to fill holes and breaks in the segmentation results, resulting in optimized voxel data. Closing uses a spherical structuring element with a radius of 2 voxels, treating different anatomical structures separately. For small blood vessels and bronchial structures, the system additionally applies connectivity analysis to remove small isolated regions (less than 50 voxels in volume) that may be caused by noise.
[0126] The optimized voxel data is binarized to generate binary voxel labels. An adaptive thresholding method is used, with different thresholding strategies applied to different anatomical structures: a fixed threshold of 0.5 is used for lesion areas; an adaptive threshold based on the Otsu algorithm is used for vascular structures; and a hybrid thresholding method combining region growing and morphological constraints is used for bronchial structures. This differentiated thresholding strategy better adapts to the characteristics of different anatomical structures and improves the binarization effect.
[0127] The Marching Cubes algorithm is used to traverse the binary voxel labels to generate a three-dimensional model that includes the lesion location, blood vessels, and bronchial structures. Specifically, the standard Marching Cubes algorithm determines how the surface passes through the unit by checking the eight vertex values of each unit in the voxel grid, and generates a triangular mesh to represent the target surface. The present invention makes adaptive changes to the Marching Cubes algorithm, including: first, introducing an adaptive sampling strategy to use denser sampling in high curvature areas to improve the reconstruction accuracy of complex structures; second, applying surface smoothing and simplification processing to reduce grid noise and control model complexity to improve rendering efficiency. The three-dimensional reconstructed model sets differentiated rendering parameters for different structures: lesions are displayed in translucent red; arteries are bright red; veins are dark blue; and bronchi are light yellow to enhance visual distinction.
[0128] Specifically, in one embodiment of the present invention, the surgical planning module performs automated intelligent planning for thoracic surgery based on the precise 3D anatomical model generated by the 3D reconstruction module. Using an improved multi-objective non-dominated sorting genetic algorithm II (NSGA-II), the algorithm simultaneously optimizes three competing objectives—surgical distance, tissue damage, and operative time—while satisfying safety constraints, generating a series of candidate surgical path plans with risk assessment.
[0129] In a specific example of the present invention, before using the multi-objective optimization algorithm to generate the surgical path, the surgical planning module can first determine its safe resection margin based on the characteristics of the lesion. This embodiment uses a combination of medical imaging features and clinical experience models to calculate the safe margin:
[0130] For benign lesions, morphological analysis is used to determine the safety margin. First, the 3D contour of the lesion is extracted, and its convex hull is calculated based on the lesion segmentation results. Then, based on the relationship between the actual lesion boundary and the convex hull, an adaptive expansion algorithm is applied to generate the safety margin. The mathematical expression of the expansion algorithm is: safe =δ r (B lesion ); among them, B lesion is the binary mask of the lesion, δ r is the morphological dilation algorithm, and r is the dilation radius.
[0131] For malignant lesions, the safety margin is calculated in combination with the pathological statistical model. Based on the multimodal features extracted from the fused images, the invasiveness and diffusion tendency of the lesions are first evaluated. The evaluation indicators include image intensity heterogeneity, boundary sharpness, enhancement characteristics, etc. The evaluation results are used to adjust the basic safety distance to ensure that highly invasive lesions obtain a larger safety margin. The calculation formula for the lesion invasiveness score ISC is: ISC = 0.3×UI+0.25×BS+0.35×EC+0.1×SC, where UI is the intensity heterogeneity score, BS is the boundary sharpness score, EC is the enhancement characteristic score, and SC is the shape complexity score. Based on the invasiveness score, the safety margin is calculated as: r m =(10 mm + ISC × 10 mm) × TF, where TF is the tissue type adjustment factor.
[0132] The calculated safety margin is displayed as a semi-transparent shell in the 3D model. As a basic constraint for surgical path planning, any resection path must completely cover this safety margin.
[0133] The multi-objective optimization algorithms in the surgical planning module include:
[0134] Three optimization objectives for the surgical path are defined in the 3D model, including: operation distance cost, whose value is determined by the total length of the surgical path; tissue damage cost, whose value is quantified by the volume of sensitive tissue traversed by the path; and operation time cost, whose value is calculated based on the path complexity and a preset operation speed model. In this embodiment, the three optimization objective functions are as follows:
[0135]
[0136] C dam =∑V tissue ×W tissue
[0137]
[0138] Where C dis is the operation distance cost, p i represents the coordinates of the i-th control point on the path, d(p i ,p i+1 ) represents the Euclidean distance, n is the total number of path control points, calculate C dis Finally, in order to balance the differences in lesions of different sizes, normalization processing is required: C dis-norm =C dis / (2×R lesion +30), R lesion is the equivalent radius of the lesion, that is, the radius of a sphere with the same volume as the lesion; C dam The cost of tissue damage, is the volume of various tissues that the path passes through, r instru is the radius of the surgical instrument, L path is the path length through the tissue, K comp is the path complexity correction coefficient, which is related to the path curvature and is calculated as the ratio of the curvature integral to the path length plus 1, W tissue is the damage weight factor of the corresponding tissue; C time is the operation time cost, t base The default setting is 5 minutes. i ) is the rotation angle θ of the device i The operating speed at the position is calculated according to the empirical formula: v(θ i )=v max ×(1-0.6×θ i / π), v max is the maximum operating speed, t pause (θ i ) is the pause time at the corner, which is used to simulate the adjustment time of the surgeon when turning, and is calculated as: t pause (θ i )=max(0,2×(θ i -π / 6)), in seconds, when θ i Dwell time is only included when the angle is >π / 6 radians.
[0139] A risk assessment index is introduced as a constraint condition. The risk assessment index is calculated jointly by the closest distance between the path and the blood vessels and neural structures and the path curvature parameters, and a maximum risk threshold is set.
[0140] Specifically, the risk assessment index is calculated as follows:
[0141] R(X)=w d ×R dis (X)+w c ×R curv (X)+w a ×R anat (X)
[0142] Among them, R dis (X) is the distance risk, R curv (X) is the curvature risk, R anat (X) is the anatomical risk, w d 、w c 、w a These are the weight coefficients for distance risk, curvature risk, and anatomical risk, which are fine-tuned by clinical experts based on the type of surgery and individual patient conditions.
[0143] Specifically, d jis the minimum distance between the path and the jth key structure, is the safety distance threshold of the corresponding structure, l j is the importance index of the structure. Different safety distance thresholds are assigned to the blood vessel diameter and bronchial diameter in path planning.
[0144] R curv (X)=max i k(p i ) / k max +σ κ / σ ref , k(p i ) is the path at point p i The curvature at max is the maximum acceptable curvature, σ κ is the standard deviation of the path curvature, σ ref The curvature is calculated using the three-point method.
[0145] V k is the volume of the kth anatomical region passed by the path, W k is the risk weight of the region, V total is the total volume of the pathway. The risk weights of different anatomical regions are determined based on their clinical importance.
[0146] In this embodiment, the risk assessment index R(X) ranges from [0,1], and the maximum risk threshold R is set. max is 0.7.
[0147] The improved NSGA-II algorithm is used for iterative search to search for non-dominated solution sets in the solution space that meets the maximum risk threshold and generate the Pareto front of candidate surgical paths.
[0148] The improved NSGA-II algorithm includes:
[0149] S1. Encode candidate surgical paths as genetically manipulated individuals and generate an initial population that satisfies basic anatomical constraints; calculate three types of optimization objective values for each candidate surgical path; and simultaneously calculate the risk assessment index for each candidate surgical path.
[0150] Specifically, the encoding method can adopt a control point representation method based on B-spline curves. During encoding, boundary constraints are imposed on the coordinates of each control point. The boundary value is automatically determined according to the boundary of the patient's chest anatomical structure, and a safety margin of 5 mm is reserved.
[0151] The initial population is generated using a constrained random sampling method to ensure that the initial path meets the basic anatomical constraints. In one example, the initial population generation process is as follows: 1. Determine the population size N pop, which is set to 100 by default. 2. For each individual, fix the starting point and end point, and use heuristic rules to generate transition points p2 and p m-1 : p2 is located 10-15 mm from the starting point along the direction of the lesion, p m-1 Located 8-12mm outward from the end point. 3. Use random walk algorithm to generate intermediate control points p3 to p m-2 , randomly sample in the space that meets the anatomical constraints. 4. Check whether the generated path passes through the prohibited area (such as large blood vessels, main bronchi, etc.), and regenerate if it does. 5. Apply the cubic B-spline interpolation algorithm to convert the control point sequence into a smooth curve path. 6. Repeat steps 2-5 until N pop A valid initial path.
[0152] In addition, 5-8 commonly used surgical approach templates can be preset according to the location of the lesion, such as anterior approach, lateral approach, posterior approach through the chest wall, etc. These template paths can be added to the initial population to guide the optimization process.
[0153] The complete process of the initialization phase is: 1. Generate an initial population P0 that meets the basic anatomical constraints, with a population size of N pop 2. For each candidate path X in the initial population, calculate three types of optimization target values: C dis (X), C dam (X) and C time (X). 3. Calculate the risk assessment index R(X) for each candidate path and eliminate those with R(X)>R max Path, the insufficient part is supplemented by random generation. 4. Use non-dominated sorting to stratify the initial population and calculate the congestion. 5. Generate the first generation of offspring population Q0 through binary tournament selection, with the same size of N pop .
[0154] S2. In each iteration t:
[0155] The parent population P t and the offspring population Q t The formation scale is 2N pop The combined population A t , traverse A t Among all candidate solutions, eliminate all those with risk assessment index R(X)>R max The solution, R max is the maximum risk threshold, and the safe solution set S is obtained t If S t Scale smaller than N pop , additional secure solutions are generated through mutation operations.
[0156] To S t Perform hierarchical sorting to generate non-dominated levels Z1, Z2, ..., Zm , where m is the number of levels and Z1 is the optimal frontier; for the solutions at the same level, calculate their original congestion degree D(X), Among them, f i represents the i-th objective function, X next and X prev are the next and previous solutions of X after the objective function is sorted, respectively, i max and f i min are the maximum and minimum values of the objective function respectively. A risk penalty term is introduced to correct the congestion:
[0157]
[0158] Where D′(X) is the modified congestion degree and α is the risk penalty coefficient. The modified congestion degree makes the riskier solutions have lower selection priority even in the same non-dominated level.
[0159] From S t Select N by level and modified congestion pop The next generation of parent population P t+1 :First, select the complete non-dominated levels Z1, Z2, ... in turn until Z is selected j After exceeding N pop For the last level Z that requires partial selection j , select in descending order according to the modified congestion degree D′(X), and fill it up to N pop .
[0160] P t+1 Perform genetic operations to generate offspring population Q t+1 :
[0161] Crossover operation: simulated binary crossover (SBX) is used, with a crossover probability p c =0.85, distribution index η c =20.
[0162] Mutation operation: polynomial mutation is used, with mutation probability p m =1 / m, m is the number of control points, distribution index η m =30.
[0163] Q t+1 For each solution in , B-spline interpolation is applied to generate a smooth path, and the objective function value and risk assessment index are recalculated.
[0164] S3. Repeat step S2 until the iteration stop condition is met, and output the Pareto optimal solution set from the optimal frontier of the final population as the candidate surgical path plan.
[0165] Specifically, the iteration stopping condition is: reaching the maximum number of iterations T max or continuous G conv There is no significant improvement in the frontier solution set.
[0166] From the final population P T Representative solutions from the optimal frontier Z1 are selected as candidate surgical path solutions. A maximum coverage strategy based on congestion is selected to ensure that the solution sets are evenly distributed in the target space. Each candidate path is smoothed and fine-tuned, and risk assessment details and optimization targets are calculated for each candidate path.
[0167] Specifically, in one embodiment of the present invention, the interactive adjustment module provides clinicians with a 3D visualization environment for evaluating, selecting, and revising system-generated candidate surgical pathways. This module utilizes a multi-view linkage design, simultaneously displaying axial, sagittal, coronal, and 3D perspectives, enabling clinicians to fully understand the spatial relationship between the patient's anatomy and the planned pathway.
[0168] The module integrates a real-time collision detection mechanism. As the surgeon adjusts the surgical path, the system calculates the minimum distance between the path and surrounding anatomical structures in real time. This mechanism uses an octree spatial index for efficient querying. When the path approaches critical structures, the system issues warnings through color changes and audible cues. Collision detection has three levels of response: safe distance (green), warning distance (yellow), and danger distance (red alert), with thresholds adjustable based on the type of surgery.
[0169] The dynamic risk calculation unit performs instant risk reassessment whenever the physician adjusts the path. The system employs an incremental calculation strategy, recalculating only the affected path segments to ensure smooth interaction. Risk assessment results are intuitively presented through path color coding, numerical dashboards, and 3D heat maps.
[0170] The module provides a variety of path adjustment tools, including dragging control points, path translation and rotation, setting mandatory points, and marking prohibited areas. Local optimization tools allow algorithmic optimization of the remaining path while retaining key points specified by the doctor, enabling collaborative human-machine decision-making.
[0171] Specifically, in one embodiment of the present invention, the report generation module is responsible for integrating key information from the surgical planning process into a standardized surgical planning report, providing a comprehensive reference for preoperative discussions and intraoperative guidance. This module utilizes a template-driven report generation architecture that supports customized output formats to meet the document standardization requirements of different hospitals and departments.
[0172] The report automatically captures key data from each module of the system, including basic patient information, lesion 3D reconstruction data, safety margin calculation results, optimal surgical path, and its risk assessment indicators. The report body is divided into four main sections: Patient and Lesion Overview, Preoperative 3D Reconstruction Results, Surgical Pathway Planning Scheme, and Risk Assessment Analysis. Each section contains a text description and visual images. The surgical path planning scheme automatically generates multi-angle screenshots of key anatomical structures and overlays surgical path and safety margin markers. The risk assessment analysis section includes quantitative risk indicators and color-coded risk distribution maps to intuitively display potential risk areas.
[0173] The report generation process supports interactive editing by physicians, allowing for personalized annotations and supplementary explanations. The system also offers intelligent content suggestions, providing reference text for the current report based on historical reports of similar cases. Reports are organized into hierarchical levels of detail, with the main report presenting core information and appendices containing detailed technical specifications and complete datasets.
[0174] Report output supports multiple formats, including medical standard DICOM SR format, PDF documents and HTML web pages, which facilitates review and sharing in different scenarios.
[0175] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. Intelligent planning system for thoracic surgery based on multimodal image fusion, characterized by: include: An image processing module is used to receive multimodal medical image data, perform preprocessing and cross-modal registration on the medical image data, and generate fused image data; A 3D reconstruction module is used to segment key anatomical structures within the chest cavity based on fused image data and construct a 3D model including lesion location, blood vessels, and bronchial structures; The surgical planning module is used to generate candidate surgical path plans with risk assessment based on the three-dimensional model and a multi-objective optimization algorithm; The interactive adjustment module provides a 3D visual interface for doctors to evaluate, select and modify candidate surgical pathways, determine the final surgical pathway and update risk assessment results in real time; The report generation module is used to integrate the final surgical path plan, three-dimensional model and risk assessment results to generate a standardized surgical planning report.
2. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 1, characterized in that: The cross-modal registration of the image processing module uses a multi-level feature fusion algorithm, including: Step 1: Based on the imaging principles and anatomical structure correspondences of the two input modal images, their correlation levels are determined using a predefined modality association rule library. A high-correlation modality combination is used for registration between anatomical imaging modalities; a low-correlation modality combination is used for registration between anatomical and functional imaging modalities. Step 2: For high-correlation modal combinations, non-rigid registration is performed using a first metric function, which includes a gradient direction consistency constraint. For low-correlation modal combinations, non-rigid registration is performed using a second metric function, which disables the gradient direction consistency constraint. The deformation field is calculated based on the registration results. Step 3: Use B-spline interpolation to constrain the deformation field to smooth and generate the registered fusion image data.
3. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 2, characterized in that: The first metric function is: M1=∑ r∈Ω [W(r)·(H(A|r)+H(B|r)-H(A,B|r))]+λ·Φ(A,B) Where Ω represents the set of key chest regions defined by a pre-segmentation template or an independent anatomical atlas, including the anatomical divisions of the lung lobes, blood vessels, and bronchi; W(r) is the clinical weight of region r, which is preset according to the anatomical structure type to which the region belongs; H(A|r) is the conditional entropy value of image A within region r; H(B|r) is the conditional entropy value of image B within region r; H(A,B|r) is the joint conditional entropy value of images A and B within region r; Φ(A,B) is the gradient direction consistency function; λ is the dynamic weight coefficient; The second metric function is the first metric function with the gradient direction consistency constraint term λ·Φ(A,B) removed.
4. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 3 is characterized in that: The gradient direction consistency function Φ(A,B) is: Where θ A (i,j) and θ B (i, j) is the gradient direction angle of image A and image B at position (i, j), which is obtained by calculating the derivatives of the image in the x and y directions and then taking the inverse tangent; and is the gradient amplitude of image A and image B at position (i, j); cos(θ A (i,j)-θ B (i,j)) is the cosine value between the gradient direction angles of the two images, and its value range is [-1,1].
5. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 3 is characterized in that: The dynamic weight coefficient λ is updated as follows: l n+1 =λ0·e -k·n +c·e Where λ n+1 represents the dynamic weight coefficient of the next iteration; λ0 is the initial weight coefficient; k is the decay rate coefficient, k0 is the static adjustment factor, σ 2 is the statistical variance of the historical registration error; c is the error sensitivity coefficient; ε is the registration error of the current iteration, calculated by the root mean square error; n represents the number of current iterations.
6. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 1, characterized in that: The 3D reconstruction module uses a multi-task deep network to segment key anatomical structures, including: A multi-task deep network was constructed, consisting of a backbone network based on a 3D residual network, a lesion detection subnetwork, and an anatomical boundary localization subnetwork. The backbone network enhanced its feature extraction capabilities for lung lobes, blood vessels, and bronchi by embedding a channel attention mechanism. The lesion detection subnetwork and the anatomical boundary localization subnetwork were deployed in parallel, interacting with each other through a spatial attention mechanism. The fused image data is preprocessed to obtain its voxel data, which is then fed into a multi-task deep network. Features are extracted through the backbone network, and candidate lesion regions are extracted using the lesion detection subnetwork. Lesion masks and classification confidence scores are then output. The anatomical boundary localization subnetwork captures vascular and bronchial boundary features and outputs an anatomical boundary probability map. The output of the multi-task deep network is processed by voxel-level fusion and three-dimensional reconstruction to generate a three-dimensional model including the lesion location, blood vessel direction and bronchial branching structure.
7. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 6, characterized in that: Voxel-level fusion and 3D reconstruction processing includes: Perform weighted fusion of the lesion mask, anatomical boundary probability map and the original fused image data to obtain fused voxel data; Perform morphological closing operations on the fused voxel data to fill holes and breaks in the segmentation results and obtain optimized voxel data. Binarize the optimized voxel data to generate binary voxel labels; The Marching Cubes algorithm is used to traverse the binary voxel labels to generate a three-dimensional model including the lesion location, blood vessels and bronchial structures.
8. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 1, characterized in that: The multi-objective optimization algorithms in the surgical planning module include: Three optimization objectives of the surgical path are defined in the 3D model, including: operation distance cost, whose value is determined by the total length of the surgical path; tissue damage cost, whose value is quantified by the volume of sensitive tissue passed by the path; and operation time cost, whose value is calculated based on the path complexity and the preset operation speed model. A risk assessment index is introduced as a constraint condition. The risk assessment index is calculated by combining the closest distance between the path and the blood vessels and neural structures and the path curvature parameters, and a maximum risk threshold is set. The improved NSGA-II algorithm is used for iterative search to search for non-dominated solution sets in the solution space that meets the maximum risk threshold and generate the Pareto front of candidate surgical paths.
9. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 8, characterized in that: The improved NSGA-II algorithm includes: S1. Encode candidate surgical paths as genetically manipulated individuals and generate an initial population that satisfies basic anatomical constraints; calculate three types of optimization objective values for each candidate surgical path; and simultaneously calculate the risk assessment index for each candidate surgical path. S2. In each iteration: The parent population P t and the offspring population Q t Merge into A t , traverse A t Among all candidate solutions, eliminate all those with risk assessment index R(X)>R max The solution, R max is the maximum risk threshold, and the safe solution set S is obtained t ; To S t Perform hierarchical sorting to generate non-dominated levels Z1, Z2, ..., Z m , where m is the number of levels and Z1 is the optimal frontier. For the solutions at the same level, the original congestion degree D(X) is calculated, and the risk penalty term is introduced to correct the congestion degree: Where D′(X) is the modified congestion degree, and α is the risk penalty coefficient; From S t The next generation of population is selected by level and corrected crowding; S3. Repeat step S2 until the iteration stop condition is met, and output the Pareto optimal solution set from the optimal frontier of the final population as the candidate surgical path plan.
10. The intelligent planning system for thoracic surgery based on multimodal image fusion according to claim 1, characterized in that: The interactive adjustment module includes: The multi-view linkage operation interface simultaneously displays the axial, sagittal, coronal and 3D perspectives of the 3D model, and supports multi-view synchronous zooming, rotation and point selection operations; Real-time collision detection mechanism, based on octree spatial hierarchical algorithm and adaptive radius search strategy, calculates the distance between candidate surgical paths and vascular and bronchial structures in real time, and issues warnings when approaching danger thresholds; The risk dynamic calculation unit calls the risk assessment function in the surgical planning module to perform immediate risk reassessment of the user-adjusted path and present it numerically.
Citation Information
Cited By
Sky-eye insight multi-mode man-machine interaction application system based on pathologist view angle
CN121034667A
Hepatic artery variation dissection rapid reconstruction and intervention oral point recommendation method
CN121962240A
A method and system for generating navigation information for assisting a neurovascular intervention procedure
CN122478629A