Pulmonary nodule real-time positioning method and device based on artificial intelligence multi-modal image fusion technology and applied to sub-pulmonary lobe resection operation
By generating a three-dimensional collapsed lung model through artificial intelligence multimodal image fusion technology, the location of lung nodules can be identified in real time. This solves the problem of insufficient accuracy of traditional intraoperative positioning methods, realizes non-invasive and accurate lung nodule positioning, and reduces the medical risks and operation time for patients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional intraoperative methods for locating lung nodules lack precision, leading to inappropriate surgical resection areas, increasing the medical risks and burden on patients, while invasive procedures also bring complications and radiation exposure issues.
By employing AI-based multimodal image fusion technology, a three-dimensional collapsed lung model is generated by acquiring medical images of the lungs and intraoperative laparoscopic images. This model can identify and accurately locate lung nodules in real time, enabling simultaneous display of two-dimensional images and three-dimensional models, thus avoiding the errors and complications associated with traditional invasive localization.
This method enables non-invasive and precise localization of lung nodules, reduces surgery time, lowers medical risks, avoids complications and radiation exposure caused by puncture, and improves the accuracy and safety of the surgery.
Smart Images

Figure CN121837580A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image processing, in particular to a lung nodule real-time positioning method and device based on artificial intelligence multi-modal image fusion technology applied to sub-lobe resection surgery. BACKGROUND
[0002] With the popularization and promotion of chest CT screening in clinical practice, the detection rate of early lung nodules has been significantly improved. Domestic and foreign literature reports that the detection rate of lung nodules in routine chest CT examination has reached 24%-31%. The results of CALGB140503 clinical trial confirmed that for peripheral early lung cancer with a diameter of ≤2cm, sub-lobe resection and traditional lobe resection had no statistical difference in survival rate, which provided a strong basis for the standardized application of sub-lobe resection in the treatment of early lung cancer. Therefore, as a core treatment method for early lung cancer and some benign lung nodules, sub-lobe resection has a unique advantage of maximizing the preservation of lung function while precisely resecting lesions, and its application in clinical practice is becoming more and more widespread.
[0003] However, sub-lobe resection still faces many key challenges in clinical implementation, of which the precise positioning of lung nodules during surgery is a core link that determines the success or failure of the operation. The traditional intraoperative positioning method generally has the problem of insufficient accuracy, with positioning errors of several millimeters or even centimeters, which may lead to excessive or insufficient resection range - excessive resection will cause additional loss of normal lung function, and incomplete resection will significantly increase the risk of local recurrence, seriously affecting the effectiveness of surgical treatment. The preoperative CT-guided percutaneous puncture positioning technique commonly used in clinical practice can mark the position of the nodule before surgery, but it has many drawbacks that cannot be ignored: this operation is an invasive intervention, with a complication rate of pneumothorax, hemorrhage, etc. as high as 10%-30%, and 15%-20% of cases will have positioning hook shedding, displacement, etc., resulting in positioning failure; patients need to endure obvious pain and discomfort during the puncture process, and multiple CT scans are needed to ensure puncture accuracy, exposing patients to a higher dose of ionizing radiation, further increasing the medical risk and patient burden. SUMMARY
[0004] The present application aims to provide a lung nodule real-time positioning method and device based on artificial intelligence multi-modal image fusion technology applied to sub-lobe resection surgery, which can assist doctors in quickly and accurately locking the position of lung nodules during lung nodule resection surgery under the premise of non-invasiveness, effectively improving the accuracy of surgical operation and shortening the operation time.
[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions: A method for real-time localization of pulmonary nodules in sublobar resection surgery based on artificial intelligence multimodal image fusion technology includes: acquiring core data information and basic model data, wherein the core data information includes pulmonary medical images and intraoperative laparoscopic dynamic images; the basic model data is a three-dimensional lung model, which includes a fused three-dimensional model of pulmonary nodules, bronchi, arteries, and veins; preprocessing the core data and basic model to generate initial data for a three-dimensional collapsed lung model; preparing data and training the three-dimensional collapsed lung model, extracting geometric features, spatial relationship features, and topological features; preprocessing the intraoperative laparoscopic dynamic images, and identifying nodules including the apex, hilum, and fissure of the lung using a trained two-dimensional lung keypoint model and a lung detector. Key features of lung structure under laparoscopy are used to determine the actual degree of lung collapse. Based on feature matching, the 3D collapsed lung model is fine-tuned to generate a target 3D collapsed lung model that adapts to the actual situation during surgery. Through artificial intelligence multimodal image fusion technology, the laparoscopic image and the target 3D collapsed lung model are merged to generate model vertex-level UV coordinates, accurately mapping the laparoscopic image onto the surface of the 3D collapsed lung model, achieving synchronous display of 2D image and 3D model. The coordinates are updated in milliseconds at a frame rate of over 30fps. Through multimodal image fusion display, transparency settings, and free adjustment of coordinate orientation, the 3D lung structure under laparoscopic image is clearly presented from various angles. The location of lung nodules is highlighted in red to determine the specific location of lung nodules in the intraoperative scene in real time.
[0006] The application relates to a lung nodule real-time positioning device based on an artificial intelligence multi-modal image fusion technology applied to sub-lobe resection surgery, which comprises the following modules: a data acquisition module for acquiring core data information and basic model data, wherein the core data information comprises lung medical images and intraoperative endoscope dynamic images; the basic model data is a lung three-dimensional model which comprises a three-dimensional model integrating a lung nodule, a bronchus, an artery and a vein; a preprocessing module for preprocessing the core data and the basic model to generate initial data of a three-dimensional collapsed lung model; a feature extraction module for data preparation and model training of the three-dimensional collapsed lung model to extract geometric features, spatial relationship features and topological features; a model fine-tuning module for preprocessing the intraoperative endoscope dynamic images, identifying key point features of endoscopic lung structures including lung apices, hilums and fissures through a trained lung two-dimensional key point model and a lung detector, and judging the actual collapse degree of lung tissue; fine-tuning the three-dimensional collapsed lung model based on feature matching to generate a target three-dimensional collapsed lung model adapted to the actual intraoperative situation; a multi-modal fusion module for fusing the endoscope image and the target three-dimensional collapsed lung model through the artificial intelligence multi-modal image fusion technology, generating model vertex level UV coordinates, accurately mapping the endoscope image to the surface of the three-dimensional collapsed lung model, and realizing the synchronous display of the two-dimensional image and the three-dimensional model; and a positioning display module for updating the coordinates at a frame rate of more than 30 fps per millisecond, freely adjusting the coordinate direction through multi-modal image fusion display, transparency setting and coordinate direction, clearly presenting the three-dimensional lung structure under the endoscope image from various angles, marking the lung nodule position with red highlights, and determining the specific position of the lung nodule in the intraoperative scene in real time.
[0007] After the technical scheme is adopted, the application has the beneficial effects that: the application provides a lung nodule real-time positioning method and device based on an artificial intelligence multi-modal image fusion technology applied to sub-lobe resection surgery, which acquires lung medical images and a lung three-dimensional model, generates a three-dimensional collapsed lung model according to surgical parameter setting, automatically identifies the features of the three-dimensional collapsed lung model, acquires intraoperative endoscope dynamic images, automatically and real-timely identifies the features of the endoscope image, fine-tunes the three-dimensional collapsed lung model, accurately fuses the endoscope image and the three-dimensional collapsed lung model through the artificial intelligence multi-modal image fusion technology, and real-timely displays the three-dimensional lung structure under the endoscope image and the accurate position of the lung nodule, thereby avoiding the errors of traditional positioning and the improper resection range, without the need for preoperative invasive puncture, avoiding the complications, pain and positioning failure risks caused by puncture, reducing radiation exposure, lowering the medical risks and patient burden, assisting doctors in quickly positioning, shortening the operation time, and better adapting to the precise treatment requirements of sub-lobe resection. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1A flow chart of a lung nodule real-time positioning method based on artificial intelligence multi-modal image fusion technology applied to sub-lobe resection surgery is provided for an embodiment of the present application. Figure 2 A control panel interface schematic diagram of endoscopic video and target three-dimensional collapsed lung model multi-modal image fusion is provided for an embodiment of the present application. Figure 3 A module schematic diagram of a lung nodule real-time positioning device based on artificial intelligence multi-modal image fusion technology applied to sub-lobe resection surgery is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0009] The preferred embodiments of the present application are described below with reference to the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present disclosure.
[0010] Figure 1 An implementation flow schematic diagram of a lung nodule real-time positioning method based on artificial intelligence multi-modal image fusion technology applied to sub-lobe resection surgery is shown, which includes: S101, acquiring core data information and basic model data.
[0011] In an embodiment, the core data information includes lung medical images and intraoperative endoscopic dynamic images; the basic model data is a lung three-dimensional model, which includes a three-dimensional model integrated by lung nodules, bronchi, arteries and veins. The lung medical images are medical image files in DICOM format, such as lung CT images and MR images, which are used to extract patient basic information (such as age and gender) and medical image related feature data (such as lung volume and bronchial bifurcation angle). The intraoperative endoscopic dynamic images are collected by a 1920x1280 high-definition thoracoscope, with a frame rate of ≥30fps, and the original images are in BGR format to restore the texture details of the lung nodules; real-time preprocessing after collection, including Gaussian filter denoising, adaptive gray scale enhancement to improve contrast, BGR to RGB color conversion, while marking invalid frames with excessive blur and occlusion and triggering resampling to ensure data continuity.
[0012] The basic model is a lung three-dimensional model, which is fused by various substructures of lung tissue, and specifically includes: Lung lobes and segments: right upper lobe (RUL, including apical segment S1, posterior segment S2, anterior segment S3), right middle lobe (RML, including lateral segment S4, medial segment S5), right lower lobe (RLL, including superior segment S6, medial basal segment S7, anterior basal segment S8, lateral basal segment S9, posterior basal segment S10), left upper lobe (LUL, including posterior apical segment S1+2, anterior segment S3, superior lingular segment S4, inferior lingular segment S5), left lower lobe (LLL, including superior segment S6, medial basal segment S7, anterior basal segment S8, lateral basal segment S9, posterior basal segment S10); Vessels and airways: right lung bronchus (RLB), right upper lobe bronchus (RULB), right middle lobe bronchus (RMLB), right lower lobe bronchus (RLLB), left lung bronchus (LLB), left upper lobe bronchus (LULB), left lower lobe bronchus (LLLB); right lung vein (RLV), left lung vein (LLV) and their respective lobar branches; right lung artery (RLA), left lung artery (LLA) and their respective lobar branches; Target lesion: pulmonary nodule (PN).
[0013] S102, preprocessing core data and basic model to generate initial data of three-dimensional atelectasis lung model.
[0014] By collecting the relevant parameters of lung medical images and three-dimensional lung models, including patient basic characteristic parameters, biomechanical parameters and surgical condition parameters, the atelectasis factor is composed. According to the three-dimensional lung model, the model center is set by point cloud technology, the distance of each vertex from the center is calculated, and the three-dimensional lung model is shrunk according to the atelectasis factor. The patient's preoperative high-resolution CT (HRCT) image is taken as the data source, and the lung three-dimensional basic model is constructed by medical image reconstruction software (such as 3dslicer). First, the HRCT image is denoised, gray corrected and tomographic image aligned, the complete anatomical structure of both lungs is retained, the spatial running of bronchus, artery and vein and the boundary of lung segment are determined; then the model is converted into a triangular mesh model, the grid unit size is set to 0.8mm, redundant vertices and invalid facets are deleted, and the surface roughness is ensured to be ≤0.1mm, and the coordinate system is adopted as the world coordinate system (X axis: left-right direction, Y axis: front-back direction, Z axis: up-down direction).
[0015] Atelectasis factor construction: Collect patient basic characteristic parameters (such as age, height, weight), biomechanical parameters and surgical condition parameters, determine the weight of each parameter by analytic hierarchy process, construct atelectasis factor F (value range 0.6-1.0, the smaller the value, the stronger the atelectasis degree), calculate the specific value according to the formula F=Σ (parameter standardized value × weight) to define the atelectasis degree of the model.
[0016] Shrinkage based on point cloud technology: Extract all vertex coordinates of the three-dimensional model of the lung by point cloud technology, calculate the geometric center (the midpoint of the line connecting the right and left hilum) of the model; traverse the vertices, calculate the distance of each vertex to the geometric center according to the Euclidean distance formula; adopt a nonlinear shrinkage algorithm, combine the atrophy factor and the vertex distance to adjust the coordinates (adjustment formula: P'(X i ’,Y i ’,Z i ’)=O+(P-O)×[F+(1-F)×(D_core / D i )]),achieve the gradient shrinkage of the lung edge area greater than the core area of the hilum.
[0017] Model accuracy verification: After shrinkage, verify the matching degree of the model and the actual atrophy state during the operation through three indicators: lung lobe volume error ≤5%, lung segment boundary position deviation ≤1.2mm, bronchus and artery and vein spatial running deviation ≤0.9mm, to ensure that the accuracy requirements of operation planning and navigation are met.
[0018] Select 1 case of lung surgery patient (male, 58 years old, height 175 cm, weight 72 kg, due to right lower lobe nodule, planned for thoracoscopic lung segment resection), based on the preoperative chest high resolution CT (HRCT) image of the patient, complete the construction of the three-dimensional basic model of the lung through medical image reconstruction software (such as 3dslicer), the specific preprocessing steps are as follows: Denoise, gray correction and tomographic image alignment are performed on the HRCT image, and the complete anatomical structure of the right lung (including upper lobe RUL, middle lobe RML and lower lobe RLL), left lung (including upper lobe LUL and lower lobe LLL), bronchus (right lung bronchus RLB, left lung bronchus LLB, etc.), artery (right lung artery RLA, left lung artery LLA, etc.), vein (right lung vein RLV, left lung vein LLV, etc.) and lung segment boundary (such as right lower lobe posterior basal segment S10, left upper lobe apicoposterior segment S1+2, etc.) are determined.
[0019] Convert the reconstructed three-dimensional model to a triangular mesh model, set the grid unit size to 0.8mm, delete redundant vertices and invalid facets, and ensure the smoothness of the model surface (surface roughness ≤0.1mm), finally get a lung basic model containing 128,640 vertices and 257,280 triangular facets, the model coordinate system uses world coordinate system (X axis: left-right direction, Y axis: front-back direction, Z axis: up-down direction).
[0020] The basic characteristic parameters, biomechanical parameters and surgical condition parameters of the patient are collected, the weight of each parameter is determined by the analytic hierarchy process, a collapse factor F (value range 0.6-1.0, the smaller the value, the stronger the collapse degree) is constructed, and the collapse factor is calculated according to the formula: F = Σ (parameter standardized value x weight). The calculated collapse factor F of the patient is 0.75, indicating that the lung basic model needs to be moderately collapsed.
[0021] The shrinkage deformation processing based on point cloud technology is as follows: 1. Model center setting: all vertex coordinates of the lung three-dimensional model are extracted by point cloud technology, the geometric center O (X0=12.3mm, Y0=-4.5mm, Z0=8.7mm) of the model is calculated, the center is located at the midpoint of the right hilum and the left hilum, which meets the anatomical center position characteristics of the lung.
[0022] 2. Vertex distance calculation: all vertices of the model are traversed, and the distance D i of each vertex P (X i ,Y i ,Z i ) to the center O is calculated according to the Euclidean distance formula. , wherein, , , are vertex coordinates, , , are geometric center coordinates, and the distance range is calculated to be 3.2mm-68.5mm, wherein the average distance D_avg of the lung edge vertex is 42.6mm, and the average distance D_core of the hilum region vertex is 15.8mm.
[0023] 3. Shrinkage deformation implementation: according to the collapse factor F and the vertex distance D i , a nonlinear shrinkage algorithm is used to adjust the coordinates of each vertex, and the adjustment formula is: P' (X i ',Y i ',Z i ') = O + (P-O) x [F + (1-F) x (D_core / D i )], wherein (P-O) is the vector difference of the vertex relative to the center, i.e. (X , , ), (D_core / D i ) is used to realize gradient shrinkage (the shrinkage amplitude of the lung edge region is greater than that of the hilum core region), and P' (X i ',Y i ',Z iF is the shrinkage of the i-th vertex, F=0 means full gradient shrinkage, F=1 means no shrinkage, all vertex coordinates remain in the initial state, 0<F<1 means toxic atrophy, the shrinkage amplitude is "basic shrinkage F" and "gradient shrinkage (1-F x (D_core / D i )”. After calculation, the average shrinkage of the lung edge vertex is 8.3mm, because the edge vertex D i is larger, the D_core / D i ratio is smaller, so the shrinkage amplitude is higher; the average shrinkage of the lung edge vertex is 2.1mm, because the edge vertex D i is smaller, the D_core / D i ratio is closer to 1, so the shrinkage amplitude is lower; the number of vertices of the shrunk model remains 128,640, the triangular facet structure is unchanged, the model volume is reduced from the initial 5.2L to 3.9L, which meets the anatomical volume characteristics of moderate atrophy lung.
[0024] The matching degree of the three-dimensional atrophy lung model after shrinkage deformation and the actual atrophy state in the operation is verified by the following indexes: 1. The lung lobe volume error is ≤5% (the actual volume of the right lower lobe is 0.82L, and the model volume is 0.78L); 2. The lung segment boundary position deviation is ≤1.2mm; 3. The spatial running deviation of bronchus, artery and vein is ≤0.9mm, which meets the accuracy requirements of operation planning and navigation.
[0025] S103, data preparation and model training are performed on the three-dimensional atrophy lung model, and geometric features, spatial relationship features and topological features are extracted.
[0026] The three-dimensional lung model data preparation, feature extraction model training, realize the recognition and positioning of three-dimensional atrophy lung; construct the feature library of three-dimensional lung, including geometric features, spatial features, topological features and multi-modal fusion features. Input the grid data of three-dimensional file, complete noise removal and coordinate system preprocessing; then call the multi-modal feature extractor and the geometric feature extractor to generate the geometric, spatial, topological and fusion features of each grid.
[0027] The multi-modal fusion network divides the structure into lung lobes, lung segments, ducts, and incision sphere according to the rules of volume, compactness, and elongation ratio; then the fusion features are input into the hierarchical classifier, and the hierarchical classifier outputs the fine classification labels of lung lobes, lung segments, bronchus, and arteries and veins; then the structure adjacency graph is constructed, and the classification results are corrected by the Topology Aware Corrector combined with the topological relationship; and then the JSON annotation file, visual model and topological relationship report are generated.
[0028] With the Total Segmentator dataset version 2 as the training data source, the geometric, spatial, and topological features of the three-dimensional lung model are extracted. The data integrity is verified and invalid samples are filtered through BatchPreparer.validate_batch_structure. The grid data of the three-dimensional file is inputted, and the noise removal and coordinate system preprocessing are completed. The feature vectors are standardized using Z-score, and the point cloud data is normalized.
[0029] A three-dimensional lung feature library is constructed, covering geometric features (28 dimensions, describing geometric morphological quantitative parameters), spatial features (15 dimensions, representing spatial position and distribution vectors), topological features (12 dimensions, reflecting structure connection relationship parameters), and multi-modal fusion features. The multi-modal feature extractor and the geometric feature extractor are called to generate various features for each grid.
[0030] The model input contains geometric / spatial / topological features and point cloud data. Parallel branch processing is used: geometric / spatial / topological features are mapped through fully connected layers, and point cloud data is extracted through PointNet or similar networks. The fusion layer uses concatenation or attention mechanism to integrate multi-source features. The output layer sets the corresponding dimension fully connected layer according to the classification / segmentation task, and the hidden layer activation function uses ReLU or LeakyReLU. The classification task output layer uses Softmax, and the regression task uses linear activation.
[0031] The training configuration is as follows: cross-entropy loss is used for classification tasks, weighted loss or FocalLoss is used for multi-label or unbalanced data, and Chamfer distance loss is added for point cloud related tasks. The optimizer uses Adam or SGD, the initial learning rate is 1e-4~1e-3, the weight decay is 1e-5, and the learning rate scheduling uses cosine annealing or step decay strategy.
[0032] The structure classification and result output are as follows: the multi-modal fusion network divides the structure into lung lobes, lung segments, ducts, and incision sphere based on volume, compactness, and elongation ratio; the fusion feature input hierarchical classifier outputs lung lobes, lung segments, bronchi, and arteries; the structure adjacency graph is constructed, and the Topology Aware Corrector is used to correct the classification results based on topological relationships, finally generating JSON annotation files, visual models, and topological relationship reports. The key hyperparameters are as follows: The optimization strategy is as follows: grid search (GridSearch) is used to select the optimal hyperparameters with 5-fold cross-validation, and checkpoint_manager.py is used to save the best model weights to avoid overfitting The data set is divided as follows: training set 70%, validation set 15%, and test set 15% (corresponding to the validation_split parameter in create_data_loaders). The classification task uses accuracy, recall, and F1 score; the point cloud task uses mean distance error.
[0033] In S104, the intraoperative endoscopic dynamic image is preprocessed, and the lung two-dimensional key point model and the lung detector are trained to identify the key point features of the endoscopic lung structure including the lung apex, hilum, and lung fissure, and to determine the actual collapse degree of the lung tissue.
[0034] The lung endoscopic video file is obtained, and it is split into five core segments of lung apex segment, lung anterior margin segment, lung posterior margin / lung root segment, lung lower margin / lung ligament segment, and lung fissure segment according to the anatomical region and surgical procedure, ensuring that each segment focuses on a single anatomical region and avoids feature cross confusion; then the key points are labeled, and the key point model is trained, which automatically identifies the lung structure features and completes the preliminary labeling of spatial coordinates and text attributes, and adjusts the positions of key points such as lung apex and lung root.
[0035] The present application focuses on the processing of lung endoscopic video and the training of lung structure key point model, and is divided into two core stages: lung endoscopic video segment splitting stage and key point labeling and model training stage. Through the combination of GPU acceleration support (based on the CUDA capabilities of PyTorch and OpenCV), efficient video processing and model training are realized to ensure accurate focus on lung structure feature analysis. The lung endoscopic video file acquisition and preprocessing process is as follows: collect clinical lung endoscopic surgery video files, formats include but are not limited to common video formats such as MP4, AVI, etc., ensure that the video is clear and has no obvious occlusion, and can fully present the surgical field of each anatomical region of the lung.
[0036] The check_gpu_support() function is called before processing to detect hardware acceleration capability, and the specific process is as follows: check PyTorch GPU support: judge whether CUDA is available through torch.cuda.is_available(), if available, record the number of GPUs, the name of each GPU (such as torch.cuda.get_device_name(i)), the memory (converted to GB units) and the CUDA version, and store them in gpu_info['pytorch_available'] and gpu_details.
[0037] Check OpenCVCUDA support: Check OpenCV's support for CUDA using cv2.cuda.getCudaEnabledDeviceCount(). The result is recorded in gpu_info['opencv_available']. If GPU support is detected, subsequent video processing and model training will preferentially use GPU acceleration to improve efficiency; if detection fails (such as no CUDA environment or library version incompatibility), use CPU processing, and output the corresponding error prompt (such as PyTorch GPU detection failure: [error information]).
[0038] Lung endoscopic video segment splitting: Based on anatomical regions and surgical procedures, the video is split into five core segments to ensure that each segment focuses on a single anatomical region. The splitting criteria are as follows: Anatomical features: Based on the anatomical structure of the lung, it is divided into five segments, including the apical segment (the upper tip area of the lung), the anterior segment (the anterior boundary area of the lung), the posterior segment / root segment (the posterior boundary and hilar root area of the lung), the inferior segment / ligament segment (the inferior boundary and ligament area connecting the pleura), and the fissure segment (the interlobar fissure area). Surgical procedure logic: Combine the operation sequence to extract continuous video segments of the corresponding anatomical region to avoid mixed frames of cross-region operations.
[0039] Video frame analysis technology (OpenCV CUDA acceleration interface can be used to improve frame processing speed, such as cv2.cuda related functions) is used to extract frames or key frames from the video. By image recognition, the features of each anatomical region are preliminarily located (such as the shape of the lung apex and the texture difference of the lung fissure), and the target segment is segmented based on the time axis information to ensure that there is no cross-mixing of other region features within the segment. Finally, five types of segment files are output, labeled as "apical segment", "anterior segment", "posterior segment_root segment", "inferior segment_ligament segment", and "fissure segment".
[0040] For the five types of segments after splitting, artificial or semi-automatic labeling of lung structure key points is performed, including but not limited to: lung apex, lung root center, lung anterior boundary point, lung posterior boundary point, lung inferior edge endpoint, lung fissure starting point and ending point, etc. Labeling information includes spatial coordinates (based on two-dimensional pixel coordinates of video frames or three-dimensional space reconstruction coordinates) and text attributes (such as "lung apex", "lung root artery attachment point", etc.).
[0041] The labeled segment frames and key point information are organized into a training set, divided into a training set, a validation set, and a test set, and the GPU acceleration capability of PyTorch (such as torch.cuda for data loading and model training) is used to improve processing efficiency. Select a deep learning model suitable for key point detection (such as a key point detection network based on CNN), input the segment frame image, and output the predicted spatial coordinates and text attributes of the key points. Train the model on a GPU that supports CUDA, optimize the parameters through back propagation, minimize the error between the predicted coordinates and the labeled coordinates, and optimize the text attribute classification accuracy.
[0042] The trained model is used to automatically recognize new thoracoscopy video segments, outputting the spatial coordinates and text attributes of key points in each anatomical region. For core key points such as the lung apex and lung root, combine anatomical prior knowledge (such as the relative position relationship between the lung apex and lung root) and the context information of the video frames to fine-tune the predicted coordinates, ensure the positioning accuracy, and reduce the deviation caused by image noise or occlusion.
[0043] Through GPU acceleration technology (PyTorch and OpenCV CUDA support), the efficiency of video processing and model training is improved throughout the process, ensuring that each link focuses on a single anatomical region, avoiding feature confusion, and providing accurate data support for the analysis and auxiliary diagnosis of thoracoscopy surgery.
[0044] S105, based on feature matching, fine-tune the three-dimensional collapsed lung model to generate a target three-dimensional collapsed lung model that adapts to the actual situation during surgery.
[0045] Import the initial parameters of the predicted three-dimensional collapsed lung model, including standard anatomical structure sizes (such as standard lung lobe volume, bronchial bifurcation angle reference value), surgical preset parameters (such as surgical approach angle, expected collapse degree threshold), and other basic configurations. Use clinical sub-lobe resection surgery case data, covering preoperative HRCT images of 100 patients, intraoperative endoscopic dynamic videos, and postoperative pathological verification data, of which 80 cases are training sets, 10 cases are validation sets, and 10 cases are test sets; the data have been ethically reviewed, and the patient's basic information (age, gender, lung function indicators) and surgical parameters (surgical approach, collapse induction conditions) are complete.
[0046] Denoise, grayscale correction, and tomographic alignment are performed on the HRCT images, a three-dimensional lung base model is reconstructed and converted to a triangular mesh format (mesh cell size 0.8mm), and the endoscopic video is divided into five types of segments according to the anatomical region, and the spatial coordinates and text attributes of 11 types of structure key points such as the lung apex and lung root are labeled, and the feature vector is standardized by Z-score, and the point cloud data is normalized.
[0047] The three-section neural network of "feature extraction-feature fusion-parameter prediction" is adopted, the input layer receives three-dimensional model grid data (vertex coordinates, topological relationship) and endoscope image feature vector (28-dimensional geometric features + 15-dimensional spatial features + 12-dimensional topological features); the hidden layer is provided with 3 full connection layers (the dimensions are 1024, 512 and 256 in sequence), and the attention mechanism (Attention) is introduced to strengthen the key structure feature weight; the output layer is a 3-class parameter prediction branch, which corresponds to the anatomical structure parameter, the dynamic response parameter and the structure correlation parameter respectively.
[0048] The Dropout layer (dropout rate 0.3) is adopted between the full connection layers to prevent overfitting, the LeakyReLU (negative slope 0.01) is selected as the activation function of the hidden layer, and the linear activation function is adopted for the output layer, so as to ensure the continuity and rationality of the parameter prediction. The hybrid loss function is adopted, the L1 loss (weight 0.4) is used to constrain the structure coordinate prediction deviation, the cosine similarity loss (weight 0.3) is used to match the dynamic response trend, and the Dice loss (weight 0.3) is used to optimize the integrity of the structure correlation, and the total loss function is: Loss = 0.4 x L1Loss + 0.3 x CosineLoss + 0.3 x DiceLoss. The Adam optimizer is selected, the initial learning rate is 1e-4, the weight decay is 1e-5, and the cosine annealing learning rate scheduling strategy (period 50 rounds, minimum learning rate 1e-6) is adopted; the batch size is set to 16, the training rounds are 200 rounds, and the early stopping strategy (stop training if the loss does not decrease for 20 consecutive rounds) is adopted based on the validation set loss.
[0049] The hyperparameter adjustment range is as follows: attention mechanism weight coefficient (0.1-0.5), Dropout rate (0.2-0.4), hidden layer dimension (512-2048), the optimal combination is determined through grid search combined with 5-fold cross-validation. The anatomical structure parameter prediction accuracy is preferentially optimized, and then the dynamic response parameter and the structure correlation parameter are iteratively adjusted; the structure positioning deviation and the morphological fit degree of the test set are calculated after each training, and the hyperparameters are updated through the Bayesian optimization algorithm to ensure that the model converges to the global optimum.
[0050] Based on the segmented feature space coordinate deviation of the network output (such as the lung tip height deviation and the lung fissure deviation), the three-dimensional coordinates and the size of the corresponding lung structure in the model are corrected, so that the average deviation of the key structure positioning of the model and the intraoperative annotation result is ≤2mm. According to the lung lower edge collapse movement range and the morphological distortion data, the nonlinear shrinkage algorithm coefficient of the model is optimized, so that the matching degree of the dynamic change of the model lung tissue and the actual collapse state in the operation is ≥95%. Combined with the visible structure types of the pulmonary hilum and the pulmonary root, the morphological changes of the pulmonary ligament, the heart notch and the left lung small tongue, the connection relationship matrix and the space arrangement weight of the lung tissue in the model are optimized, so as to ensure that the model structure integrity is consistent with the intraoperative annotation feature attributes.
[0051] The adjusted model output result is compared with the feature data labeled by the intraoperative endoscopic image, and three types of indexes are calculated, including structure coordinate deviation (Euclidean distance), morphological fitting degree (Dice coefficient), and dynamic response error (inter-frame displacement difference). If the structure positioning deviation is greater than 3 mm or the morphological fitting degree is less than 95%, return to the parameter adjustment step to re-optimize; if the indexes meet the requirements after three consecutive iterations, and the average structure positioning deviation of the test set is less than or equal to 2.5 mm, the morphological fitting degree is greater than or equal to 96%, and the dynamic response error is less than or equal to 1 mm / frame, stop iteration, and finally obtain the intraoperative three-dimensional collapsed lung model.
[0052] S106, by means of artificial intelligence multi-modal image fusion technology, fusing endoscopic images and target three-dimensional collapsed lung model, generating model vertex level UV coordinates, accurately mapping endoscopic images to the surface of three-dimensional collapsed lung model, realizing synchronous display of two-dimensional images and three-dimensional model.
[0053] Load the three-dimensional collapsed lung model and endoscopic video stream, establish the mapping relationship between the model coordinate system and the video coordinate system: through VTK, computer vision technology, simultaneously use two-dimensional feature library, three-dimensional feature library and feature relationship library, make the endoscopic dynamic image real-time and three-dimensional collapsed lung fusion registration, output clear and smooth multi-modal fusion body, that is, clear three-dimensional collapsed lung model under endoscopic dynamic image, including lung nodule, bronchus, artery and vein and clear lung segment boundary. Load the STL format model file of each subdivided structure of the lung by the load_lung_model method, including lung lobe, lung segment, bronchus, artery and vein, lung nodule, etc. Assign different colors and transparencies to each structure during loading, which is convenient for subsequent visualization and differentiation.
[0054] Combine all loaded sub-models into a complete lung model (self.lung_STL), convert the combined STL model to PLY format by vedo.write function and save it to the specified path, in order to optimize the efficiency of subsequent texture mapping and feature matching; load the PLY format model and clone the original mesh data (self.original_mesh) and vertex coordinates (self.original_points) as the reference data in the fusion process. Set the initial color of the model to light blue (lightblue) and the transparency to 0.5, add the model by the Plotter tool of VEDO and start the visualization rendering, enable the coordinate axis display (axes=1) and interactive display mode, and ensure that the model loading effect can be checked in real time.
[0055] Access to the intraoperative endoscopic video stream through a dedicated data interface, establish a circular video frame buffer (buffer capacity is set to 30 frames, matching the video capture frame rate), realize the real-time acquisition, frame extraction and cache management of video stream, ensure the continuous output of each frame of image, provide stable two-dimensional image data source for the subsequent dynamic fusion with three-dimensional model.
[0056] Through VTK technology and computer vision algorithm, the mapping relationship between the three-dimensional collapsed lung model coordinate system (world coordinate system: X axis left and right, Y axis front and back, Z axis up and down) and the endoscopic video coordinate system (pixel coordinate system: the upper left corner as the origin) is established; the two-dimensional feature library (endoscopic image key point features), three-dimensional feature library (model geometry / space / topological features) and feature relationship library are called to provide feature matching basis for real-time fusion and registration of endoscopic dynamic images and three-dimensional model, ensuring that the fusion can clearly show lung nodules, bronchi, arteries and veins, and lung segment boundaries.
[0057] The technology of superimposing two-dimensional images on three-dimensional images uses complex curved surface texture algorithm to generate vertex-level UV coordinates of the model: if the three-dimensional collapsed lung model has no obvious holes, select the least squares conformal mapping (LSCM) algorithm, set the parameters, call the function to generate vertex-level UV coordinates, if the model contains airway holes, select the angle domain unfolding (ABF) algorithm, set the parameters, call the function. Extract two-dimensional feature points such as bronchial bifurcation points, blood vessel intersection points, and lung nodule edges from the endoscopic images, and extract corresponding three-dimensional feature points from the three-dimensional collapsed lung model; call the two-dimensional feature library, three-dimensional feature library and feature relationship library for accurate matching through the spatial position and topological relationship of the feature points; with the help of VTK technology and computer vision algorithm, calculate the conversion matrix from the video coordinate system to the three-dimensional model coordinate system, and update the matrix in real time according to the dynamic changes of the intraoperative endoscopic images, to ensure the accuracy of dynamic registration.
[0058] Vertex-level UV coordinate generation: curved surface texture algorithm is used to generate vertex-level UV coordinates of the model to realize accurate mapping of two-dimensional endoscopic images and three-dimensional model: if the three-dimensional collapsed lung model has no obvious holes, select the least squares conformal mapping (LSCM) algorithm, fix any two points on the model surface as the mapping boundary, set the convergence threshold to 1e-6 and the maximum iteration number to 1000, and call the algorithm function to generate UV coordinates to ensure the angle-preserving and continuity of the mapping.
[0059] For models without obvious holes, the least squares conformal mapping (LSCM) algorithm is used, the model surface is fixed with any two points as the mapping boundary, the convergence threshold is set to 1e-6 and the maximum iteration number is set to 1000, the algorithm function is called to generate UV coordinates, which ensures the angle-preserving mapping effect from the model surface to the two-dimensional plane, avoiding texture distortion.
[0060] For the airway hole model, the angle domain expansion (ABF) algorithm is used, with angle constraint weight 0.5, area constraint weight 0.5, and hole boundary smoothing technology. The algorithm function is called to generate UV coordinates, adapt to the texture mapping needs of complex hole structure, and ensure mapping continuity.
[0061] The endoscopic image in the mapping area is positionally offset, scaled, and rotationally transformed by the texture transformation layer: set the texture scaling factor, calculate the scaled size, and set the position offset and rotation angle to complete fine adjustment of the texture. Based on the video frame size and the scaling factor, the mapping area size is calculated, and the center position offset is calculated to ensure that the mapping area is displayed in the center. The mapping area size = video frame size x 0.6; the center offset = (video frame size - mapping area size) / 2.
[0062] The boundary protection mechanism is turned on to handle the texture synthesis boundary, and the processed endoscopic image is mapped to the surface of the three-dimensional collapsed lung model to realize real-time matching: the final placement position of the texture is calculated, and the texture synthesis is completed through position boundary protection and size boundary protection. The endoscopic image in the mapping area is geometrically transformed by the texture transformation layer, the texture scaling factor is set, and the scaled size is calculated: scaled size = mapping area size x scaling factor; the position offset is accurately adjusted by the UI slider, as follows: self.x_value.setText(f"{self.x_slider.value() / 100.0:.1f"); self.y_value.setText(f"{self.y_slider.value() / 100.0:.1f")。
[0063] The rotation angle is set to adjust the direction of the texture, as follows: self.rotation_value.setText(f"{self.rotation_slider.value()°")。
[0064] The transparency parameter of the texture is adjusted as needed, as follows: self.alpha_value.setText(f"{self.alpha_slider.value() / 100.0:.1f")。
[0065] The boundary protection mechanism is as follows: position boundary protection: ensure that the texture after transformation does not exceed the range of the model surface; size boundary protection: prevent deformation caused by excessive stretching or compression of the texture. Texture synthesis is as follows: calculate the final placement position of the texture to ensure matching with the three-dimensional model surface. Apply a smooth transition algorithm to process the texture edge to avoid obvious splicing marks. The processed endoscopic image is mapped to the surface of the three-dimensional collapsed lung model through UV coordinates.
[0066] Real-time tracking of endoscopic image changes, updating feature point matching relationships, dynamically adjusting coordinate conversion matrices, real-time registration of endoscopic dynamic images and three-dimensional models, multi-modal fusion body output, fusion results including endoscopic dynamic images, three-dimensional collapsed lung models, lung nodules, bronchi, arteries and veins, and pulmonary segment boundaries. Real-time rendering and display of the fusion results are achieved through the _update_vedo_display method, as follows: #update display example as follows: img_array=self.vp.screenshot(asarray=True); q_img=QImage(img_array.data,width,height,bytes_per_line,QImage.Format_RGB888); pixmap=QPixmap.fromImage(q_img); self.vedo_label.setPixmap(pixmap.scaled(200,200,Qt.KeepAspectRatio,Qt.SmoothTransformation))。
[0067] Interactive control provides three-dimensional model rotation, scaling, translation, and other interactive operations, supports real-time adjustment of fusion parameters, and optimizes fusion effects. Through the above process, the system realizes precise fusion of endoscopic images and three-dimensional collapsed lung models, providing intuitive and precise visualization support for surgical planning and navigation.
[0068] S107, update coordinates at a frame rate of 30fps or above with millisecond-level updates, display through multi-modal image fusion, transparency setting, and coordinate orientation free adjustment, clearly present the three-dimensional lung structure under the endoscopic image from various angles, highlight the lung nodule position in red, and determine the specific position of the lung nodule in the intraoperative scene in real time.
[0069] After fusion, the three-dimensional structure of the collapsed lung can be viewed in real time under dynamic endoscopic images, clearly showing lung segments, trachea, arteries and veins, and nodules. A deep learning model is used to detect the location information of lung nodules in the endoscopic frames in real time, outputting the coordinates of the lung nodules and verifying which region of the three-dimensional structure they fall within. The location of the lung nodules is then highlighted in red on the three-dimensional model, and the localization results are updated in real time as the surgical procedure changes.
[0070] like Figure 2 As shown, the current interface parameters and coordinate system are explained. The 3D model in the figure uses the surgical space coordinate system (matching the laparoscopic field of view): The X-axis represents the left and right directions (the horizontal direction of the model in the figure). The Y-axis represents the front-to-back direction (the depth direction of the model in the figure). The Z-axis is the vertical direction (the vertical direction of the model in the figure). Based on the laparoscopic lens viewpoint (0° is the direction the lens is pointing towards the lung tissue). Taking the nodule in the lower lobe of the right lung in the image as an example, the interface control parameters (“Control Parameters” on the left) are: Opacity: 0.8 (the 3D model is semi-transparent and can be overlaid with endoscopic images); Endoscopic control: 0.7 (the degree of fusion between the endoscopic image and the 3D model); Brightness / Contrast: 0.5 / 0.7 (to ensure clear lung tissue texture); Endoscopic angle (“Endoscopic Camera” on the right): 80° (the lens is facing the lower lobe region of the right lung).
[0071] 3D model center coordinates (geometric center of the model in the figure): O( =15.2mm, =-6.3mm, =9.1mm).
[0072] Using a deep learning model (real-time analysis of laparoscopic frames), a lung nodule was detected in the lateral basal segment (S9) of the right lower lobe. The pixel coordinates of the nodule in the output laparoscopic image were (420, 315) (corresponding to the red area in the right laparoscopic image). After conversion using the image-3D model registration algorithm, the 3D coordinates of the nodule in the surgical space coordinate system are obtained: P( =22.7mm, =-8.5mm, =5.4mm).
[0073] The regional verification information is as follows: Calculate the distance from P to the boundary of each lung segment, confirm that it is located in the outer basal segment (RLL-LB, S9) of the right lower lobe (3.2 mm from the center of the lung segment) and close to the branch of the right lower lobe bronchus (RLLB).
[0074] The angle matching information is as follows: the current angle of the endoscope lens is 80°, and the visual angle of the nodule relative to the lens is: horizontal angle a = 15° (the lens is right-biased by 15°), vertical angle β = -10° (the lens is downward-biased by 10°), which is consistent with the visual position of the nodule in the endoscope picture.
[0075] When the surgical instrument pulls the lung tissue (the lower lobe of the right lung is slightly displaced), the interface parameters are updated: the endoscope angle is adjusted to 85°, and the opacity is adjusted to 0.9; the deep learning model is re-detected, and the three-dimensional coordinates of the nodule are updated to P'(X_P'=23.1mm, Y_P'=-8.1mm, Z_P'=5.6mm).
[0076] As shown in Figure 3 A lung nodule real-time positioning device based on artificial intelligence multi-modal image fusion technology applied to sub-lobe resection surgery, comprising: A data acquisition module 301 is configured to acquire core data information and basic model data, wherein the core data information includes lung medical images and intraoperative endoscopic dynamic images; and the basic model data is a three-dimensional lung model, which includes a three-dimensional model integrating a lung nodule, a bronchus, an artery, and a vein. A preprocessing module 302 is configured to preprocess the core data and the basic model to generate initial data of a three-dimensional collapsed lung model. A feature extraction module 303 is configured to perform data preparation and model training on the three-dimensional collapsed lung model, and extract geometric features, spatial relationship features, and topological features. A model fine-tuning module 304 is configured to preprocess the intraoperative endoscopic dynamic images, identify key point features of the endoscopic lung structure including the lung apex, the lung gate, and the lung fissure through a trained lung two-dimensional key point model and a lung detector, and judge the actual collapse degree of the lung tissue; based on feature matching, the three-dimensional collapsed lung model is fine-tuned to generate a target three-dimensional collapsed lung model that adapts to the actual situation during the operation. A multi-modal fusion module 305 is configured to fuse the endoscopic image and the target three-dimensional collapsed lung model through artificial intelligence multi-modal image fusion technology, generate model vertex-level UV coordinates, accurately map the endoscopic image to the surface of the three-dimensional collapsed lung model, and realize synchronous display of the two-dimensional image and the three-dimensional model. A positioning display module 306 is configured to update the coordinates at a frame rate of 30fps or higher and a millisecond level, display through multi-modal image fusion, set the transparency, and freely adjust the coordinate direction, clearly present the three-dimensional lung structure under the endoscopic image from various angles, mark the lung nodule position with red highlight, and determine the specific position of the lung nodule in the intraoperative scene in real time.
[0077] The embodiments of the present disclosure provide a computing device for executing a lung nodule real-time positioning method based on multi-modal image fusion technology applied to sub-lobe resection surgery, comprising: Memory: for storing computer program instructions corresponding to the full-process execution logic of the foregoing positioning device, specifically including: multi-source data acquisition logic corresponding to the acquisition module, data analysis and model matching logic corresponding to the processing module, and result output logic corresponding to the output module; wherein the data acquisition logic covers execution codes for lung DICOM image (CT / MR sequence) acquisition, 1920x1280 high-definition endoscope dynamic image (30 fps and above frame rate, BGR / RGB format) acquisition, STL / OBJ / PLY format three-dimensional model file reception, and model weight file reception; the data analysis and model matching logic covers execution codes for DICOM data analysis, endoscope image preprocessing (Gaussian filtering, grayscale enhancement, color conversion), feature extraction (11 types of lung structure feature recognition, JSON annotation file generation), three-dimensional collapsed lung model debugging (anatomy / dynamic response / structure correlation parameter adjustment), and multi-modal image matching (LSCM / ABF algorithm UV coordinate generation, texture transformation, boundary protection); and the result output logic covers execution codes for multi-modal fusion image, lung nodule 3D positioning information (deviation ≤3 mm), and lung structure topology relationship output. Processor: in communication connection with the memory, for calling and executing the computer program instructions stored in the memory. When the computer program instructions are executed by the processor, the modules of the foregoing positioning device are triggered to work cooperatively, and the following positioning method is specifically executed: Lung DICOM medical images, endoscope dynamic images, three-dimensional model files, and large model weight files are acquired through multi-source data acquisition logic; Patient basic features, image anatomy parameters, endoscope image features, and three-dimensional collapsed lung model features are extracted through data analysis logic to generate a structured parameter set and an annotation file; Three-dimensional collapsed lung model parameters are debugged through model matching logic, a coordinate system mapping relationship is established, and real-time fusion matching of endoscope images and three-dimensional models is completed; Fusion images, positioning information, and other parameters are output in real time to an interactive display screen through result output logic, and painless and non-invasive real-time positioning of lung nodules in sub-lobe resection surgery is achieved.
[0078] The method and embodiments in the embodiments of the disclosure can be implemented as a computer software program. For example, the embodiments of the disclosure include a computer program product, which includes: a computer readable medium and a computer program. The computer readable medium is used to carry the computer program; and the computer program includes program codes for executing the full-process functions of the foregoing positioning device, specifically including: The data acquisition module code controls the computer device to interface with a medical image storage and transmission system (PACS), acquires a DICOM format sequence file, and completes format analysis and integrity checking; controls a high-definition thoracoscope to acquire a BGR format endoscope image, performs Gaussian filter denoising, adaptive grayscale enhancement, color space conversion (BGR→RGB), and invalid frame marking and resampling; receives a three-dimensional model file and performs integrity checking, noise filtering, and coordinate system normalization processing; receives a pre-trained weight file of the model, and supports version checking and updating; The data processing module code analyzes a DICOM file, extracts image parameters, imports a three-dimensional lung model, completes lung tissue segmentation and grid optimization, splits an endoscope image into five types of core anatomical segments, identifies 11 types of lung structure features, corrects spatial coordinates, supplements text attributes, and generates a JSON annotation file; pre-processes three-dimensional model grid data, calls a multi-modal / geometric feature extractor to generate fused features, completes structure classification through a multi-modal fusion network and a hierarchical classifier, corrects the classification results, and generates a visual model and a topology report; The model matching module code: through three-dimensional technology and computer vision technology, a coordinate system mapping between the three-dimensional model and the endoscope image is established, UV coordinates are generated according to the model hole state by selecting LSCM / ABF algorithm, a central mapping area (1152x768) with a scale factor of 0.6 is defined, texture scaling (default 1.2), offset, rotation (default 15°) adjustment and boundary protection are performed, and texture mapping is completed; the model prediction data and the actual data during the operation are compared, and the anatomical structure / dynamic response / structure correlation parameters are iteratively adjusted until the structure positioning deviation is less than or equal to 3 mm and the morphological fit degree is greater than or equal to 95%; The result output module code controls an interactive display screen to output a multi-modal fusion image, lung nodule 3D positioning information, lung structure information, and deviation checking report in real time. When the computer program is executed by the processor of the computer device, all the functions of the foregoing positioning device are fully implemented, and the clinical application requirements of sub-lobe resection surgery are met.
[0079] It should be noted that the computer readable medium described in the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium is a tangible medium capable of storing program codes, and specific examples include but are not limited to: an electrical connection device with one or more conductive lines; a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); an optical fiber, a portable compact disk read-only memory (CD-ROM); an optical storage device, a magnetic storage device; and any suitable combination thereof; the computer readable signal medium includes a data signal carrying computer readable program codes in a baseband or as part of a carrier wave, and the signal form can include electromagnetic signals, optical signals, etc. In the embodiments of the present disclosure, the computer readable medium is any medium that contains or stores the computer program described above, which can be used by or in combination with an instruction execution system (such as a processor of a computing device), an apparatus or a device to implement the real-time positioning function of the pulmonary nodule.
[0080] The above merely illustrates the specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present disclosure, which should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for real-time localization of pulmonary nodules in sublobar resection surgery based on artificial intelligence multimodal image fusion technology, characterized in that, include: Acquire core data and basic model data. The core data includes lung medical images and intraoperative laparoscopic dynamic images; the basic model data is a three-dimensional lung model, which includes a three-dimensional model of lung nodules, bronchi, arteries and veins fused together. The core data and basic model are preprocessed to generate initial data for a three-dimensional collapsed lung model; Data preparation and model training were performed on a three-dimensional collapsed lung model, and geometric features, spatial relationship features, and topological features were extracted. The intraoperative laparoscopic dynamic images are preprocessed, and the key features of the lung structure under laparoscopy, including the lung apex, hilum and lung fissure, are identified by the trained two-dimensional lung key point model and lung detector to determine the actual degree of lung tissue collapse. Based on feature matching, the three-dimensional collapsed lung model is fine-tuned to generate a target three-dimensional collapsed lung model that adapts to the actual situation during surgery. By using artificial intelligence multimodal image fusion technology, endoscopic images and target three-dimensional collapsed lung models are fused to generate model vertex-level UV coordinates, and the endoscopic images are accurately mapped onto the surface of the three-dimensional collapsed lung model, realizing the synchronous display of two-dimensional images and three-dimensional models. The coordinates are updated in milliseconds at a frame rate of 30fps or higher. Through multimodal image fusion display, transparency settings, and free adjustment of coordinate orientation, the three-dimensional lung structure under the endoscopic image is clearly presented from various angles. The location of lung nodules is marked with red highlights, and the specific location of lung nodules in the intraoperative scene is determined in real time.
2. The method according to claim 1, characterized in that, The core data and basic model are preprocessed to generate initial data for a three-dimensional collapsed lung model, including: Using point cloud technology, the center of the three-dimensional lung model is set, the distance of each vertex from the center is calculated, and the three-dimensional lung model is shrunk and deformed according to the collapse factor. The collapse factor is set according to the surgical parameters and the actual lung characteristics of the patient.
3. The method according to claim 1, characterized in that, Data preparation and model training were performed on a three-dimensional collapsed lung model, extracting geometric features, spatial relationship features, and topological features, including: The grid data of the three-dimensional collapsed lung model is preprocessed, and model features are generated by a multimodal feature extractor and a geometric feature extractor. The structural classification is completed by a multimodal fusion network and a hierarchical classifier, and the results are corrected by combining the topological relationship. Data preparation and feature extraction were performed on the intraoperative laparoscopic dynamic images. BGR format images were acquired using a 1920×1280 high-definition thoracoscope at a frame rate of 30fps or higher. Gaussian filtering was used to denoise the images, and adaptive grayscale enhancement was used to improve the contrast. BGR to RGB color conversion was completed, and invalid frames that were blurred or had excessive occlusion were marked and re-acquisition was triggered.
4. The method according to claim 1, characterized in that, Intraoperative laparoscopic dynamic images are preprocessed, and key features of laparoscopic lung structures, including the apex, hilum, and fissure, are identified using a trained two-dimensional lung key point model and a lung detector to determine the actual degree of lung tissue collapse, including: By using computer vision technology to train a two-dimensional key point model of the lung, a lung detector is constructed to identify key features of lung structures under endoscopy, including the lung apex, hilum, lung root, anterior / posterior / lower margin, horizontal fissure, and oblique fissure, and to determine the actual degree of lung tissue collapse.
5. The method according to claim 1, characterized in that, Based on feature matching, the 3D collapsed lung model is fine-tuned to generate a target 3D collapsed lung model adapted to the actual intraoperative situation, including: The three-dimensional collapsed lung model was fine-tuned, and the matching relationship between the two-dimensional feature library and the three-dimensional feature library was established. Based on the registration data, the anatomical structural parameters, dynamic response parameters, and structural correlation parameters of the model were adjusted in a targeted manner. The model was corrected by scaling, rotating, and moving operations until the structural positioning deviation was ≤3mm and the morphological fit was ≥95%, ensuring that the model closely matched the actual situation during the operation.
6. A real-time lung nodule localization device based on artificial intelligence multimodal image fusion technology applied to sublobar resection surgery, characterized in that, The device includes: The data acquisition module is used to acquire core data information and basic model data. The core data information includes lung medical images and intraoperative laparoscopic dynamic images; the basic model data is a three-dimensional lung model, which includes a three-dimensional model of lung nodules, bronchi, arteries and veins fused together. The preprocessing module is used to preprocess the core data and basic model to generate initial data for the three-dimensional collapsed lung model. The feature extraction module is used to prepare data and train the three-dimensional collapsed lung model, extracting geometric features, spatial relationship features, and topological features. The model fine-tuning module is used to preprocess intraoperative laparoscopic dynamic images. Through the trained two-dimensional lung key point model and lung detector, it identifies key features of lung structures under laparoscopy, including the lung apex, hilum, and lung fissure, and judges the actual degree of lung tissue collapse. Based on feature matching, it fine-tunes the three-dimensional collapsed lung model to generate a target three-dimensional collapsed lung model that is adapted to the actual intraoperative situation. The multimodal fusion module is used to fuse endoscopic images with a target 3D collapsed lung model through artificial intelligence multimodal image fusion technology, generate model vertex-level UV coordinates, and accurately map the endoscopic images onto the surface of the 3D collapsed lung model to achieve synchronous display of 2D images and 3D models. The positioning and display module is used to update coordinates at millisecond levels at frame rates of 30fps or higher. Through multimodal image fusion display, transparency settings, and free adjustment of coordinate orientation, it clearly presents the three-dimensional lung structure under the endoscopic image from various angles, and highlights the location of lung nodules in red to determine the specific location of lung nodules in the intraoperative scene in real time.
7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; Wherein, the first processor is configured to execute, by executing the executable instructions, a method for real-time localization of pulmonary nodules based on artificial intelligence multimodal image fusion technology applied to sublobar resection surgery, as described in any one of claims 1 to 5.
8. A computing device, the device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, When the computer program instructions are executed by the processor, the device is triggered to execute the method for real-time localization of pulmonary nodules in sublobar resection surgery based on artificial intelligence multimodal image fusion technology as described in any one of claims 1 to 5.