Processing method and system for predicting lung function based on PET / CT (positron emission tomography / computed tomography) multi-mode image
Through the Siamese neural network integrating PET/CT multimodal imaging, the problem of difficulty in accurately predicting preoperative lung function in lung cancer patients in the prior art is solved, multi-dimensional lung function evaluation and accurate prediction are achieved, and breath holding dependence is overcome.
Patent Information
- Application Number
- CN202510232484.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The prior art is difficult to accurately predict the preoperative lung function of lung cancer patients, especially for patients with respiratory limitations or physical exhaustion. Traditional methods rely on high breath holding requirements and lack effective lung function evaluation algorithms and models.
Using a multimodal image-based processing method based on PET/CT, the Siamese neural network integrates the characteristics of free breath CT, breath holding CT and PET images to reduce the dependence on breath holding and realize multi-dimensional evaluation of lung function.
Accurate prediction of preoperative lung function in patients with lung cancer is achieved, the breath holding dependence of traditional methods is overcome, and the applicability and prediction accuracy are improved under different clinical conditions.
Smart Images

Figure CN120131046A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent medical treatment, and more specifically, to a processing method and system for predicting lung function based on PET / CT multimodal images. Background Art
[0002] Lung cancer is a malignant tumor with relatively high incidence and fatality rates globally. Preoperative and postoperative lung function assessment is crucial in the treatment of lung cancer. For lung cancer patients, preoperative lung function prediction can not only help doctors formulate personalized surgical plans to ensure the safety of the surgery, but also provide valuable references for postoperative recovery plans, complication prevention, and follow-up management. However, traditional lung function measurement methods, such as forced vital capacity measurement (FVC), one-second rate (FEV1 / FVC), etc., are highly dependent on the patient's subjective cooperation and breathing, which is particularly difficult for elderly, weak, patients with limited respiratory function, or those with severe conditions. These patients often cannot complete the test smoothly due to physical weakness or respiratory limitations, resulting in inaccurate or incomplete lung function assessment results, thus affecting surgical decisions and postoperative management. In addition, traditional lung function tests also have the problem of being inconvenient for continuous postoperative monitoring. Patients often face challenges such as pain and dyspnea after surgery and are difficult to repeat lung function tests, which limits doctors' timely understanding of the patient's lung function recovery and affects the adjustment of subsequent treatment plans.
[0003] As an imaging tool that combines functional metabolic imaging and anatomical structure imaging, PET / CT examination plays an important role in the precise diagnosis, staging, efficacy evaluation, and prognosis monitoring of lung cancer. PET stands for Positron Emission Computed Tomography, and PET images can reflect the metabolic activity of tumor tissues, helping to identify the active areas and potential metastatic foci of tumors. CT stands for Computed Tomography, which is a medical imaging technology. CT images provide high-resolution anatomical structure information, helping to accurately delineate the tumor boundary and the surrounding organ structures. For lung cancer patients who have undergone PET / CT examination before surgery, using these imaging data for pulmonary function prediction can reduce the burden of additional examinations and provide a non-invasive and convenient means of pulmonary function assessment. However, the current PET / CT examination method still faces challenges in preoperative pulmonary function prediction. Especially under preoperative conditions, many lung cancer patients have difficulty holding their breath for a long time due to respiratory restriction or physical weakness, which seriously affects the quality of PET / CT images and thus limits the accuracy of pulmonary function prediction. In addition, the existing PET / CT image analysis methods mainly focus on the tumor itself and lack effective algorithms and models for pulmonary function assessment, making it a technical problem to directly extract pulmonary function information from PET / CT images.
[0004] In related technologies, for example, Chinese Patent CNCN123456789A provides a method for evaluating pulmonary function based on CT images. This method indirectly evaluates pulmonary function by analyzing the density distribution and volume changes of lung CT images, but this method does not consider the metabolic information provided by PET images and has a high requirement for breath-holding, so it is not applicable to preoperative pulmonary function prediction. Although it has promoted the development of pulmonary function assessment technology to a certain extent, it does not give any technical inspiration on how to combine PET / CT images for non-invasive and convenient preoperative pulmonary function prediction and overcome the limitations such as breath-holding difficulties. Summary of the Invention
[0005] 1. Technical problems to be solved
[0006] Aiming at the problem of how to accurately predict the preoperative pulmonary function of lung cancer patients in the existing technology, the present invention provides a processing method and system for predicting pulmonary function based on PET / CT multimodal images. It can utilize the complementary advantages of the information of PET and CT images, and through image processing and machine learning algorithms, achieve accurate prediction of the preoperative pulmonary function of lung cancer patients, while reducing the dependence on breath-holding, and providing a more reliable and convenient means of pulmonary function assessment for the personalized treatment of lung cancer.
[0007] 2. Technical solutions
[0008] The object of the present invention is achieved by the following technical solutions.
[0009] The content part of this application is used to briefly introduce ideas, which will be described in detail in the specific implementation part later. The content part of this application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0010] Some embodiments of this application propose a processing method and system for predicting lung function based on PET / CT multimodal images to solve the technical problems mentioned in the above background art part.
[0011] As the first aspect of this application, some embodiments of this application provide a processing method for predicting lung function based on PET / CT multimodal images, including the following steps: obtaining image data from a database and performing preprocessing to obtain the image data input into the Siamese network; constructing a Siamese network to extract multimodal features from the image data; setting a feature fusion module at the end of the Siamese network to fuse the extracted multimodal features to obtain a multimodal feature vector; using the fused multimodal feature vector as the input of a prediction model to generate a prediction result through the prediction model.
[0012] Furthermore, the image data includes CT image data and PET image data, and the CT image data includes free-breathing CT image data and breath-hold CT image data.
[0013] Furthermore, the preprocessing process includes image format conversion, and metadata in the image data is retained during the image format conversion process; the image format conversion includes NIfTI format conversion, and a direction cosine matrix is used for spatial coordinate transformation. The process is as follows:
[0014] T′ = M × T;
[0015] T represents the original spatial coordinates, M is the direction cosine matrix, and T′ is the transformed spatial coordinates.
[0016] Furthermore, the preprocessing process also includes a process of eliminating differences. For CT image data, a linear normalization method is used, and for PET image data, a standardization method is used. The process includes: traversing the CT image data to find the minimum pixel value min(f CT ) and the maximum pixel value max(f CT ); for each pixel value f CT (x) in the CT image data, perform normalization processing:
[0017]
[0018] f CT ′(x) represents the result after mapping the original pixel value f CT (x) to the interval [0, 1] through a linear normalization formula;
[0019] Traverse the PET image data and calculate the mean μ and standard deviation σ of the pixel values;
[0020] For each pixel value f PET (x) in the PET image data, perform normalization:
[0021]
[0022] f PET ′(x) refers to the result after converting the original pixel value f(x) to a normal distribution with a mean of 0 and a standard deviation of 1 through a standardization formula.
[0023] Furthermore, the Siamese network includes 3 branch structures with shared weights, and each branch is responsible for processing image data of one modality; the front end of each branch includes 4 convolutional layers, and the convolutional layers use 3×3 convolutional kernels.
[0024] Furthermore, the fusion of multi-modal features includes weighted fusion, and the process includes: inputting the feature vectors F CT自由 of the free-breathing CT image data, the feature vectors F CT屏气 of the breath-hold CT image data, and the feature vectors F PET of the PET image data extracted by the Siamese network;
[0025] Set the learning parameters α, β, and γ, and perform weighted fusion on the multi-modal feature vectors to obtain the result value F 加权融合 of the weighted fusion:
[0026] F 加权融合 = αF CT自由 + βF CT屏气 + γF PET ;
[0027] Among them, α + β + γ = 1.
[0028] Furthermore, the initial weights of the learning parameters are set as: α = β = γ = 1 / 3, and during the training process, the weights are automatically adjusted according to the feedback of the loss function to optimize the fused multi-modal feature vectors.
[0029] Furthermore, the fusion of multi-modal features also includes: concatenating the fused multi-modal feature vectors: obtaining the finally fused feature vector F 最终融合 :
[0030] F 最终融合= Concat(Weighted fusion feature or attention fusion feature, F CT自由 , F CT屏气 , F PET );
[0031] Input F 最终融合 into the fully connected layer, perform feature compression and abstraction through the non-linear activation function ReLU to obtain the deep fusion feature F 输出 , as the input of the prediction model:
[0032] F 输出 = ReLU(WFC·F 最终融合 + b);
[0033] WFC is the weight matrix of the fully connected layer, and b is the bias term.
[0034] Furthermore, the weighted mean square error is used as the loss function to measure the contribution of each modality feature to the prediction result, expressed as:
[0035]
[0036] w i is the weight of the i-th modality; y i is the true lung function value; is the predicted value.
[0037] As the second aspect of the present application, some embodiments of the present application provide a system for the processing method of predicting lung function based on PET / CT multimodal images, including a data module: obtaining image data from a database and performing preprocessing to obtain the image data input into the Siamese network; a network construction module: constructing a Siamese network to extract multimodal features from the image data; setting a feature fusion module at the end of the Siamese network to fuse the extracted multimodal features to obtain a multimodal feature vector; a prediction module: using the fused multimodal feature vector as the input of the prediction model and generating a prediction result through the prediction model.
[0038] 3. Beneficial effects
[0039] Compared with the prior art, the advantages of the present invention are as follows: By adopting a Siamese neural network architecture and combining multi-modal inputs of free-breathing CT, breath-hold CT, and PET images, various imaging features can be effectively integrated through the Siamese network; at the end of the Siamese network, feature vectors of different modalities are fused to achieve multi-dimensional and refined evaluation of lung function; the Siamese network can identify and integrate the metabolic features of PET images and the anatomical features of free-breathing and breath-hold CT, enhancing the accuracy of the prediction model; through this multi-modal feature fusion, the present invention can not only effectively overcome the limitation of relying on breath-holding in traditional methods, but also achieve accurate lung function prediction even when free-breathing CT or breath-hold CT images are input alone. This enables the processing method of the present invention to adapt to data acquisition limitations under different clinical conditions, and at the same time, through the design of combining multi-modal and single-modal inputs, the applicability under different clinical conditions is improved, and lung function-related information is captured more comprehensively and accurately, especially providing more robust prediction results under imperfect clinical imaging conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic flowchart of a processing method for predicting lung function based on PET / CT multi-modal images in an embodiment of the present invention;
[0041] Figure 2 It is a schematic diagram of data augmentation in an embodiment of the present invention;
[0042] Figure 3 It is a schematic diagram of contrastive loss in an embodiment of the present invention;
[0043] Figure 4 It is a schematic diagram of the process of converting free-breathing CT images from DICOM format to NIfTI format in an embodiment of the present invention;
[0044] Figure 5 It is a schematic diagram of the comparison between the original and adjusted breath-hold CT images in an embodiment of the present invention;
[0045] Figure 6 It is a schematic diagram of the comparison of the accuracy and loss function curves of experimental schemes A and B in an embodiment of the present invention;
[0046] Figure 7 It is a schematic diagram of the comparison of the accuracy and loss function curves between the Siamese network and the traditional convolutional network during the training process in an embodiment of the present invention;
[0047] Figure 8 It is a schematic diagram showing the influence of weight initialization (mean setting) and adaptive weight adjustment on the prediction accuracy in an embodiment of the present invention;
[0048] Figure 9Schematic diagram of the accuracy change after adding the contrast learning method in an embodiment of the present invention;
[0049] Figure 10 Schematic diagram of the result of lung function prediction by the technical solution of the present invention in an embodiment of the present invention. Detailed implementation manners
[0050] The present invention will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0051] Combined with Figures 1 to 10 , a processing method for predicting lung function based on PET / CT multimodal images of the present invention includes the following steps:
[0052] S1. Obtain image data and preprocess it:
[0053] Obtain image data from a medical image database and preprocess the obtained image data to obtain image data that meets the input requirements of the Siamese network.
[0054] Specifically, the image data is used to store patient information and is usually stored in the DICOM (Digital Imaging and Communications in Medicine) format.
[0055] In a specific embodiment, the image data includes CT (Computed Tomography) image data and PET (Positron Emission Tomography) image data. The CT image data includes free-breathing CT image data and breath-hold CT image data. The information recorded in the image data includes patient ID, scan parameters, and direction cosine matrix. The scan parameters include scan time, voltage, current, etc., and the direction cosine matrix is used to describe the spatial direction of the image.
[0056] In order to adapt to different deep learning frameworks and deep learning models, it is necessary to preprocess the image data. The preprocessing process includes image format conversion, resolution adjustment, slice selection, image registration, and processing to eliminate differences. Through the preprocessing process, it can be ensured that the data meets the input requirements of the Siamese network, improving the stability and prediction accuracy of the deep learning model; at the same time, the preprocessing process helps to reduce noise and redundant information in the data, improving the quality and usability of the data.
[0057] In a specific embodiment, the specific process of preprocessing is as follows:
[0058] S101. Image format conversion
[0059] Convert the DICOM - formatted image data into a standard format suitable for the Siamese network by using medical image - processing tools.
[0060] In this embodiment, the medical image - processing tool can be dcm2nii, and the standard format is the NIfTI format or the PNG format. The NIfTI format is suitable for three - dimensional processing, and the PNG format is suitable for two - dimensional processing.
[0061] First, during the image - format conversion process, the metadata in the DICOM - format file is retained. The metadata includes the orientation matrix, pixel spacing, and patient information. Ensuring the complete retention of the metadata is to ensure the maintenance of the spatial consistency and dimensional accuracy of the image.
[0062] Second, according to the analysis target, that is, according to the subsequent processing requirements for three - dimensional modeling or two - dimensional analysis, select the corresponding format. For example: Select the NIfTI format to correspond to the three - dimensional processing scenario it applies to, retaining the spatial information and gray - level distribution; or select the PNG format to correspond to the two - dimensional processing scenario it applies to or the scenario that requires fast loading, so as to meet different requirements.
[0063] Specifically, when performing the NIfTI - format conversion, use the orientation matrix for spatial - coordinate transformation to ensure that the converted image exactly coincides with the original image in space. The NIfTI - format conversion process is expressed as follows:
[0064] T′ = M×T;
[0065] where T represents the original spatial coordinates, M is the orientation matrix, and T′ is the converted spatial coordinates.
[0066] In a specific embodiment, use the DICOM - format data of free - breathing CT image data and convert it into the NIfTI format through the dcm2nii tool, and verify whether the output file retains spatial consistency. After the conversion is completed, the pixel distribution of the image data should be consistent with the original data, and the orientation matrix has no offset.
[0067] As Figure 4 shown, for the process of converting free - breathing CT image data from the DICOM format to the NIfTI format, the spatial consistency between the converted image data and the original image data is verified by comparing the Volume Information in 3D Slicer. The image data after being converted by the dcm2nii tool demonstrates the spatial consistency before and after the conversion. The data shows that the converted NIfTI - format image data is spatially consistent with the original DICOM - format image data, without any offset or distortion.
[0068] S102. Image resolution adjustment
[0069] The bilinear interpolation method is adopted to adjust the image resolution of the image data. Bilinear interpolation is based on the values of four adjacent pixel points and calculates the value of the target pixel point through linear interpolation.
[0070] During the bilinear interpolation process, the following steps are required: Determine the pixel range that needs to be interpolated according to the target resolution and the original resolution; For each target pixel point, calculate its relative distance relative to the original pixel point; Use the bilinear interpolation formula to calculate the value of the target pixel point according to the values of the adjacent pixel points and the relative distance; Perform interpolation calculations on all target pixel points to generate new image data.
[0071] In a specific embodiment, the calculation process of bilinear interpolation is as follows:
[0072] f(x,y) = (1 - a)(1 - b)f(i,j) + a(1 - b)f(i + 1,j) + (1 - a)bf(i,j + 1) + abf(i + 1,j + 1);
[0073] Among them, (x,y) is the coordinate of the target pixel point, (i,j), (i + 1,j), (i,j + 1) and (i + 1,j + 1) are the coordinates of the four adjacent pixel points around the target pixel point, (a,b) is the relative distance of the target point relative to these four points (the ratio in the horizontal and vertical directions), and f(i,j) is the original pixel value.
[0074] Specifically, after adjusting the resolution, it is necessary to perform consistency verification on the adjusted image data to ensure that its quality meets the requirements of subsequent processing. Use a visualization tool to check the adjusted image data and observe whether it is consistent with the original data. In particular, pay attention to whether the details of key areas such as the lungs are clearly distinguishable.
[0075] In a specific embodiment, the breath-hold CT image data is adjusted from 512×512 to 256×256 pixels, and it is verified that the detailed features of the lung structure are not lost.
[0076] In a specific embodiment, the breath-hold CT image data is adjusted from 512×512 pixels to 256×256 pixels, and the following verification is carried out:
[0077] (1) Detail feature inspection: Check the adjusted image data through a visualization tool and find that the detailed features of the lung structure are well preserved without obvious blurring or distortion. Such as Figure 5As shown, it is a comparison between the original and adjusted breath-hold CT image data. The comparison between the original image data and the adjusted image data in key regions (such as the lungs, bronchi, etc.) shows that the details of the adjusted image data in key regions (such as the lungs, bronchi) are well preserved, and the signal-to-noise ratio and contrast metrics have not changed significantly, effectively guaranteeing the image quality of the image data.
[0078] (2) Quantitative evaluation results: The signal-to-noise ratio and contrast metrics of the image data before and after adjustment were calculated, as Figure 5 shown, these metrics remained at a relatively high level before and after adjustment, indicating that the image quality of the adjusted image data was effectively guaranteed.
[0079] Therefore, through the interpolation method and verification steps, the effective adjustment of the image resolution can be achieved, ensuring that the adjusted image data meets the requirements of subsequent processing in terms of quality and details.
[0080] S103. Slice selection
[0081] CT image data usually contains multiple slices, and the multiple slices of CT image data contain rich anatomical and pathological information. The key information of the lungs is very likely to be concentrated in the slices of specific layers. When processing lung CT image data, accurately selecting representative slices is crucial for subsequent analysis. The process of selecting representative slices is as follows:
[0082] First, use a deep learning model to segment the CT slices to accurately locate the main regions of the lungs, such as the apex of the lungs, the base of the lungs, and the middle lung lobe part.
[0083] In this embodiment, the deep learning model is a convolutional neural network. The segmentation methods include but are not limited to threshold segmentation and region growing method.
[0084] Specifically, threshold segmentation selects the lung region according to the gray value range of the lung tissue in the CT image data. The region growing method starts from the seed points and gradually expands to the regions that meet the conditions according to the gray value similarity, so as to more accurately capture the lung boundary.
[0085] Deep learning segmentation: Use the trained CNN model to perform pixel-level classification on the CT slices to accurately segment the lung region. This method combines image features and context information, and has higher segmentation accuracy and robustness.
[0086] Secondly, select slices. On the basis of automatic segmentation, further select the slices covering the main lung regions according to the lung position.
[0087] In a specific embodiment, the selected slices should meet the following conditions:
[0088] (1) It can cover the complete lung lobe and segment structures, ensuring the comprehensiveness of subsequent analysis;
[0089] (2) Avoid interference factors such as artifacts and noise, ensuring the clarity and accuracy of the slices;
[0090] (3) Select slices that can reflect the overall lung structure and pathological features, such as regions containing major blood vessels, bronchi, and lung parenchyma.
[0091] Specifically, for 3D modeling, 5 - 10 consecutive slices can be stacked to enhance the expression and visualization of lung features.
[0092] In a specific embodiment, a segmentation model is used to select lung CT slices from free - breathing CT image data, which cover the complete lung lobe and segment structures and are ultimately used for lung function feature extraction.
[0093] In a specific embodiment, a segmentation model is used to select lung CT slices from free - breathing CT image data. These slices cover the complete lung lobe and segment structures and can ultimately be used for lung function feature extraction.
[0094] Therefore, by optimizing steps such as CT image data pre - processing, automatic lung region segmentation, and slice selection, key information in lung CT slices can be more accurately selected and utilized.
[0095] S104. Alignment and Registration
[0096] In this embodiment, for the registration between CT and CT, that is, the registration between free - breathing CT and breath - held CT, a feature - point - based registration method is adopted. The feature points include anatomical key points or corner points. First, identify and extract the anatomical key points or corner points in the two CT image data. These feature points are usually located at prominent positions of the lung structure, such as the apex of the lung, the base of the lung, and the bifurcations of the main bronchi. Subsequently, by calculating the transformation matrix (including rotation and translation transformations), precise alignment of the two CT image data is achieved. This process ensures a high degree of anatomical consistency between free - breathing CT and breath - held CT. During the registration process of PET and CT, due to the relatively low resolution and large dynamic range of PET images, the mutual - information - based registration method is used for calculation.
[0097] Specifically, the mutual - information - based registration method evaluates the similarity between CT and PET by calculating the mutual - information value between them. The larger the mutual - information value, the higher the matching degree between the two images. Therefore, by optimizing the mutual - information value and continuously adjusting the transformation parameters (such as rotation, translation, and scaling) of the PET image until the best alignment effect is achieved. Thus, the spatial consistency between PET image data and CT image data is improved through this process.
[0098] In a specific embodiment, the above steps can be used to align PET and CT using the mutual information registration method. Through the careful adjustment and optimization of this process, the highly consistent multimodal images in the lung region can be finally achieved, which also verifies the effectiveness of the mutual information registration method in PET and CT registration.
[0099] S105. Eliminate differences
[0100] In order to eliminate the processing inconsistency caused by the difference in pixel value ranges between different modalities, in this embodiment, the image data is normalized to ensure the consistency and effectiveness of different modality image data in subsequent processing:
[0101] Specifically, the pixel values of CT image data usually have a relatively fixed range, but different scanning parameters or devices may cause differences in pixel value ranges. To eliminate this difference and facilitate subsequent processing, in this embodiment, the linear normalization method is adopted to ensure the consistency and effectiveness of different modality image data in subsequent processing. The process is as follows:
[0102] First, traverse the entire CT image data to find the minimum value min(f CT ) and the maximum value max(f CT );
[0103] For each pixel value f CT (x) in the CT image data, the following formula is used for normalization processing:
[0104]
[0105] Among them, f CT ′(x) represents the result after mapping the original pixel value f CT (x) to the interval [0, 1] through the linear normalization formula.
[0106] Thus, the normalized pixel value f CT ′(x) will fall within the interval [0, 1].
[0107] More specifically, the pixel values of PET image data usually have a large dynamic range, and the pixel value distributions under different scanning conditions may have significant differences. To handle this difference, in this embodiment, the standardization method is adopted to normalize the pixel values of PET image data to the standard normal distribution (mean is 0, standard deviation is 1). The specific process is as follows:
[0108] First, traverse the entire PET image data to calculate the mean μ and standard deviation σ of the pixel values;
[0109] For each pixel value f in the PET image dataPET (x) is normalized using the following formula:
[0110]
[0111] where f PET ′(x) refers to the result after converting the original pixel value f(x) to a normal distribution with a mean of 0 and a standard deviation of 1 through the normalization formula.
[0112] Thus, the normalized pixel value f PET ′(x) will conform to the standard normal distribution, that is, with a mean of 0 and a standard deviation of 1.
[0113] In a deep learning model (such as a convolutional neural network for feature extraction), normalization processing is crucial for improving the stability and convergence speed of the model. Especially for multi-modal image data (such as CT and PET), normalization can eliminate the differences in the pixel value ranges between different modalities, enabling the deep learning model to process data from different modalities more consistently. Normalizing the pixel values of PET image data to a normal distribution with a mean of 0 and a standard deviation of 1 can ensure processing consistency in the deep learning network; at the same time, for CT image data, using a linear normalization method to map its pixel values to the interval [0, 1] can reduce the differences between different modalities. Through normalization processing, the performance and stability of the deep learning model in processing multi-modal image data can be effectively improved.
[0114] S106. Data augmentation
[0115] Random transformations are performed on the image data using data augmentation to improve the generalization ability of the deep learning model. The process of data augmentation includes random rotation, flipping, and random scaling. Specifically, random rotation refers to randomly rotating the image data by an angle to help the deep learning model learn the feature representations of the image data at different angles; flipping includes horizontal flipping and vertical flipping, and the flipping operation is used to increase the diversity of the image data; random scaling is to perform random scale transformation on the image data, and the scaling ratio in this embodiment is 90% - 110%, to simulate images taken at different distances, which can help the deep learning model learn the feature representations of the image data at different scales, thereby enhancing the robustness of the deep learning model to different image data acquisition conditions.
[0116] S107. Slice stacking and multi-dimensional input
[0117] The selected two-dimensional slices are stacked into a three-dimensional input to capture the three-dimensional structural features of the organ; at the same time, combined with the 3D volume information of the PET image data, multi-modal input data is constructed.
[0118] Specifically, the specific process of slice stacking and multi-dimensional input is as follows:
[0119] First, the selected two-dimensional slices are stacked in the order in which they appear in space to form a three-dimensional volume data. This three-dimensional data preserves the spatial relationships between the slices, thus enabling the reflection of the three-dimensional structure of the organ.
[0120] In addition, the 3D volume information of the PET image data is combined to construct multi-modal input data. The PET image data provides additional information about organ function (such as metabolic activity), which complements the structural information of the CT image data.
[0121] Then, the constructed three-dimensional input data and multi-modal input data are used as the input of the deep learning model.
[0122] The preprocessed image data meets the requirements of the deep learning model in terms of format, resolution, slice selection, registration, and normalization. The preprocessing process can ensure the quality and usability of the image data, providing strong support for subsequent deep learning model training and prediction.
[0123] S2. Siamese network feature extraction:
[0124] A Siamese network is constructed to extract features from the preprocessed image data to achieve the extraction of representative multi-modal features.
[0125] Specifically, the Siamese network consists of three branch structures with shared weights, and each branch corresponds to one type of modal image data (free-breathing CT image data, breath-hold CT image data, and PET image data). The front end of each branch in the Siamese network consists of multiple convolutional layers, which are used to extract the edge, texture, and morphological features of the image. The three-branch structure with shared weights is the core of the Siamese network, and each branch shares parameters starting from the convolutional layer to ensure the consistency of the feature space.
[0126] In this embodiment, 4 convolutional layers are set to capture multi-scale spatial features. The convolutional layer uses a 3×3 convolutional kernel; in addition, a max-pooling layer is used to reduce the dimensionality of the features to reduce the computational complexity and retain the main spatial information, enhancing the robustness of the features.
[0127] More specifically, in order to fuse features of different modalities, a feature fusion module (MLP) is set at the end of the Siamese network in this embodiment. This module is used to extract and fuse multi-modal features, thus ensuring the consistency of features of different modalities and enabling the Siamese network to learn the correlation information between different modalities. The feature vectors fused in the feature fusion layer mentioned in step S3 are the multi-modal features extracted and combined through the feature fusion module.
[0128] In a specific embodiment, a Siamese network is built using the PyTorch framework; after constructing the Siamese network, training and optimization are carried out. During the training process, the Adam optimizer is used to update the weights of the Siamese network, and the initial learning rate is set to 0.001.
[0129] Specifically, during the training of the Siamese network, the number of training epochs is set to 50 to ensure that the network can fully converge. At the same time, the Contrastive Loss function is used to measure the feature similarity between different modalities, and the network performance is optimized by weighting the loss contributions of different modalities. The Contrastive Loss function can encourage the feature vectors of similar samples to be close to each other, while the feature vectors of different samples are far from each other.
[0130] Based on the constructed Siamese network, the preprocessed image data is input, that is, the preprocessed free-breathing CT image data, breath-hold CT image data, and PET image data; the feature vectors extracted by each branch are output through the Siamese network.
[0131] More specifically, a data augmentation strategy is adopted for optimization. The input image data is randomly rotated (0 - 15 degrees), flipped, and scaled (±10%) to improve the adaptability of the Siamese network to different image conditions.
[0132] In a specific embodiment, different network parameter configurations are compared and verified during the experiment. Among them, experimental scheme A uses 3 convolutional layers and an average pooling layer; in addition, experimental scheme B, as an optimized scheme, uses 4 convolutional layers and a max pooling layer. The experimental results show that the structure with 4 convolutional layers and a max pooling layer can more effectively capture the spatial features of lung images, thereby improving the performance of the deep learning model. The accuracy of experimental scheme B reaches 92.3%, which is higher than the accuracy of experimental scheme A, which is 87.5%. As Figure 6 shown, it is a comparison of the accuracy and loss function curves of experimental scheme A and scheme B. Among them, scheme B (4 convolutional layers and a max pooling layer) performs better than scheme A (3 convolutional layers and an average pooling layer) in terms of both accuracy and convergence speed. It can be seen that experimental scheme B shows a faster convergence speed and higher accuracy during the training process, further verifying the effectiveness of the optimized scheme.
[0133] Through in-depth analysis of the experimental results, it can be found that increasing the number of convolutional layers and using a max pooling layer can significantly improve the feature extraction ability of the network. As Figure 7As shown, it is a comparison of the accuracy and loss function curves between the Siamese network and the traditional convolutional network during the training process. It can be seen that the Siamese network (blue curve) shows higher accuracy and a more stable decline in loss during the training process, verifying its advantages in feature extraction and convergence speed. According to Figure 7 The comparison results show that the accuracy of the Siamese network finally reached 95.0%, higher than 92.3% of the traditional convolutional network. At the same time, the loss function of the Siamese network shows a more stable and rapid decline during the training process. These results prove that the increase in the number of convolutional layers and the optimization of the pooling layer can effectively improve the network performance. At the same time, operations such as weighted loss function and data augmentation also contribute to improving the stability and generalization ability of the Siamese network.
[0134] By constructing the Siamese network and conducting experimental verification, it can be proven that it can effectively extract the shared features and modality-specific information of free-breathing CT image data, breath-hold CT image data, and PET image data. These high-quality feature inputs provide a solid foundation for subsequent fusion and prediction. At the same time, the optimized design of the Siamese network (such as weighted loss function, data augmentation, increasing the number of convolutional layers, and using max pooling layer) further improves the stability and generalization ability of the network.
[0135] S3. Multi-modal feature fusion:
[0136] In step S2, the multi-modal features of free-breathing CT image data, breath-hold CT image data, and PET image data are extracted through the Siamese network. In the next step, in order to integrate the features from different modalities and further improve the prediction ability of the deep learning model, the present invention introduces a feature fusion layer, which is responsible for effectively fusing the multi-modal features to obtain the final multi-modal feature vector. The feature fusion layer can enable the data of different modalities to complement each other and strengthen the comprehensive utilization ability of the deep learning model for different information. The multi-modal feature fusion includes weighted fusion, attention mechanism, feature splicing and full connection layer processing, and introducing a contrast learning module. The specific process includes:
[0137] S301. Weighted fusion
[0138] Weighted fusion is carried out according to the importance of different modal features.
[0139] First, the feature vectors of free-breathing CT images, breath-hold CT images, and PET images extracted from the Siamese network are denoted as F CT自由 、F CT屏气 and F PET, these feature vectors will be used as inputs and sent to the feature fusion layer for weighted fusion processing to obtain the fused multi-modal feature vectors.
[0140] Introduce learning parameters α, β, and γ, satisfying α + β + γ = 1; perform weighted fusion on the three multi-modal feature vectors to obtain the weighted fusion result value F 加权融合 , and thus set the weights:
[0141] F 加权融合 = αF CT自由 + βF CT屏气 + γF PET ;
[0142] Specifically, the initial weights are set to the mean: α = β = γ = 1 / 3, and the optimal weights are automatically learned through the backpropagation algorithm subsequently.
[0143] During the training process, the weighted mean square error (WMSE) loss function is used to evaluate the prediction error of the deep learning model. This loss function guides the training of the deep learning model by measuring the difference between the predicted value and the true value, thereby optimizing the fused feature vectors. Specifically, the calculation formula of the weighted mean square error loss function is as follows:
[0144]
[0145] where w i is the weight of the i-th modality, y i is the true lung function value, is the predicted lung function value.
[0146] In a specific embodiment, as Figure 8 shown, it demonstrates the influence of weight initialization and adaptive weight adjustment on the prediction accuracy, further verifying the advantages of automatically learning weights. The comparison results between weight initialization (mean setting) and automatic learning show that the adaptive weights can improve the prediction accuracy by approximately 5%. This indicates that allowing the network to automatically learn the weight allocation of different modality features can more effectively integrate multi-modal information.
[0147] S302. Attention mechanism
[0148] An adaptive attention mechanism is introduced in the feature fusion layer to dynamically allocate the importance of modality features. The specific implementation process is as follows:
[0149] First, calculate the attention weights. Calculate the attention weight Ai of each modality feature through the Softmax function, that is, Ai = Softmax(WiFi). Where Wi is the attention weight matrix and Fi is the feature vector of the i-th modality.
[0150] Secondly, according to the calculated attention weights, the feature vectors of each modality are weighted and fused to obtain the fused feature vector F 注意力融合 , which is expressed as follows:
[0151]
[0152] The attention mechanism can dynamically select the optimal modality information according to the image quality and task requirements. Especially in the presence of noise or incomplete images, this mechanism can more effectively utilize the useful modality information, thereby improving the prediction accuracy.
[0153] S303. Feature Concatenation and Fully Connected Layer
[0154] First, perform feature concatenation. Concatenate the fused multi-modal feature vectors: According to the needs of the task, weighted fusion or attention fusion can be selected to fuse the modality features.
[0155] In a specific embodiment, the fused modality features are achieved through weighted fusion. The feature vector obtained by weighted fusion (weighted fusion feature) is concatenated with the feature vectors of the original free-breathing CT image, breath-hold CT image, and PET image to obtain the finally fused feature vector F 最终融合 .
[0156] In a specific embodiment, the fused modality features are achieved through attention fusion. The feature vector obtained by attention fusion (attention fusion feature) is concatenated with the feature vectors of the original free-breathing CT image, breath-hold CT image, and PET image to obtain the finally fused feature vector F 最终融合 .
[0157] Specifically, the concatenation process is expressed as:
[0158] F 最终融合 = Concat(weighted fusion feature or attention fusion feature, F CT自由 , F CT屏气 , F PET );
[0159] Then, the concatenated feature vector is input into the fully connected layer (FC), and feature compression and abstraction are performed through a non-linear activation function. In this embodiment, the non-linear activation function ReLU is adopted, and the process is expressed as follows:
[0160] F 输出 = ReLU(WFC·F 最终融合 + b);
[0161] Among them, WFC is the weight matrix of the fully connected layer, and b is the bias term.
[0162] The final output is the deeply fused feature F 输出 , which is used as the input of the prediction model.
[0163] S304. Multi-modal contrastive learning
[0164] To strengthen the similarity and complementarity of features between modalities, contrastive learning is introduced. The specific process is as follows:
[0165] Construct positive and negative sample pairs. The positive sample pairs come from different modal features of the same object, and the negative sample pairs come from the features of different objects.
[0166] Adopt the contrastive loss function to maximize the similarity of positive sample pairs (i.e., reduce the Euclidean distance between them) while minimizing the similarity of negative sample pairs (i.e., increase the Euclidean distance between them). The contrastive loss function is expressed as follows:
[0167] L = (1 - Y)·max(0, m - D) + Y·D2;
[0168] Among them, Y is the sample label (positive / negative sample), D is the Euclidean distance between features, and m is the distance threshold.
[0169] The multi-modal feature vector obtained by fusing through the above steps is used as the input of the lung function prediction model and trained in combination with the labels of the training data. The prediction model can make more effective use of the relationship information between different modalities.
[0170] Experimental results show that, combined with weighted fusion and injecting contrastive learning, the pre-training accuracy has increased by 8.5%, verifying the effectiveness and robustness of this scheme in improving the accuracy of the prediction model. By enhancing the connection between modalities through contrastive learning, the accuracy of multi-modal feature fusion in lung function prediction is further improved. As Figure 9 shown, after adding the contrastive learning method, the accuracy shows a significant improvement during the training process.
[0171] S4. Prediction:
[0172] According to the multi-modal feature vector fused in step S3, the fused feature vector is used as the input of the prediction model to generate the corresponding lung function evaluation result. The specific process is as follows:
[0173] In the present invention, it is necessary to predict a continuous value (such as a pulmonary function index) from image data, so the pulmonary function assessment is regarded as a regression prediction task. The regression task requires the prediction model to be able to output continuous numerical predictions, rather than just classification labels.
[0174] To more accurately measure the contribution of different modality features to the prediction result, the weighted mean square error (WMSE) is used as the loss function, and its expression is:
[0175]
[0176] where w i is the weight of the i-th modality, indicating the importance of each modality feature to the prediction result; y i is the true pulmonary function value; is the predicted pulmonary function value.
[0177] By optimizing the weight w i the prediction model can balance the information contribution of different modality features, thus ensuring the efficient utilization of fused features.
[0178] Specifically, to further enhance the collaboration of multi-modal features, a contrastive loss function is introduced in this embodiment to measure the feature similarity between modalities. The expression of the contrastive loss is:
[0179]
[0180] where is the similarity label of the sample pair; D is the Euclidean distance between features; margin is the threshold used to control the similarity learning range between features. By introducing the contrastive loss, the feature consistency among free-breathing CT images, breath-hold CT images, and PET images can be effectively improved, thereby enhancing the prediction performance of the prediction model.
[0181] Specifically, to prevent the prediction model from overfitting and improve the generalization ability of the prediction model, an L2 regularization term is added to the loss function in this embodiment. The specific expression is as follows:
[0182]
[0183] where λ is the regularization coefficient used to limit the magnitude of the weight, and w j is the weight parameter in the prediction model. By adding the L2 regularization term, it can effectively prevent the weights of the prediction model from being too large, thus avoiding the occurrence of overfitting.
[0184] In a specific embodiment, to quantitatively evaluate the accuracy of the prediction result, the following metrics are used for quantitative evaluation in this embodiment:
[0185] (1) Mean Squared Error (MSE): Measures the overall difference between the predicted values and the true values. The smaller the value, the better the prediction effect. It is expressed as follows:
[0186]
[0187] (2) Root Mean Squared Error (RMSE): Restores the actual dimension of the error. The closer the value is to 0, the better the prediction effect. It is expressed as follows:
[0188]
[0189] (3) Mean Absolute Error (MAE): Avoids sensitivity to large errors and is suitable for stable error performance. It is expressed as follows:
[0190]
[0191] (4) Coefficient of Determination (R 2 ): Measures the explanatory ability of the prediction model for the data variance. The closer the value is to 1, the more superior the performance of the prediction model. It is expressed as follows:
[0192]
[0193] Specifically, a processing system for predicting lung function based on PET / CT multimodal images includes:
[0194] Data module: Obtains image data from the database and performs preprocessing to obtain the image data input into the Siamese network;
[0195] Network construction module: Constructs a Siamese network to extract multimodal features from the image data; A feature fusion module is set at the end of the Siamese network to fuse the extracted multimodal features to obtain a multimodal feature vector;
[0196] Prediction module: Uses the fused multimodal feature vector as the input of the prediction model and generates a prediction result through the prediction model.
[0197] In a specific embodiment, free-breathing CT image data, breath-hold CT image data, and PET image data are used for training and validating the prediction model. The prediction results generated through the above steps show that the prediction model has high prediction accuracy and stability and can effectively evaluate the lung function status of patients. Through steps such as a weighted regression prediction model, introducing contrast loss, regularization and overfitting prevention, and prediction result evaluation, accurate prediction and evaluation of lung function can be achieved.
[0198] To facilitate the intuitive interpretation of the prediction results, in this embodiment, the pulmonary function prediction results are presented in the forms of scores, charts, and visualization analysis graphs. Meanwhile, the contribution weights of each modal feature are output to explain the importance of different modal features in the prediction results. Figure 10 The figure shows the results of pulmonary function prediction applying this technical solution. On the left is the input data graph of the Siamese network (including free-breathing CT, breath-hold CT, and PET images). On the right, the comparison between the pulmonary function indicators (such as FVC and FEV1) predicted by the prediction model and the true values is shown, further demonstrating the accuracy and robustness of the pulmonary function prediction model.
[0199] By adopting the Siamese neural network architecture, combining the multi-modal inputs of free-breathing CT, breath-hold CT, and PET images, and effectively integrating various image features through the Siamese network, the technical solution of the present invention realizes the multi-dimensional and refined evaluation of pulmonary function. It effectively overcomes the limitation of traditional methods relying on breath-holding and can still achieve accurate pulmonary function prediction when the free-breathing CT or breath-hold CT images are input alone.
[0200] The present invention and its implementation manners are schematically described above. This description is not restrictive. Without departing from the spirit or basic features of the present invention, the present invention can be implemented in other specific forms. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Any reference signs in the claims should not limit the claims involved. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments to this technical solution without creative efforts without departing from the purpose of this creation, they shall fall within the protection scope of this patent. In addition, the word "including" does not exclude other elements or steps, and the word "a" before an element does not exclude including "a plurality of" such elements. The plurality of elements stated in the product claims can also be implemented by one element through software or hardware. The words such as "first" and "second" are used to represent names and do not represent any specific order.
Claims
1. A method for predicting lung function based on PET / CT multimodal images, comprising the following steps: Obtain image data from the database and perform preprocessing to obtain image data input into the Siamese network; Construct a Siamese network to extract multimodal features from image data; set a feature fusion module at the end of the Siamese network to fuse the extracted multimodal features and obtain a multimodal feature vector; The fused multimodal feature vector is used as the input of the prediction model, and the prediction result is generated through the prediction model.
2. The method for predicting lung function based on PET / CT multimodal images according to claim 1, characterized in that: The image data includes CT image data and PET image data, and the CT image data includes free-breathing CT image data and breath-holding CT image data.
3. The method for predicting lung function based on PET / CT multimodal images according to claim 2, characterized in that: The preprocessing process includes image format conversion, during which metadata in the image data is retained; the image format conversion includes NIfTI format conversion, which uses a direction cosine matrix for spatial coordinate transformation. The process is expressed as follows: T′=M×T; T represents the original space coordinates, M is the direction cosine matrix, and T′ is the transformed space coordinates.
4. The method for predicting lung function based on PET / CT multimodal images according to claim 2, characterized in that: The preprocessing process also includes a process of eliminating differences, using a linear normalization method for CT image data and a standardization method for PET image data, and the process includes: Traverse the CT image data and find the minimum pixel value min(f CT ) and the maximum value max(f CT ); For each pixel value f in the CT image data CT (x), and normalize it: f CT ′(x) represents the original pixel value f CT (x) The result after being mapped to the interval [0, 1] by the linear normalization formula; Traverse the PET image data and calculate the mean μ and standard deviation σ of the pixel values; For each pixel value f in the PET image data PET (x), normalized: f PET ′(x) refers to the result of converting the original pixel value f(x) to a normal distribution with a mean of 0 and a standard deviation of 1 through a standardization formula.
5. The method for predicting lung function based on PET / CT multimodal images according to claim 1, characterized in that: The Siamese network includes three branch structures with shared weights, and each branch is responsible for processing image data of one modality; the front end of each branch includes four convolutional layers, and the convolutional layers use 3×3 convolution kernels.
6. The method for predicting lung function based on PET / CT multimodal images according to claim 2, characterized in that: The fusion of the multimodal features includes weighted fusion, and the process includes: inputting the feature vector F of the free breathing CT image data extracted by the Siamese network CT自由 , the feature vector F of breath-hold CT image data CT屏气 and the feature vector F of the PET image data PET ; Set the learning parameters α, β and γ, perform weighted fusion on the multimodal feature vectors, and obtain the weighted fusion result value F 加权融合 : F 加权融合 =αF CT自由 +βF CT屏气 +γF PET ; Among them, α+β+γ=1.
7. The method for predicting lung function based on PET / CT multimodal images according to claim 6, characterized in that: The initial weights of the learning parameters are set to: α = β = γ = 1 / 3. During the training process, the weights are automatically adjusted according to the feedback of the loss function to optimize the fused multimodal feature vector.
8. The method for predicting lung function based on PET / CT multimodal images according to claim 6 or 7, characterized in that: The fusion of the multimodal features further includes: concatenating the fused multimodal feature vectors to obtain a final fused feature vector F 最终融合 : F 最终融合 =Concat(weighted fusion feature or attention fusion feature, F CT自由 , F CT屏气 , F PET ); F 最终融合 The input is sent to the fully connected layer, and the feature is compressed and abstracted through the nonlinear activation function ReLU to obtain the deep fusion feature F 输出 , as the input of the prediction model: F 输出 =ReLU(WFC·F 最终融合 +b); WFC is the weight matrix of the fully connected layer, and b is the bias term.
9. The method for predicting lung function based on PET / CT multimodal images according to claim 8, characterized in that: The weighted mean square error is used as the loss function to measure the contribution of each modal feature to the prediction result, which is expressed as: w i is the weight of the i-th mode; y i is the true lung function value; is the predicted value.
10. A system based on the processing method for predicting lung function based on PET / CT multimodal images according to any one of claims 1 to 9, characterized in that: Including data module: obtain image data from the database and pre-process it to obtain image data input into the Siamese network; Network construction module: construct a Siamese network to extract multimodal features from image data; set a feature fusion module at the end of the Siamese network to fuse the extracted multimodal features and obtain a multimodal feature vector; Prediction module: The fused multimodal feature vector is used as the input of the prediction model to generate prediction results through the prediction model.
Citation Information
Patent Citations
Lung image processing method and device, electronic equipment and storage medium
CN115063433A
Analysis method and device for dyspnea, electronic equipment and storage medium
CN115295144A
Multi-modal remote sensing data change detection method and system based on twin U-Net neural network
CN117372885A
Lung brain image processing method and device, equipment and storage medium
CN117710332A
Head and neck squamous cell carcinoma recurrence rate prediction method, system and equipment based on multi-modal fusion and medium
CN118675731A