Multi-mode pulmonary nodule benign and malignant judgment method

Through multimodal fusion technology, combined with three-dimensional medical imaging and methylation sequence data, a model for determining benign and malignant pulmonary nodules is constructed, solving the problem of relying on subjective experience and single-modal analysis in traditional methods, achieving higher diagnostic accuracy and lower misdiagnosis rate.

CN120107157APending Publication Date: 2025-06-06CHONGQING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510045747.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The traditional method of determining benign and malignant pulmonary nodules depends on the subjective experience of doctors, lack of quantitative data support, and the accuracy is difficult to guarantee. Single modal analysis cannot fully reflect the biological characteristics of the tumor, which can easily lead to missed diagnosis and false positives.

Method used

Through multimodal fusion and retraining a specific network model, combining three-dimensional medical imaging data and methylated sequence data, imaging omics and methylated sequence features are extracted, and feature selection is performed through cable regression to construct a multimodal fusion judgment model.

Benefits of technology

It significantly improves the accuracy of the determination of benign and malignant pulmonary nodules, reduces the rate of misdiagnosis and misdiagnosis, improves the generalization ability of the model, and reduces the false positive rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107157A_ABST
    Figure CN120107157A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode pulmonary nodule benign and malignant judgment method, which comprises the following steps: data acquisition: acquiring three-dimensional image data of a pulmonary nodule by utilizing medical imaging equipment; meanwhile, circulating free DNA is extracted from a blood sample of the patient, methylation is performed, and methylation sequence data is obtained; data preprocessing: carrying out noise reduction processing on the acquired three-dimensional image data, and accurately marking the pulmonary nodules from surrounding tissues; meanwhile, the methylation sequence data are cleaned, and invalid or wrong data are removed; feature extraction: successively carrying out three-dimensional medical image feature extraction, methylation sequence feature extraction, inhaul cable regression feature extraction and other modal data processing; and constructing a multi-modal fusion judgment model, and performing training optimization on the fused data judgment model by using the labeled benign and malignant pulmonary nodule sample data of the sample. And a specific network model judgment algorithm is retrained through multi-modal fusion, so that the accuracy of judging benign and malignant pulmonary nodules is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image analysis, and in particular provides a multimodal method for determining whether a lung nodule is benign or malignant. Background Art

[0002] With the rapid development of medical imaging technology, especially the advancement of computed tomography (CT) and magnetic resonance imaging (MRI), the field of medical image analysis has ushered in revolutionary changes. In this field, the determination of benign and malignant pulmonary nodules is an important research direction, because pulmonary nodules are one of the most common manifestations of early lung cancer, and lung cancer is one of the leading causes of cancer death worldwide.

[0003] In recent years, multimodal image fusion technology has been increasingly used in the determination of benign and malignant pulmonary nodules. Multimodal fusion technology combines the advantages of different imaging methods, such as the high resolution of CT, the soft tissue contrast of MRI, and the metabolic information of positron emission tomography (PET) to improve the accuracy of diagnosis. In addition, with the development of artificial intelligence technology, machine learning algorithms have been widely used in medical image analysis to improve the detection rate and diagnostic accuracy of pulmonary nodules.

[0004] In the field of benign and malignant determination of lung nodules, traditional methods and single modalities have many limitations. Traditional methods mainly rely on doctors' subjective experience to conduct qualitative analysis of images, lack quantitative data support, and accuracy is difficult to guarantee. Moreover, they are mostly based on a single modality, such as relying solely on image morphology to judge, ignoring information at the molecular level of the tumor.

[0005] Although single-modality CT or MRI image analysis can provide certain morphological and structural information, it does not adequately reflect early microlesions and tumor biological characteristics. For example, some early lung cancers have atypical imaging manifestations, which can easily lead to missed diagnosis. Moreover, single-modality analysis is easily limited by the imaging technology itself, such as the relatively limited resolution of CT for soft tissues. Summary of the invention

[0006] Based on this, the present invention provides a multimodal method for determining the benign or malignant nature of lung nodules, which effectively integrates data such as methylation sequence data and three-dimensional medical imaging data by retraining a specific network model determination algorithm through multimodal fusion, thereby improving the accuracy of determining the benign or malignant nature of lung nodules.

[0007] In order to achieve the above object, the present invention provides a multimodal method for determining benign or malignant pulmonary nodules, comprising:

[0008] Data collection: Medical imaging equipment is used to obtain three-dimensional imaging data of lung nodules. At the same time, circulating free DNA is extracted from patient blood samples and methylated to obtain methylation sequence data.

[0009] Data preprocessing: noise reduction of the collected 3D image data to accurately mark the lung nodules from the surrounding tissues; at the same time, the methylation sequence data is cleaned to remove invalid or erroneous data;

[0010] Feature extraction: 3D medical image feature extraction, methylation sequence feature extraction, Lasso regression feature extraction and other modality data processing are performed successively;

[0011] Construct a multimodal fusion judgment model: compare and select prediction models, splice and fuse the three-dimensional medical image features and methylation sequence features with the patient information modality data, construct a multimodal fusion judgment model, and use the sample's labeled benign and malignant lung nodule sample data to train and optimize the fused data judgment model.

[0012] Furthermore, the 3D medical image data preprocessing process includes:

[0013] Calculate the weighted average of the pixels of the 3D medical image and the pixels in its neighborhood to remove noise, maintain image details and effectively reduce noise interference;

[0014] The linear normalization method is used to map the image pixel values ​​of three-dimensional medical images to a specific interval, so that the image data of different patients are in the same numerical range, which is convenient for subsequent feature extraction and analysis.

[0015] Furthermore, in the process of extracting three-dimensional medical image features, the morphological features, texture features and three-dimensional structure-based features of the lung nodules are extracted based on the collected and preprocessed three-dimensional image data. The morphological features of the lung nodules include size, shape and edge features, the texture features include grayscale co-occurrence matrix and grayscale run-length matrix, and the three-dimensional structure-based features include volume and surface area.

[0016] Furthermore, in the process of methylation sequence feature extraction, the methylation degree vector is obtained as the initial data of the methylation sequence feature. The methylation sequence data is self-learned through Lasso regression to mine key biomarkers, identify the feature subset that has the greatest impact on the judgment of benign and malignant lung nodules, and achieve feature selection and extraction.

[0017] Furthermore, for the extracted 3D medical imaging omics features and methylation sequence data, Lasso regression is used for self-learning to identify the feature subset that has the greatest impact on the determination of benign and malignant pulmonary nodules, achieve feature selection, and compress some unimportant feature coefficients to zero, thereby achieving sparse modeling and feature selection. This method helps to deal with multicollinearity between features and improves the generalization ability of the model.

[0018] Compared with the prior art, the technical advantages of the multimodal method for determining benign or malignant pulmonary nodules provided by the present invention are at least reflected in:

[0019] 1. Multimodal data fusion: By combining radiomics features, methylation sequence features and clinical information, the complementary information of different modalities is fully utilized to improve the accuracy of benign and malignant judgment of lung nodules.

[0020] 2. Improve diagnostic accuracy: Compared with traditional methods and single modality, the multimodal fusion judgment scheme adopted by the present invention shows higher accuracy in judging the benign and malignant nature of lung nodules.

[0021] 3. Strong generalization ability: In the independent validation set and single-center test set, it showed stable performance, and its prediction sensitivity for benign and malignant lung nodules was as high as 99.92%, indicating that the model has good generalization ability.

[0022] 4. Reduce misdiagnosis and missed diagnosis: By integrating molecular and imaging parameters through deep learning, the developed complex classifier is both accurate and robust in distinguishing benign and malignant lung nodules, significantly higher than the performance of individual methylation, protein and imaging models.

[0023] 5. Reduce false positive rate: Although AI is significantly more sensitive than radiologists in detecting lung nodules, the joint diagnosis of AI and radiologists can improve the specificity of diagnosis and reduce the false positive rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0025] Figure 1 The data processing flow chart of the provided multi-modality lung nodule benign and malignant determination method;

[0026] Figure 2 A flowchart of preprocessing each modality of multimodal data of the provided multimodal lung nodule benign or malignant determination method;

[0027] Figure 3 The multimodal fusion training flow chart of the provided multimodal lung nodule benign and malignant determination method;

[0028] Figure 4 The present invention provides a flowchart for a specific example implementation of the multimodal lung nodule benign or malignant determination method. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] The present invention aims to provide a multimodal method for determining whether a lung nodule is benign or malignant. By retraining a specific network model determination algorithm through multimodal fusion, methylation sequence data and three-dimensional medical imaging data are effectively integrated to improve the accuracy of determining whether a lung nodule is benign or malignant. The implementation process is as follows:

[0031] 1. Data Collection

[0032] Three-dimensional medical imaging data acquisition uses advanced CT or MRI equipment to perform high-resolution scans of the patient's lungs to obtain three-dimensional imaging data of lung nodules. Ensure that the scanning process follows strict medical standards to ensure the accuracy and completeness of the imaging data. Three-dimensional medical imaging technologies, such as CT and MRI, provide three-dimensional information of lung nodules, which helps to more accurately evaluate the size, shape, growth rate and other characteristics of nodules. Combining methylation sequence data and three-dimensional medical imaging data, a more accurate model for determining whether lung nodules are benign or malignant can be constructed. This multimodal fusion determination scheme can not only improve the early detection rate of lung cancer, but also reduce the overdiagnosis and treatment of benign lung nodules, thereby optimizing clinical diagnosis and treatment decisions.

[0033] In the 3D radiomics feature extraction, advanced 3D medical imaging technologies, such as CT and MRI, are used to extract the radiomics features of lung nodules. Radiomics features can provide a quantitative description of lung nodules, including information such as shape, texture, and signal intensity distribution. By using pyradiomics tools, all types of radiomics features can be extracted, including first-order statistical features, shape features, and texture features. At the same time, the feature extractor is set, including image normalization, resampling, etc., to adapt to the image data and improve the accuracy of feature extraction.

[0034] Methylation sequence data collection: Circulating free DNA (cfDNA) is extracted from the patient's blood samples, and professional methylation detection technologies, such as methylation-specific PCR (MSP) or bisulfite sequencing, are used to detect the methylation level in cfDNA and obtain methylation sequence data.

[0035] As a biomarker, methylation sequence data shows great potential in the early diagnosis and prognosis assessment of tumors. Studies have shown that the detection of tumor DNA methylation markers has important clinical application value in the early screening and auxiliary diagnosis of lung cancer. By analyzing the methylation level in circulating free DNA (cfDNA), a diagnostic prediction model for early lung cancer can be constructed to improve the diagnostic accuracy of lung cancer.

[0036] (II) Data preprocessing

[0037] like Figure 1 As shown in the figure, 3D medical image data preprocessing: noise reduction processing is performed on the collected 3D image data to remove noise interference in the image and improve image quality. Image annotation is performed to accurately mark the lung nodules from the surrounding tissues to obtain mask data for subsequent feature extraction. The image is then normalized to make the image data of different patients comparable.

[0038] (1) Noise reduction processing

[0039] Principle: Adopt the Non-Local Means Filtering algorithm. This algorithm is based on the principle that similar structures in an image have similar grayscale values. It removes noise by calculating the weighted average of pixels and pixels in its neighborhood, effectively reducing noise interference while maintaining image details.

[0040] Formula: Let the input image be I(x), x be the pixel coordinates in the image, and the calculation formula of the filtered image J(x) be:

[0041] Where N(x) represents the search neighborhood centered at , w(x,y) is the weight function that measures the similarity between the pixel and , and C(x) = ∑ y∈N(x) w(x,y) is a normalization constant. The weight function is defined as:

[0042]

[0043] in represents the weighted Euclidean distance of the gray value vector of pixels in the neighborhood N(x) and N(y) centered on and , a is the standard deviation of the Gaussian kernel, and h is the filtering parameter that controls the filtering intensity.

[0044] (2) Normalization

[0045] Principle: Linear normalization method is used to map image pixel values ​​to a specific interval, so that the imaging data of different patients are in the same numerical range, which is convenient for subsequent feature extraction and analysis.

[0046] Formula: Let the pixel value range in the original image I be [Imax ,I min ], to normalize to the interval [0,1], the calculation formula for the pixel value x of the normalized image J is:

[0047]

[0048] (3) Method for marking nodule areas in NRRD 3D images

[0049] Data preparation and visualization tool selection: First, make sure that medical imaging data containing lung nodules in nrrd format have been obtained. Select appropriate medical imaging visualization software, such as 3D Slicer or ITK-SNAP. These software provide a wealth of functions for image loading, visualization, and annotation operations. Import the nrrd file into the selected visualization software. Adjust the window width and window position of the image in the software so that the lung nodule area can be clearly displayed. This helps to accurately identify and annotate nodules.

[0050] Initial positioning and rough annotation: Use the image browsing tools provided by the software to observe the image in different views such as axial, coronal and sagittal planes to preliminarily locate the position of the lung nodules.

[0051] Select an appropriate annotation tool, such as the freehand tool or the automatic segmentation tool (if applicable). For the freehand tool, manually draw an outline that roughly covers the nodule area based on the morphology of the nodule in different views. During the drawing process, try to stick to the actual boundary of the nodule, but it does not need to be too precise at this time. The purpose is to preliminarily frame the scope of the nodule.

[0052] Through the above denoising, labeling and normalization processing steps, the quality of three-dimensional medical imaging data can be effectively improved, laying a solid foundation for the subsequent accurate extraction of lung nodule features.

[0053] Methylation sequence data preprocessing: Clean the methylation sequence data to remove invalid or erroneous data points. Perform standardization to ensure that the methylation data of different samples are analyzed on the same scale.

[0054] 3.3.3D Medical Image Feature Extraction

[0055] like Figure 2 As shown, three-dimensional medical image feature extraction: Based on three-dimensional image data, multiple features are extracted, including morphological features of lung nodules (such as size, shape, edge features, etc.), texture features (such as gray-level co-occurrence matrix, gray-level run-length matrix, etc.) and features based on three-dimensional structures (such as volume, surface area, etc.).

[0056] (1) 3D medical image feature extraction, operation steps:

[0057] First, make sure the pyradiomics library and its dependencies are properly installed.

[0058] Determine the file path, and then load the 3D medical image data in nrrd format and the corresponding mask data into the feature extractor of pyradiomics.

[0059] Use the functions provided by the pyradiomics library to configure feature extraction parameters. For example, set the parameters for calculating texture features such as gray level co-occurrence matrix (GLCM) and gray level run length matrix (GLRLM), including distance, angle and other parameters to meet different image feature extraction requirements.

[0060] Perform feature extraction operations to obtain feature vectors including morphological features (such as equivalent diameter and sphericity of lung nodules), texture features (such as contrast and correlation calculated by GLCM algorithm, short area emphasis and long area emphasis calculated by GLRLM) and three-dimensional structural features (such as volume and surface area of ​​lung nodules). The feature vector of the three-dimensional medical image obtained after extraction is X=[x 1 ,.....,x n ], where n is the number of features, n is generally 100-150, and is 139 in the present invention.

[0061] (2) Methylation sequence feature extraction

[0062] The methylation degree vector (each of which has a size of (1x697)) has been obtained and is used as the initial data of the methylation sequence feature. Let the methylation sequence feature vector be Y = [y 1 ,......,y n ], n is the dimension of each methylation feature vector, which is 697 in the present invention.

[0063] (3) Lasso regression feature selection (3D medical image features)

[0064] Principle: Lasso Regression adds L to the loss function 1 The regularization term shrinks the coefficients of some features to zero, thereby achieving the purpose of feature selection, avoiding overfitting and improving the generalization ability of the model.

[0065] Formula and operation steps:

[0066] Construct the Lasso regression model, and set the loss function of Lasso regression to be:

[0067]

[0068] Where m is the sample size (the number of cases in this invention), x ijis the jth 3D medical image feature of the i-th sample, β j is the corresponding coefficient, and θ is the regularization parameter, which is used to control the strength of regularization.

[0069] Use a suitable optimization algorithm (the coordinate descent method is used in the present invention) to solve the above loss function and obtain a coefficient vector.

[0070] Standardization: Before Lasso regression, it is very important to standardize the feature data. This is because different features may have different dimensions and value ranges. If they are not standardized, features with larger value ranges may dominate the model training, affecting the accuracy and stability of the model, and also affecting the regularization effect.

[0071] Operation flow: Use the Python library function preprocessing.StandardScaler() to create a standardizer object stand_ (stand_ is the object name, which is a standardizer). This object will calculate the mean and standard deviation of the feature data X. Call the library function stand_.fit_transform(X) method to standardize X. This method first calculates the mean and standard deviation of each feature based on the training data X, and then transforms each feature of each sample.

[0072] The standardized feature data is stored in X_stand (X_stand is the object name, indicating the result of X data after stand_ standardization), with a mean of 0 and a standard deviation of 1.

[0073] Lasso regression model parameter setting and cross-validation, parameter meaning and setting:

[0074] alpha parameter: This is the regularization parameter in Lasso regression. The purpose of regularization is to prevent the model from overfitting and to limit the complexity of the model by adding a penalty term to the loss function. In Lasso regression, the regularization term is the norm of the coefficient vector (i.e., the sum of the absolute values ​​of the coefficients). Smaller values ​​mean weaker regularization, and the model may be more complex and prone to overfitting; larger values ​​lead to stronger regularization, more coefficients may be compressed to zero, and the model may be underfitting. The present invention uses the np.logspace(-3,2,200) of the numpy library function of python to generate 200 logspaces at 10 -3 to 10 2 The alpha value of the logarithmic equidistant distribution is used for subsequent cross-validation to select the optimal one.

[0075] cv parameter: represents the number of cross-validation folds. The present invention sets it to 15, that is, the data set is divided into 15 parts, 14 of which are selected as training sets and 1 as validation sets each time, and training and validation are performed in turn, and finally 15 model performance evaluation indicators (such as mean square error, etc.) are obtained. By averaging these evaluation indicators, the performance of the model on different data subsets can be more robustly evaluated, thereby selecting the optimal model parameters.

[0076] max_iter parameter: specifies the maximum number of iterations of the Lasso regression model during the training process. Since the optimization problem of Lasso regression may be more complicated, multiple iterations are required to converge to a better solution. The present invention is set to 250,000 to ensure that the model has enough iterations to achieve a better convergence effect. If the number of iterations is set too small, the model may not converge to the optimal solution, resulting in poor model performance; but if it is set too large, it may increase the calculation time and have limited improvement in model performance.

[0077] Model training and cross-validation process, use LassoCV (alphas = alpha, cv = 15, max_iter = 250000) to create a Lasso regression cross-validation model object lasso_cv (lasso_cv is the object name, indicating the cross-validation model). This object will use different values ​​internally to train and validate the standardized data set X_stand and the corresponding target variable Y (assuming that Y is a one-dimensional array representing the target value of each sample, such as the benign and malignant labels of lung nodules).

[0078] Call lasso_cv.fit(X_stand,Y) to start the training and cross-validation process. In each round of cross-validation, for each alpha value:

[0079] Use the training set data to train the Lasso regression model and calculate the predicted value of the model on the validation set.

[0080] Calculate evaluation indicators (such as mean square error, etc.) based on the predicted value and the true value of the validation set.

[0081] After cross validation is completed, the lasso_cv object will select the alpha value with the best performance (usually the smallest evaluation index, such as the smallest mean square error) as the regularization parameter of the final model based on the model performance evaluation indicators under different alpha values. At the same time, the model object will also save the corresponding optimal model coefficients and other information.

[0082] Through the above standardization and Lasso regression cross-validation process, the most relevant features for the target variable (the present invention is for the determination of benign and malignant lung nodules) can be selected to a certain extent (by compressing some coefficients to zero), and the appropriate regularization parameters are selected through cross-validation, which improves the generalization ability and prediction accuracy of the model, and provides a reliable model basis for subsequent analysis and prediction. In practical applications, these parameters can also be further adjusted according to specific problems and data set characteristics to obtain better model performance.

[0083] After the above steps, the present invention screens the 139-dimensional medical image features to obtain 9-dimensional features related to benign and malignant lung nodules, so as to facilitate fusion training in the later stage.

[0084] After Lasso self-learning feature selection, the extracted high-dimensional imaging features are self-learned using Lasso regression to identify the feature subset that has the greatest impact on the determination of benign and malignant pulmonary nodules, achieve feature selection, and compress some unimportant feature coefficients to zero, thereby achieving sparse modeling and feature selection. This method helps to deal with multicollinearity between features and improve the generalization ability of the model.

[0085] (4) Lasso regression feature selection (methylation sequence features)

[0086] The Lasso regression model is also constructed for methylation sequence feature selection. The loss function form is similar to the above, but the variables are related to the methylation sequence features:

[0087]

[0088] where z i is the benign or malignant label corresponding to the methylation sequence feature (0 indicates benign, 1 indicates malignant), y ij is the methylation sequence feature of the th sample, γ j is the corresponding coefficient.

[0089] The coefficient vector [y 1 ,......,y n ] and screen out important subsets of methylation sequence features based on the threshold.

[0090] After the above steps, the present invention screens the 697-dimensional methylation sequence features to obtain 11-dimensional features related to benign and malignant pulmonary nodules, so as to facilitate fusion training in the later stage.

[0091] Other modal data processing

[0092] The present invention also has a patient information mode, the specific content of which is the patient's age, gender, etc.

[0093] Since the methylation features and radiomics contents are numerical values, and the patient's gender is text, the present invention uses 1 to represent male and 0 to represent female. In addition, age is also normalized, with 0-85 corresponding to 0-1 in equal proportions.

[0094] The data of other auxiliary modalities can be processed using similar methods and can be fused after being standardized in the final modal fusion.

[0095] For methylation sequence data, the present invention also uses Lasso regression for self-learning to extract features useful for judging benign and malignant pulmonary nodules. This method can identify key biomarkers from high-dimensional methylation data and provide molecular information for early diagnosis of lung cancer.

[0096] (IV) Construction of multimodal fusion judgment model

[0097] Model selection: After comparing multiple models, the present invention finally adopts random forest as the prediction model to train the multiple modal features processed previously.

[0098] Fusion strategy: adopt the splicing fusion strategy (LMF (minimum rank multimodal fusion algorithm), TFN (orthogonal multimodal fusion algorithm) and other methods, the present invention adopts splicing), and fuse the processed three-dimensional medical image features and methylation sequence features obtained in the third step with the patient information modality data. Splice the three standardized features into a joint feature vector and input it into the model for training;

[0099] Model training: Use the sample data of pulmonary nodules with annotated benign and malignant characteristics (including 3D medical imaging data, patient information data, and methylation sequence data) to train the fused data judgment model. By optimizing the model parameters, the model can accurately predict the benign and malignant characteristics of pulmonary nodules based on the input features. Figure 3 As shown, the specific implementation method is as follows:

[0100] (1) Data reading and preprocessing

[0101] Data reading: First determine the path of the data file processed in step 4, and then use pd.read_excel(path) in the python panda library to read the data in the Excel file into the T data frame.

[0102] Separation of features and labels: Because the fused features also contain a column called label, whose value is benign or malignant, it is necessary to delete a specific column (['label']) from T to obtain feature data X (X is the feature name, indicating the feature separated from the label), and use the label column in T (original data) as the target label Y.

[0103] Data division: Use the library function train_test_split to divide the data set into training set (X_train, Y_train) and test set (X_test, Y_test) according to the ratio of test_size = 0.2, and set the random seed random_state = 88 to ensure the consistency of each division result, which is convenient for model comparison and debugging. At the same time, save the test set feature X_test as X_test_stand (X_test_stand is the object name, which means the X_test data after stand_ standardization) for subsequent prediction.

[0104] (2) Model creation and training

[0105] The present invention uses a random forest model to train and predict the fused features. Random forest is an ensemble learning method that improves the accuracy and robustness of prediction by building multiple decision trees and outputting average results. The model can process high-dimensional data and has a good ability to capture the interaction between features, which is suitable for the analysis of multimodal data.

[0106] Create a random forest model: Use Python to install the sklearn library, then instantiate the RandomForestClassifier object model and set a series of parameters, including:

[0107] n_estimators=4000: specifies the number of decision trees in the forest to be 4000. A larger number of trees can increase the complexity and stability of the model, but it will also increase the computational cost.

[0108] criterion = 'entropy': Entropy is selected as the criterion for splitting nodes in the decision tree. It is used to measure the impurity of nodes. The smaller the entropy, the purer the node is and the better the classification effect may be.

[0109] max_depth=2: Limit the maximum depth of the decision tree to 2 to prevent overfitting caused by excessive tree depth. By controlling the depth of the tree, the complexity and generalization ability of the model are balanced.

[0110] min_samples_split = 3: The minimum number of samples required for internal node splitting is 3, avoiding splitting when the number of samples is too small and reducing the impact of noise on the model.

[0111] min_samples_leaf=1: The minimum number of samples required for a leaf node is 1, which controls the number of samples in a leaf node and affects the smoothness and generalization ability of the model.

[0112] max_features = 'log2': The number of features considered when finding the best split is logarithmic (base 2). The number of features is automatically selected to balance model diversity and computational cost.

[0113] bootstrap = True: Use random sampling with replacement (bootstrap) to build a decision tree, increase sample diversity, and improve the generalization ability of the model.

[0114] oob_score = True: Use out-of-bag (OOB) samples to estimate the generalization error of the model, and a certain degree of model evaluation can be performed without an additional validation set.

[0115] n_jobs = -1: Use all available CPU cores for parallel computing to speed up model training (if the computer supports multi-core parallel computing).

[0116] random_state=42: Set the random seed to ensure that the randomness during model training is reproducible and to facilitate comparison of different experimental results.

[0117] verbose=1: Controls the verbosity of log output to 1. During the training process, you can view some basic information, such as the training progress of each tree.

[0118] Model training: Use model.fit(X_train, Y_train) to train the random forest model, input the training set features X_train and labels Y_train into the model, so that the model can learn the patterns and relationships in the data.

[0119] After obtaining the key features of radiomics and methylation sequences, the present invention standardizes these features to eliminate the impact of different dimensions between features so that they can be compared and fused at the same scale. Multimodal feature fusion is performed by combining radiomics features and methylation sequence features as well as basic personal information and family medical history (if any), so that the complementary information of multiple modalities can be fully utilized to improve the accuracy of benign and malignant judgment of pulmonary nodules.

[0120] (3) Model prediction and evaluation

[0121] Prediction: Use the trained model to predict the test set feature X_test_stand, and call model.predict(X_test_stand) to get the predicted label Y_pred (Y_pred is the object name, indicating the predicted output sequence).

[0122] Evaluation and calculation of prediction accuracy: The accuracy between the predicted label Y_pred and the true label Y_test is calculated through the library function accuracy_score(Y_test,Y_pred) to evaluate the overall performance of the model.

[0123] Generate classification report: Use the library function classification_report(Y_test, Y_pred) to generate a detailed classification report, including indicators such as precision, recall, F1-score, etc., and evaluate each category (if it is a multi-classification problem) to understand the classification effect of the model on different categories.

[0124] Construct a confusion matrix: Use confusion_matrix(Y_test,Y_pred) to construct a confusion matrix to intuitively display the correspondence between the model prediction results and the actual results, and help analyze the misclassification of the model, such as the number of true positive examples (TruePositive), false positive examples (FalsePositive), true negative examples (TrueNegative), and false negative examples (FalseNegative).

[0125] (4) Model saving (main function part)

[0126] Use joblib.dump(model,'random_model.pkl') to save the trained random forest model to the file random_model.pkl for subsequent loading and use in other programs.

[0127] Similarly, use joblib.dump(scaler,'scaler_random.pkl') to save the normalizer object scaler to the file scaler_random.pkl for the same normalization process on the new data.

[0128] So far, the model of the present invention has been obtained, and the following is a method for implementing prediction of new data.

[0129] New data prediction (predict_new_data function part)

[0130] Load the model and normalizer: First, use the library functions joblib.load(model_path) and joblib.load(scaler_path) to load the previously saved random forest model and normalizer respectively, where model_path and scaler_path are the paths to the model and normalizer files.

[0131] Load new data: Use pd.read_excel(new_path) to read the new data file into the new_T data frame. If the new data contains a label column, use it as the new target label new_Y; otherwise, assume that the Malignancyclassification column in the data contains Benign and Malignant categories, map it to a numerical value (Benign is mapped to 0, Malignant is mapped to 1) and convert it to float64 type as new_Y. Then, call the frametofloat(new_T) function (assuming that this function is used to convert the data in the data frame to a floating-point type, the specific function may need to be determined according to the actual situation) to process the new_T data frame.

[0132] Data preprocessing and prediction, including:

[0133] Make sure that the new data new_X only contains the same columns as the training set and in the same order, select the same columns as the training set from new_T (train_columns, assuming that the column name list of the training set has been correctly passed in when calling the function) to get new_X (new_X is the object name, indicating the new test data input).

[0134] Use the loaded normalizer scaler to standardize the new data new_X, and call scaler.transform(new_X) to get the standardized new data new_X_stand.

[0135] Use the loaded model model to predict the standardized new data new_X_stand, and call the library function model.predict(new_X_stand) to get the new prediction label new_Y_pred.

[0136] Calculate the prediction accuracy: Use accuracy_score(new_Y, new_Y_pred) to calculate the accuracy between the new data prediction label new_Y_pred and the true label new_Y to evaluate the performance of the model on the new data. Finally, return the new prediction label new_Y_pred and the accuracy.

[0137] like Figure 4 As shown, the method of the present invention is described in detail below in conjunction with specific embodiments.

[0138] Example 1: Data collection and preprocessing

[0139] A certain number (e.g., 100) of patients who were initially clinically diagnosed with pulmonary nodules were selected, and their lung CT image data and blood samples were collected.

[0140] The CT image data are preprocessed using professional image processing software, such as the image processing toolbox in Matlab, to perform noise reduction, segmentation, and normalization operations.

[0141] Methylation detection was performed on cfDNA in blood samples to obtain methylation sequence data, which was then cleaned and standardized.

[0142] Example 2: Feature extraction and model building

[0143] Based on the preprocessed CT image data, the feature extractor of Python's pyraidomics library is used to extract morphological, texture and 3D structural features, and a total of [X] features are obtained (see point 3 of this article for details).

[0144] Extract [Y] features associated with lung nodules from methylation sequence data.

[0145] Lasso regression was used on methylation sequence features and three-dimensional image features to obtain features that were only related to the judgment of benign and malignant status of lung nodules.

[0146] The random forest model is selected as the basic model, and the feature-level fusion strategy is adopted to normalize multiple features and splice them into a joint feature vector, which is input into the model for training. 80% of the sample data is used as the training set and 20% as the test set (refer to point 4 of this article for specific parameter settings).

[0147] Example 3: Model evaluation and optimization

[0148] The trained model is evaluated on the test set to calculate the accuracy, recall and F1 value.

[0149] If the model performance does not meet expectations, adjust the model parameters (such as the maximum depth of the tree and the splitting criteria), increase the amount of training data, and retrain and evaluate.

[0150] Example 4: Clinical application and verification

[0151] The optimized model was applied to the clinical diagnosis of another 50 patients with pulmonary nodules and compared with the traditional diagnostic method (relying solely on imaging diagnosis).

[0152] The accuracy and misdiagnosis rate of statistical model-assisted diagnosis are used to evaluate the effectiveness of the model in clinical practice.

[0153] Through the detailed description of the above embodiments, it can be seen that the feasibility and effectiveness of the methylation sequence enhanced lung nodule image analysis method of the present invention in practical applications can provide a more accurate and reliable means for determining the benign and malignant nature of lung nodules, and is expected to be widely used in clinical practice.

[0154] 5. Clinical application

[0155] Improve the detection rate of early malignant pulmonary nodules: By integrating cfDNA methylation biomarkers, clinical and imaging features, the established multimodal joint diagnosis model can significantly improve the detection rate of early malignant pulmonary nodules. This helps to achieve early detection and early treatment of lung cancer and improve the survival rate of patients.

[0156] Avoid overdiagnosis and treatment of benign pulmonary nodules: This model helps to distinguish benign from malignant pulmonary nodules, reduce unnecessary surgery and treatment for benign pulmonary nodules, and reduce the economic burden and physical harm to patients.

[0157] Assisting clinical diagnosis and treatment decisions: The multimodal fusion model provides a valuable tool to assist doctors in identifying benign and malignant lung nodules and optimizing clinical diagnosis and treatment decisions in the era of precision medicine.

[0158] In summary, the multimodal fusion judgment algorithm provided by this patent, that is, the methylation sequence data and three-dimensional medical image fusion judgment scheme, is of great significance for improving the early diagnosis rate of lung cancer, reducing the misdiagnosis rate and optimizing treatment strategies. With the continuous advancement of technology, the application prospects of multimodal fusion technology in the determination of benign and malignant lung nodules are broad, and it is expected to provide more accurate diagnosis and treatment services for lung cancer patients.

[0159] After clinical application, the provided method improves the detection rate of early malignant pulmonary nodules: by integrating cfDNA methylation biomarkers, clinical and imaging features, the established multimodal joint diagnosis model can significantly improve the detection rate of early malignant pulmonary nodules. This helps to achieve early detection and early treatment of lung cancer and improve the survival rate of patients. Avoid overdiagnosis and treatment of benign pulmonary nodules: The model helps to distinguish between benign and malignant pulmonary nodules, reduce unnecessary surgery and treatment of benign pulmonary nodules, and reduce the economic burden and physical harm of patients. Assist clinical diagnosis and treatment decisions: The multimodal fusion model provides a valuable tool to assist doctors in identifying the benign and malignant nature of pulmonary nodules in the era of precision medicine and optimize clinical diagnosis and treatment decisions.

[0160] A person skilled in the art should understand that the specific implementation modes of the present invention can still be modified or some technical features can be replaced by equivalents without departing from the spirit of the technical solution of the present invention, which should be included in the scope of the technical solution for protection of the present invention.

Claims

1. A multimodal method for determining whether a pulmonary nodule is benign or malignant, characterized in that: include: Data collection: Medical imaging equipment is used to obtain three-dimensional imaging data of lung nodules. At the same time, circulating free DNA is extracted from patient blood samples and methylated to obtain methylation sequence data. Data preprocessing: noise reduction of the collected 3D image data to accurately mark the lung nodules from the surrounding tissues; at the same time, the methylation sequence data is cleaned to remove invalid or erroneous data; Feature extraction: 3D medical image feature extraction, methylation sequence feature extraction, Lasso regression feature extraction and other modality data processing are performed successively; Construct a multimodal fusion judgment model: compare and select prediction models, splice and fuse the three-dimensional medical image features and methylation sequence features with the patient information modality data, construct a multimodal fusion judgment model, and use the sample's labeled benign and malignant lung nodule sample data to train and optimize the fused data judgment model.

2. The multimodal method for determining benign or malignant pulmonary nodules according to claim 1, characterized in that: The 3D medical imaging data preprocessing process includes: Calculate the weighted average of the pixels of the 3D medical image and the pixels in its neighborhood to remove noise, maintain image details and reduce noise interference; The linear normalization method is used to map the image pixel values ​​of three-dimensional medical images to a specific interval, so that the image data of different patients are in the same numerical range, which is convenient for feature extraction and analysis.

3. The multimodal method for determining benign or malignant pulmonary nodules according to claim 2, characterized in that: In the process of extracting three-dimensional medical image features, the morphological features, texture features and three-dimensional structure-based features of lung nodules are extracted based on the collected and preprocessed three-dimensional image data. The morphological features of the lung nodules include size, shape and edge features, the texture features include grayscale co-occurrence matrix and grayscale run-length matrix, and the three-dimensional structure-based features include lung nodule volume and surface area.

4. The multimodal method for determining benign or malignant pulmonary nodules according to claim 3, characterized in that: In the process of methylation sequence feature extraction, the methylation degree vector is obtained as the initial data of the methylation sequence feature. The methylation sequence data is self-learned through Lasso regression to mine key biomarkers, identify the feature subset that has the greatest impact on the judgment of benign and malignant lung nodules, and realize feature selection and extraction.

5. The multimodal method for determining whether a pulmonary nodule is benign or malignant according to claim 3 or 4, characterized in that: For the extracted three-dimensional medical imaging omics features and methylation sequence data, Lasso regression is used for self-learning to identify the feature subset that has the greatest impact on the judgment of benign and malignant pulmonary nodules, thereby achieving feature selection.

6. The multimodal method for determining benign or malignant pulmonary nodules according to claim 5, characterized in that: Use Lasso regression feature selection to select 3D medical image features, including: Construct the Lasso regression model, and set the loss function of Lasso regression to be: Where m is the sample size (the number of cases in this invention), x ij is the jth 3D medical image feature of the i-th sample, β j is the corresponding coefficient, and θ is the regularization parameter, which is used to control the strength of regularization.

7. The multimodal method for determining benign or malignant pulmonary nodules according to claim 5, characterized in that: Lasso regression feature selection for methylation sequence features includes: A Lasso regression model is constructed for methylation sequence feature selection. The loss function form is similar to the above, but the variables are related to the methylation sequence features: where z i is the benign or malignant label corresponding to the methylation sequence feature (0 indicates benign, 1 indicates malignant), y ij is the methylation sequence feature of the th sample, γ j is the corresponding coefficient.

8. The multimodal method for determining benign or malignant pulmonary nodules according to claim 3, characterized in that: The random forest model is used to train and predict the fused three-dimensional image data and methylation sequence data features, and the prediction accuracy and robustness are improved through integrated learning.

9. The multimodal method for determining benign or malignant pulmonary nodules according to claim 8, characterized in that: The process of building a multimodal fusion judgment model includes: Read and preprocess three-dimensional image data and methylation sequence data; Create and train a random forest model, input the training set features and labels into the random forest model to learn patterns and relationships in the data; Use the trained random forest model to predict the test set features, obtain the predicted labels, evaluate the prediction accuracy and the overall performance of the model, generate a classification report and construct a confusion matrix; Save the trained random forest model.

10. The multimodal method for determining benign or malignant pulmonary nodules according to claim 9, characterized in that: After the multimodal fusion judgment model is built, new data is predicted, including: Use the loaded normalizer to standardize the new data to obtain the standardized new data; use the loaded random forest model to predict the standardized new data to obtain a new prediction label; Calculate the accuracy between the predicted labels and the true labels of the new data, evaluate the performance of the model on the new data, and return the new predicted labels and accuracy.

Citation Information

Cited By

  • Lung solid nodule identification method based on multi-modal data fusion

    CN121998963A