A method for detecting lung nodule pathological types and storage medium
By using urinary protein expression profile data and classifier technology, the invasiveness and inconsistent diagnostic results of traditional lung nodule detection have been resolved, enabling non-invasive and reliable pathological type detection and supporting personalized treatment.
Patent Information
- Application Number
- CN202411020775.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-07-29
AI Technical Summary
Traditional methods for detecting pathological types of pulmonary nodules cause physical harm and psychological burden to patients, and the diagnostic results rely on the doctor's personal experience, leading to inconsistent results.
By acquiring patients' urinary protein expression profile data, preprocessing it, and inputting it into a pre-built classifier, artificial intelligence models such as random forest or support vector machine are used for feature matching and prediction to output pathological type classification results.
It achieves non-invasive testing, reduces the physical and psychological burden on patients, provides more consistent and reliable diagnostic results, and supports the development of personalized treatment plans.
Smart Images

Figure CN118942552B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical technology, and more specifically to a method and storage medium for detecting pathological types of pulmonary nodules. Background Technology
[0002] Lung cancer is a major global health threat. Lung nodules are closely related to lung cancer. For patients with lung nodules, the nodules may be benign or malignant; malignant lung nodules are a manifestation of early-stage lung cancer. Therefore, detecting the pathological type of lung nodules is of great significance for the early detection, precise treatment, and prognostic assessment of lung cancer.
[0003] Traditional methods for diagnosing the pathological type of pulmonary nodules primarily rely on histopathological examinations, such as bronchoscopy, fine-needle aspiration biopsy, and open lung biopsy. These traditional histopathological examinations are all invasive procedures and may lead to complications such as bleeding, infection, and pneumothorax (lung collapse). Open lung biopsy, in particular, may also involve risks associated with general anesthesia, bleeding, and postoperative recovery problems, causing harm to the patient's body. Furthermore, invasive procedures can impose a significant psychological burden on patients.
[0004] Furthermore, traditional diagnostic results rely on the individual skill level of doctors, and the results are prone to variation due to the doctor's experience level, failing to provide a uniform and reliable diagnostic outcome.
[0005] The above analysis shows that traditional methods for diagnosing the pathological types of pulmonary nodules have drawbacks, such as causing physical harm and psychological burden to patients, as well as inconsistent diagnostic results. Summary of the Invention
[0006] The purpose of this invention is to provide a method and storage medium for detecting the pathological type of pulmonary nodules, so as to solve the problems of physical harm and psychological burden on patients and inconsistent diagnostic results caused by the above-mentioned traditional methods for detecting the pathological type of pulmonary nodules.
[0007] To achieve the above objectives, in a first aspect, the present invention provides a method for detecting the pathological type of pulmonary nodules, comprising the following steps:
[0008] Acquire urinary protein expression profile data of pulmonary nodules in the individual sample to be analyzed; preprocess the urinary protein expression profile data of pulmonary nodules to obtain standardized data based on the summation of quantitative data; input the standardized data into a pre-constructed classifier; and the pre-constructed classifier performs feature matching and prediction on the input data to output the pathological type classification result of the individual sample to be analyzed.
[0009] Furthermore, the urinary protein expression profile data of the lung nodules were preprocessed to obtain standardized quantitative data based on the summation, including:
[0010] Proteins with frequencies lower than a preset frequency or proteins with frequencies higher than a preset frequency in the urinary protein expression profile data of pulmonary nodules are removed to eliminate redundant proteins and obtain a candidate protein set. Proteins with iFOT values ranked at a preset position in the candidate protein set are selected to obtain a protein set for screening, wherein the iFOT value of each protein is the ratio of the iBAQ value of that protein to the sum of the iBAQ values of all proteins in the candidate protein set. The sum of iFOT values of the protein set for screening is calculated, and then the iFOT value of each protein in the protein set for screening is divided by the sum of iFOT values to obtain the standardized data of quantitative data based on the sum.
[0011] Furthermore, the pre-built classifier is a lung nodule pathological type classifier, which includes a lung nodule pathological benign / malignant classifier or a lung nodule pathological subtype classifier.
[0012] Furthermore, the construction of the pulmonary nodule pathological type classifier includes the following steps:
[0013] A dataset of urinary protein expression profiles of pulmonary nodules is obtained, which is labeled with pulmonary nodule pathological type tags. A portion of the data in the dataset is selected as a training set. The training set is preprocessed to obtain standardized quantitative data based on the sum of data. Each standardized data point is a feature. The features are sorted in a certain way, and k features are selected, where k is a pre-set number of features to be selected. Based on the k features and the pulmonary nodule pathological type tags, a classifier is constructed to obtain the pulmonary nodule pathological type classifier.
[0014] Furthermore, when the constructed classifier is a benign or malignant lung nodule pathology classifier, the lung nodule pathology type label includes benign lung nodules and malignant lung nodules; when the constructed classifier is a lung nodule pathology subtype classifier, the lung nodule pathology type label includes high-risk subtypes and intermediate-low-risk subtypes.
[0015] Furthermore, the features are sorted in a certain way, and k features are selected from them, including:
[0016] The SelectKBest class is used to select features. The F-value, i.e., the variance ratio, of each feature is obtained by ANOVA f_classif analysis. After obtaining the F-value of each feature, the features are sorted from high to low according to their F-values, and then the top k features with the highest F-values are selected.
[0017] Furthermore, the pulmonary nodule urinary protein expression profile dataset includes a multi-center cohort sample dataset and an independent external cohort sample dataset. To prevent model overfitting, the classifier uses five-fold cross-validation, that is, 4 / 5 of the multi-center cohort sample dataset is selected as the training set for classifier training, the remaining 1 / 5 of the multi-center cohort sample dataset is selected for internal classifier validation, and the independent external cohort sample dataset is selected for external classifier validation.
[0018] Furthermore, when constructing the classifier, an artificial intelligence model is used, preferably a random forest or a support vector machine.
[0019] Furthermore, the construction of the classifier also includes the evaluation of the classifier, with evaluation metrics including accuracy, specificity, and AUC, and the optimal classifier is retained based on these metrics.
[0020] Furthermore, the k features include at least one of the following proteins: CBR1, CDH15, DSC3, FRK, GSTP1, HSPD1, LGALS3, LGALS3BP, MYH9, NEU1, PAH, PPIA, SCNN1G, SRC, TKT, TUBA3C, TXN, TUBA1A, VAMP8, MGAM, WASL, KL, HDAC6, TOM1, NAMPT, TUBA1B, ST3GAL6, OLFM4, RAB21, CHMP2B, ATP6V1D, TUBA8, MIOX, ALDH8A1, AGMAT, TUBA1C, TUBA3E, TUBA3D, CLRN3, and LRRK2.
[0021] Furthermore, the k features include at least one of the following proteins: APOC2, APOD, ATP6V1A, B2M, CALML3, CAT, CD9, CDH13, CRABP2, CST3, GLO1, IGFBP6, LTA4H, PAH, PCCB, SERPINI1, PRKACB, PSAP, PTMA, CXCL12, SLC4A1, TSN, UMOD, STK24, ATRN, PTER, NAPSA, ADAMTS1, DPP3, TSPAN9, FRMPD1, RAB21, AHCYL2, QPCT, CHMP2A, CD248, SLC44A2, ITFG1, TMEM192, and UBE2NL.
[0022] In a second aspect, the present invention provides a machine-readable storage medium storing instructions that cause a machine to perform any of the methods described above for detecting the pathological type of pulmonary nodules.
[0023] Thirdly, the present invention provides a computer program product comprising instructions for causing a machine to perform any of the methods described above for detecting pathological types of pulmonary nodules.
[0024] Fourthly, the present invention provides an apparatus for detecting the pathological type of pulmonary nodules, the apparatus comprising a processor and a memory, the memory storing instructions for causing the processor to execute any of the methods described above for detecting the pathological type of pulmonary nodules in this application.
[0025] By employing the above technical solution, the present invention has the following beneficial effects compared with the prior art:
[0026] 1. This invention utilizes urinary protein to detect the pathological type of pulmonary nodules, requiring only a urine sample from the patient, thus achieving a truly non-invasive test. Compared to traditional invasive procedures, this method avoids physical harm and psychological burden on patients, making the testing process more comfortable and safer.
[0027] 2. This invention classifies the pathological types of individual samples to be analyzed by a pre-built classifier, avoiding reliance on the doctor's diagnostic level, helping to reduce diagnostic differences caused by the doctor's experience level, and providing more unified and reliable classification results.
[0028] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0029] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings:
[0030] Figure 1 This is a flowchart illustrating a method for detecting pathological types of pulmonary nodules provided in an embodiment of the present invention.
[0031] Figure 2 This is a schematic diagram illustrating the construction process of the lung nodule pathology type classifier provided in this embodiment of the invention.
[0032] Figure 3 This is a schematic diagram of the internal validation set ROC evaluation of the lung nodule pathology benign and malignant classifier provided in the embodiments of the present invention.
[0033] Figure 4 This is a schematic diagram of the external validation set ROC evaluation of the lung nodule pathology benign and malignant classifier provided in this embodiment of the invention.
[0034] Figure 5This is a schematic diagram of the internal validation set ROC evaluation of the lung nodule pathological subtype classifier provided in this embodiment of the invention.
[0035] Figure 6 This is a schematic diagram of the external validation set ROC evaluation of the lung nodule pathological subtype classifier provided in this embodiment of the invention.
[0036] Figure 7 This is a schematic diagram illustrating the accuracy evaluation of the internal validation set of the lung nodule pathology benign / malignant classifier provided in this embodiment of the invention.
[0037] Figure 8 This is a schematic diagram illustrating the accuracy evaluation of the internal validation set of the lung nodule pathological subtype classifier provided in this embodiment of the invention.
[0038] Figure 9 This is a schematic diagram of the external validation set parameter evaluation of the lung nodule pathology benign and malignant classifier provided in the embodiments of the present invention.
[0039] Figure 10 This is a schematic diagram of the external validation set parameter evaluation of the lung nodule pathological subtype classifier provided in this embodiment of the invention.
[0040] Figure 11 This is a schematic diagram of the external validation set confusion matrix of the lung nodule pathology benign and malignant classifier provided in an embodiment of the present invention.
[0041] Figure 12 This is a schematic diagram of the external validation set confusion matrix of the lung nodule pathological subtype classifier provided in this embodiment of the invention. Detailed Implementation
[0042] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0043] Figure 1 This is a flowchart illustrating a method used to detect the pathological type of pulmonary nodules. (For example...) Figure 1 As shown, the present invention provides a method for detecting pathological types of pulmonary nodules, comprising the following steps:
[0044] S101: Obtain urinary protein expression profile data of lung nodules in the individual sample to be analyzed;
[0045] S102: Preprocess the urinary protein expression profile data of the lung nodules to obtain standardized quantitative data based on the summation;
[0046] S103: Input the standardized data into a pre-built classifier; and
[0047] S104: The pre-built classifier performs feature matching and prediction on the input data and outputs the pathological type classification results of the individual samples to be analyzed.
[0048] It should be noted that feature matching also includes logarithmic transformation of the features.
[0049] Human urine contains over 8,000 proteins, approximately 40% of which originate from plasma proteins. Healthy individuals contain over 1,800 proteins significantly expressed in the lungs. Therefore, this invention uses proteins present in urine to detect the pathological type of pulmonary nodules in patients. Compared to other testing methods, urine collection provides a non-invasive and readily available way to obtain large samples, causing less harm to the patient's body, making the testing process more comfortable and safer, and avoiding physical harm and psychological burden on the patient.
[0050] Traditional diagnostic methods for classifying pulmonary nodule pathological types rely heavily on the physician's skill level, with less experienced physicians exhibiting significantly lower accuracy. This invention, by constructing a pulmonary nodule pathological type classifier, effectively identifies and utilizes proteins in urine, improving the accuracy of pulmonary nodule pathological type classification. This helps reduce diagnostic discrepancies caused by physician experience levels, providing more consistent and reliable classification results, which is of great significance for early detection, precision treatment, and prognostic assessment.
[0051] By classifying the pathological types of patients' lung nodules, doctors can better develop treatment plans based on the classification results, such as deciding whether surgery is needed, assisting in preoperative planning of the surgical scope, identifying high-recurrence groups after surgery, and assisting in the development of postoperative management plans, thereby achieving personalized treatment.
[0052] Preferably, the urinary protein expression profile data of the pulmonary nodules is preprocessed to obtain standardized quantitative data based on the summation, including:
[0053] Proteins with a frequency lower than a preset frequency or proteins with a frequency higher than a preset frequency are removed from the urinary protein expression profile data of the lung nodules, and redundant proteins are removed to obtain a candidate protein set.
[0054] Proteins with iFOT values ranking high in the candidate protein set are selected to obtain a set of proteins for screening. The iFOT value of each protein is the ratio of its iBAQ value to the sum of the iBAQ values of all proteins in the candidate protein set.
[0055] Calculate the sum of iFOT for the protein set to be screened, and then divide the iFOT of each protein in the protein set to be screened by the sum of iFOT to obtain the standardized quantitative data based on the sum.
[0056] It should be noted that the iBAQ value of a protein is the sum of the peak areas of all corresponding peptides of that protein divided by the theoretical number of peptides. It is calculated using the label-free quantitative iBAQ method based on peak area.
[0057] It is understandable that the number of data points contained in the standardized data based on the summation is the same as the number of proteins in the protein set for screening.
[0058] In this embodiment, when removing proteins, proteins with an occurrence frequency of less than 10% are preferentially removed.
[0059] Preferably, the pre-built classifier is a lung nodule pathological type classifier, which includes a lung nodule pathological benign / malignant classifier or a lung nodule pathological subtype classifier.
[0060] Figure 2 This is a flowchart illustrating the construction process of a lung nodule pathology type classifier. For example... Figure 2 As shown, the construction of the pulmonary nodule pathology type classifier includes the following steps:
[0061] S201: Obtain the urinary protein expression profile dataset of pulmonary nodules, which is labeled with the pathological type of pulmonary nodules;
[0062] S202: Select a portion of the data from the urinary protein expression profile dataset of the lung nodules as a training set;
[0063] S203: Preprocess the training set to obtain the quantitative data standardized based on the sum of the training set;
[0064] S204: Each standardized data point is a feature; the features are sorted in a certain way, and k features are selected; where k is a pre-set number of features to be selected; and
[0065] S205: Based on the k features and the pathological type label of the lung nodules, construct a classifier to obtain the pathological type classifier of the lung nodules.
[0066] Each standardized data point is a feature, and each feature corresponds to a protein. These proteins corresponding to the features are also called markers.
[0067] Preferably, when the constructed classifier is a benign or malignant lung nodule pathology classifier, the lung nodule pathology type label includes benign lung nodules and malignant lung nodules;
[0068] When the constructed classifier is a pulmonary nodule pathological subtype classifier, the pulmonary nodule pathological type label includes high-risk subtypes and intermediate-low-risk subtypes.
[0069] In one embodiment of the present invention, when constructing a classifier for benign and malignant pathological changes of pulmonary nodules, the selected k features include at least one of the following proteins: CBR1, CDH15, DSC3, FRK, GSTP1, HSPD1, LGALS3, LGALS3BP, MYH9, NEU1, PAH, PPIA, SCNN1G, SRC, TKT, TUBA3C, TXN, TUBA1A, VAMP8, MGAM, WASL, KL, HDAC6, TOM1, NAMPT, TUBA1B, ST3GAL6, OLFM4, RAB21, CHMP2B, ATP6V1D, TUBA8, MIOX, ALDH8A1, AGMAT, TUBA1C, TUBA3E, TUBA3D, CLRN3, and LRRK2.
[0070] In one embodiment of the present invention, when constructing a classifier for pulmonary nodule pathological subtypes, particularly for invasive pulmonary adenocarcinoma, the selected k features include at least one of the following proteins: APOC2, APOD, ATP6V1A, B2M, CALML3, CAT, CD9, CDH13, CRABP2, CST3, GLO1, IGFBP6, LTA4H, PAH, PCCB, SERPINI1, PRKACB, PSAP, PTMA, CXCL12, SLC4A1, TSN, UMOD, STK24, ATRN, PTER, NAPSA, ADAMTS1, DPP3, TSPAN9, FRMPD1, RAB21, AHCYL2, QPCT, CHMP2A, CD248, SLC44A2, ITFG1, TMEM192, and UBE2NL.
[0071] It is understandable that both the lung nodule pathology benign / malignant classifier and the lung nodule pathology subtype classifier are binary classifiers. The main difference lies in the different pathological type labels and the different k features selected, but the construction steps of the two are the same.
[0072] Preferably, sorting the features in a certain way and selecting k features from them includes:
[0073] Feature selection is performed using the SelectKBest class, and the F-value (variance ratio) for each feature is obtained through ANOVA's f_classif analysis.
[0074] After obtaining the F-value for each feature, the features are sorted from highest to lowest according to their F-values, and then the top k features with the highest F-values are selected.
[0075] Understandably, the F-value reflects the importance of each feature, i.e., the statistical significance of each feature for the pathological label. In this embodiment, the F-value for each feature is calculated using ANOVA's f_classif analysis. Each F-value corresponds to an index. Sorting the F-values involves sorting their indices, and then finding the F-value based on the index. The pre-defined range for the number of features selected, k, is [40, 1500]. In this embodiment, k = 40 is selected.
[0076] Preferably, the pulmonary nodule urinary protein expression profile dataset includes a multicenter cohort sample dataset and an independent external cohort sample dataset;
[0077] To prevent overfitting, the classifier employs five-fold cross-validation, which involves selecting 4 / 5 of the multi-center cohort sample dataset as the training set for classifier training, selecting the remaining 1 / 5 of the multi-center cohort sample dataset for internal classifier validation, and selecting an independent external cohort sample dataset for external classifier validation.
[0078] It should be noted that in practice, the multi-center cohort sample dataset is divided into five subsets of equal size. Four-fifths of the dataset (four subsets) are selected for classifier training, and the remaining subset is used for internal validation. This training process is repeated five times, each time selecting a different subset for internal validation, resulting in five evaluation results used to evaluate the classifier.
[0079] Preferably, when constructing the classifier, an artificial intelligence model is used, preferably a random forest or a support vector machine.
[0080] Preferably, the construction of the classifier also includes the evaluation of the classifier, and the evaluation metrics include accuracy, specificity and AUC, based on which the optimal classifier is retained.
[0081] In constructing the classifiers, this embodiment constructed a total of 7 types, including Support Vector Machine (SVM), Multilayer Perceptron (MLP), Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, and K-Nearest Neighbors.
[0082] Figure 3 This is a schematic diagram of the internal validation set ROC evaluation of the lung nodule pathology benign / malignant classifier. Figure 4 This is a schematic diagram of the ROC evaluation of the external validation set for a lung nodule pathology benign / malignant classifier. Figure 5This is a schematic diagram of the internal validation set ROC evaluation of the pulmonary nodule pathological subtype classifier. Figure 6 This is a schematic diagram of the ROC evaluation of the external validation set for a lung nodule pathological subtype classifier. When selecting the optimal classifier, the AUC (Area Under the Curve) is the primary consideration. The AUC is the area under the ROC curve and the coordinate axes; the larger the area, the higher the AUC score, indicating a better classification performance. Figure 3 , Figure 4 , Figure 5 and Figure 6 As can be seen from the data, when using two validation sets for validation, the AUC of the support vector machine shows the highest score for both the lung nodule pathology benign / malignant classifier and the lung nodule pathology subtype classifier, and its classification performance is better than other classifiers.
[0083] Figure 7 This is a schematic diagram illustrating the accuracy evaluation of the internal validation set of the lung nodule pathology benign / malignant classifier. Figure 8 This is a schematic diagram illustrating the accuracy evaluation of the internal validation set of the pulmonary nodule pathological subtype classifier. Figure 7 and Figure 8 All data are box plots. In this embodiment, the median is primarily used for judgment; a higher median indicates higher classifier accuracy. It can be understood that the median is the median of the five evaluation results obtained from five-fold cross-validation. Figure 7 and Figure 8 As can be seen, the support vector machine still exhibits the highest accuracy.
[0084] It should be noted that even if the AUC is excellent, it should be discarded when other indicators perform poorly. Figure 9 This is a schematic diagram illustrating the evaluation of external validation set parameters for a lung nodule pathology benign / malignant classifier. Figure 10 This is a schematic diagram illustrating the evaluation of external validation set parameters for a lung nodule pathological subtype classifier. From... Figure 9 and Figure 10 As can be seen, the Support Vector Machine (SVM) achieved high scores across all metrics. Therefore, in this embodiment, the SVM is selected as the optimal classifier to be retained. However, the invention is not limited to this; if a classifier with better performance across all metrics is constructed in other embodiments, other types of classifiers can also be retained.
[0085] Figure 11 This is a schematic diagram of the confusion matrix of the external validation set for a lung nodule pathology benign / malignant classifier. Figure 12 This is a schematic diagram of the confusion matrix of the external validation set for the pulmonary nodule pathological subtype classifier. Figure 11 and Figure 12The classifiers used are all support vector machines, where malignant tumors are represented as malignant pulmonary nodules, benign nodules as benign pulmonary nodules, moderately to well-differentiated pulmonary nodules as low- to intermediate-risk subtypes, and poorly differentiated pulmonary nodules as high-risk subtypes. From the two confusion matrices, it can be calculated that the optimal classifier for malignant / benign pulmonary nodules can identify malignant pulmonary nodules with 85% sensitivity and maintain 75% specificity, while the classifier for pulmonary nodule pathological subtypes can identify high-risk subtypes with 66.67% sensitivity and maintain 79.17% specificity, demonstrating excellent classification performance.
[0086] This invention improves the accuracy of the constructed classifier by evaluating it and selecting the optimal classifier. This effectively identifies and utilizes proteins in urine, which is crucial for early detection, precision treatment, and prognostic assessment. Precise pathological type detection helps doctors choose the most appropriate treatment method, avoiding unnecessary surgery and treatment, reducing medical costs, and improving treatment outcomes and patient survival rates.
[0087] This invention also provides a machine-readable storage medium storing instructions that cause a machine to perform the method for detecting pathological types of pulmonary nodules described in this application.
[0088] Those skilled in the art will understand that this application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0089] This application is described with reference to flowchart illustrations and / or block diagrams of methods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0090] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0091] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0092] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0093] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0094] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0095] It should also be noted that the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0096] The acquisition, transmission, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.
[0097] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0098] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for detecting the pathological type of pulmonary nodules, characterized in that, Includes the following steps: Obtain urinary protein expression profile data of lung nodules in the individual samples to be analyzed; The urinary protein expression profile data of the lung nodules were preprocessed to obtain standardized quantitative data based on the summation. The standardized data is then input into a pre-built classifier; as well as A pre-built classifier performs feature matching and prediction on the input data, and outputs the pathological type classification results of the individual samples to be analyzed. The pre-constructed classifier is a pulmonary nodule pathological type classifier, which includes a pulmonary nodule pathological benign / malignant classifier or a pulmonary nodule pathological subtype classifier. The urinary protein expression profile data of the lung nodules were preprocessed to obtain standardized quantitative data based on summation, including: Proteins with a frequency lower than a preset frequency or proteins with a frequency higher than a preset frequency are removed from the urinary protein expression profile data of the lung nodules, and redundant proteins are removed to obtain a candidate protein set. Proteins with iFOT values ranking high in the candidate protein set are selected to obtain a set of proteins for screening. The iFOT value of each protein is the ratio of its iBAQ value to the sum of the iBAQ values of all proteins in the candidate protein set. The sum of iFOT values for the protein set to be screened is calculated. Then, the iFOT value for each protein in the protein set to be screened is divided by the sum of iFOT values to obtain standardized quantitative data based on the sum. The construction of the lung nodule pathology type classifier includes: acquiring a lung nodule urinary protein expression profile dataset labeled with lung nodule pathology type tags, including a multi-center cohort sample dataset and an independent external cohort sample dataset; selecting a portion of data from the lung nodule urinary protein expression profile dataset as a training set; preprocessing the training set to obtain standardized quantitative data based on summation; each standardized data point is a feature, the features are sorted in a certain way, and k features are selected, where k is a pre-set number of feature selections; and constructing a classifier based on the k features and the lung nodule pathology type tags to obtain the lung nodule pathology type classifier. The process of sorting the features in a certain way and selecting k features includes: using the SelectKBest class to select features, calculating the F-value (variance ratio) of each feature through ANOVA's f_classif analysis; and after obtaining the F-value of each feature, sorting the features from highest to lowest F-value, and then selecting the k features with the highest F-values. When the constructed classifier is a benign and malignant lung nodule pathology classifier, the lung nodule pathology type label includes benign lung nodules and malignant lung nodules; When the constructed classifier is a pulmonary nodule pathological subtype classifier, the pulmonary nodule pathological type label includes high-risk subtypes and intermediate-low-risk subtypes. The method further includes, to prevent model overfitting, the classifier uses five-fold cross-validation, that is, selecting 4 / 5 of the multi-center cohort sample dataset as the training set for classifier training, selecting the remaining 1 / 5 of the multi-center cohort sample dataset for internal classifier validation, and selecting an independent external cohort sample dataset for external classifier validation.
2. The method according to claim 1, characterized in that, Artificial intelligence models are used when building the classifier.
3. The method according to claim 1, characterized in that, The construction of the classifier also includes the evaluation of the classifier, with evaluation metrics including accuracy, specificity, and AUC, and the optimal classifier is retained based on these metrics.
4. The method according to claim 1, characterized in that, The k features include at least one of the following proteins: CBR1, CDH15, DSC3, FRK, GSTP1, HSPD1, LGALS3, LGALS3BP, MYH9, NEU1, PAH, PPIA, SCNN1G, SRC, TKT, TUBA3C, TXN, TUBA1A, VAMP8, MGAM, WASL, KL, HDAC6, TOM1, NAMPT, TUBA1B, ST3GAL6, OLFM4, RAB21, CHMP2B, ATP6V1D, TUBA8, MIOX, ALDH8A1, AGMAT, TUBA1C, TUBA3E, TUBA3D, CLRN3, and LRRK2.
5. The method according to claim 1, characterized in that, The k features include at least one of the following proteins: APOC2, APOD, ATP6V1A, B2M, CALML3, CAT, CD9, CDH13, CRABP2, CST3, GLO1, IGFBP6, LTA4H, PAH, PCCB, SERPINI1, PRKACB, PSAP, PTMA, CXCL12, SLC4A1, TSN, UMOD, STK24, ATRN, PTER, NAPSA, ADAMTS1, DPP3, TSPAN9, FRMPD1, RAB21, AHCYL2, QPCT, CHMP2A, CD248, SLC44A2, ITFG1, TMEM192, and UBE2NL.
6. A machine-readable storage medium having instructions stored thereon for causing a machine to perform the method described in any one of claims 1-5 of this application.
Citation Information
Patent Citations
Pancreaticobiliary ampulla carcinoma classification model generation method and image classification method
CN113762395A
Gastric cancer proteomic typing framework identification method based on deep learning feature extraction
CN114550831A