An early lung cancer screening information management system

By designing an early screening information management system for lung cancer and using random forest algorithms and cluster analysis to identify rare lung cancer patterns, the problem of difficulty in accurately identifying rare lung cancer in the existing technology is solved, and the accuracy and efficiency of screening are improved.

CN119833127BActive Publication Date: 2025-06-27THE AFFILIATED HOSPITAL OF PUTIAN UNIV (THE SECOND HOSPITAL OF PUTIAN CITY)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510307592.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify rare lung cancer patterns in early screening of lung cancer, resulting in frequent misdiagnosis and missed diagnosis, delaying the treatment opportunity, and reducing the patient's survival rate.

Method used

A system for early screening information management of lung cancer is designed, including a prediction subsystem, feature substitution subsystem and information summary subsystem. The prediction subsystem uses a random forest algorithm to analyze the probability that characteristic values ​​belong to the rare lung cancer pattern. The feature substitution subsystem recognizes the rare lung cancer pattern through cluster similarity analysis and abnormal score calculation. The information summary subsystem generates information evaluation with the help of a table and uploads it to the hospital information system.

Benefits of technology

It has improved the screening ability of rare lung cancer, accurately identified rare lung cancer patterns, reduced misdiagnosis and misdiagnosis, improved the efficiency and quality of medical services, and provided efficient support for early screening of lung cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119833127B_ABST
    Figure CN119833127B_ABST
Patent Text Reader

Abstract

The present invention discloses an early lung cancer screening information management system, which relates to the field of medical technology. By combining the random forest algorithm, the system can accurately calculate the probability that the characteristic value belongs to the rare lung cancer pattern, construct the probability vector P and complete the feature extraction, just like accurately screening out the features highly related to rare lung cancer from a large amount of data, providing a key basis for subsequent diagnosis. The feature substitution subsystem locates the target patients according to the feature set, and through the clustering similarity analysis and the abnormal score YF calculation, keenly discovers the clustering groups that may represent the rare lung cancer pattern. For example, patients with similar features are grouped together, and the clustering groups with high abnormal scores are very likely to be related to rare lung cancer. On this basis, a combined condition test is performed to accurately identify the probability corresponding to the rare lung cancer pattern, further improving the screening ability of rare lung cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical technology, and particularly to an information management system for early screening of lung cancer. Background Art

[0002] In the broad and crucial field of medical informatization, the information management system for early screening of lung cancer occupies a very key position. Lung cancer, as a malignant tumor with extremely high incidence and mortality rates globally, early screening is of decisive significance for improving the survival rate and cure rate of patients. By calling the data of storage devices in the hospital and carefully sorting out the characteristics of patients, it has become the core support tool in the process of early diagnosis and intervention of lung cancer, and plays an indispensable role in the formulation of medical decisions and the planning of patient treatment plans.

[0003] Currently, early screening of lung cancer faces many challenges. In terms of diagnostic models, ordinary screening methods have serious deficiencies in the ability to identify rare lung cancer patterns, and misdiagnosis and missed diagnosis often occur. For example, the characteristics of many rare lung cancers have subtle differences from those of common types of lung cancer or normal physiological states. Conventional diagnosis is difficult to accurately capture the information contained in these characteristic combinations, making it difficult for patients to obtain timely and accurate diagnosis in the early stage, delaying the treatment opportunity, which will reduce the survival rate of patients, bring a heavy burden to the patient's family and society, and seriously affect the overall effect of early screening of lung cancer and the improvement of the quality of medical services. Summary of the Invention

[0004] In view of the above problems existing in the prior art, the present application provides an information management system for early screening of lung cancer.

[0005] The embodiments of the present disclosure provide an information management system for early screening of lung cancer, including: a prediction subsystem, a feature substitution subsystem, and an information aggregation subsystem;

[0006] The prediction subsystem is used to collect relevant data of each early-stage lung cancer patient in the hospital by using the storage devices in the hospital, generate a lung cancer database, and combine with the random forest algorithm to analyze the probability that the corresponding feature values belong to rare lung cancer patterns, so as to complete feature extraction and construct a feature set;

[0007] The feature substitution subsystem is used to extract the target early-stage lung cancer patients in the hospital, extract the feature sets of the target early-stage lung cancer patients, perform clustering similarity analysis, traverse to obtain the abnormal score YF, and based on the abnormal score YF, execute a combined condition test operation to identify the probability of the occurrence of the corresponding rare lung cancer pattern under the corresponding conditions;

[0008] The information aggregation subsystem is used to perform information aggregation after the test to generate an information evaluation reference table and upload it to the hospital information system.

[0009] Optionally, the prediction subsystem includes a data collection unit, a soft classification unit, and a feature extraction unit;

[0010] The data collection unit is used to collect the basic information, diagnosis information, examination information, and medication records of each early-stage lung cancer patient according to the in-hospital hospital information system, and perform data sharing and interaction with the picture archiving and communication system, laboratory information system, and electronic medical record system respectively according to the call function possessed by the hospital information system, so as to fill in the medical images, pathology, laboratory data, and electronic medical records of each early-stage lung cancer patient. After statistics, a lung cancer database is generated.

[0011] Optionally, the soft classification unit is used to construct a random forest model based on the random forest algorithm, divide the data in the lung cancer database into a training set, a test set, and a validation set, and preset an initial range of the number of decision trees. Use the training set to train the random forest model, and then use the validation set to perform an accuracy A operation on the random forest model to evaluate the performance of the random forest model under different numbers of decision trees. By comparing the accuracy A values under different numbers of decision trees, select the number of decision trees corresponding to the maximum accuracy A value as the optimal number of decision trees M in the random forest;

[0012] According to the optimal number of decision trees, input each feature in the lung cancer database into the root node of each decision tree, and determine the direction of each feature through the feature judgment conditions of each internal node of the decision tree. After traversing, reach the leaf node of the corresponding decision tree. After statistics, obtain the number of features N of the training set included in the leaf node of the corresponding decision tree, and determine the number of features n of the training set included in the leaf node of the corresponding decision tree that belong to each rare lung cancer pattern according to the number of features N of the training set included in the leaf node of the corresponding decision tree. According to the number of features N of the training set included in the leaf node of the corresponding decision tree and the number of features n of the training set included in the leaf node of the corresponding decision tree that belong to each rare lung cancer pattern, analyze the probability situation of the corresponding feature value belonging to the rare lung cancer pattern, and construct a probability vector P, specifically:

[0013] ;

[0014] ;

[0015] In the formula, represents the decision tree 's probability estimate value that the feature x in the lung cancer database belongs to the rare lung cancer pattern ; represents the number of features in the training set included in the leaf node of the decision tree that belong to the rare lung cancer pattern ; represents the decision tree The probability vector of feature x belonging to different rare lung cancer patterns in the lung cancer database; , and respectively represent the probability estimation values of decision tree for feature x belonging to a rare lung cancer pattern in the lung cancer database, the probability estimation values of decision tree for feature x belonging to a rare lung cancer pattern in the lung cancer database, and the probability estimation values of decision tree for feature x belonging to a rare lung cancer pattern in the lung cancer database; m represents the number of types of rare lung cancer patterns; i represents the decision tree number; j represents the rare lung cancer pattern number.

[0016] Optionally, based on the probability vector P, the majority voting method is used to predict the prediction result corresponding to the corresponding feature value, specifically:

[0017] ;

[0018] In the formula, represents the prediction result corresponding to feature x; represents the set of rare lung cancer patterns, = ; represents the average probability given by all decision trees regarding feature x belonging to a rare lung cancer pattern ;

[0019] According to the prediction result and the actual result corresponding to the corresponding feature value, and shuffling the features in the test set, a feature importance score ZP is constructed to analyze the importance of each feature for rare lung cancer screening, specifically:

[0020] ;

[0021] In the formula, represents the importance score of feature x; represents the classification error of the decision tree without shuffling the value of feature x; represents the classification error of the decision tree when shuffling the value of feature x.

[0022] Optionally, the feature extraction unit is used to set an importance threshold according to the feature importance score ZP of different features, compare the feature importance score ZP of the corresponding feature with the importance threshold. If the feature importance score ZP of the corresponding feature exceeds the importance threshold, the corresponding feature is included in the feature set. If the feature importance score ZP of the corresponding feature does not exceed the importance threshold, the corresponding feature is not included in the feature set.

[0023] Optionally, the feature substitution subsystem includes a clustering unit, an anomaly analysis unit, and a recombination analysis unit;

[0024] The clustering unit is used to extract early-stage lung cancer patients in the hospital who contain relevant features in the feature set according to the feature set in the feature extraction unit, label the extracted early-stage lung cancer patients as target early-stage lung cancer patients, and based on the target early-stage lung cancer patients, extract relevant data of the target early-stage lung cancer patients from the lung cancer database. According to the DBSCAN algorithm, cluster the relevant data of the target early-stage lung cancer patients to obtain several clustering groups, and generate clustering information according to the corresponding clustering groups. The clustering information includes the features in clustering group z, the number of features in clustering group z , the average number of features of all clusters , the average distance from clustering group z to its nearest neighbor cluster and the average distance from all clusters to their nearest neighbor clusters .

[0025] Optionally, the anomaly analysis unit is used to traverse all clustering groups according to the clustering information to calculate and obtain an anomaly score YF. The anomaly score YF is specifically obtained through the following formula:

[0026] ;

[0027] In the formula, represents the anomaly score of clustering group z; represents the average number of features of all clusters; represents the number of features in clustering group z; represents the average distance from clustering group z to its nearest neighbor cluster; represents the average distance from all clusters to their nearest neighbor clusters; and represents the weight coefficient, which is used to adjust the importance of the number of features and distance factors in the anomaly score.

[0028] Optionally, the recombination analysis unit is used to obtain the anomaly score YF of the corresponding clustering group respectively according to the acquisition method of the anomaly score YF, extract the clustering group corresponding to the maximum anomaly score YF as the target clustering group, randomly combine the features in the target clustering group, and test the combined conditions after random combination to identify the probability of the occurrence of the corresponding rare lung cancer pattern under the corresponding conditions. Specifically:

[0029] ;

[0030] In the formula, represents the probability of developing the rare lung cancer pattern under the combined condition ; Indicates the probability of a combined condition in the case of a known rare lung cancer pattern ; Indicates the proportion of the rare lung cancer pattern in the overall cases in the historical data; Indicates the probability of a combined condition ;

[0031] Optionally, the information summary subsystem is used to sort the test results according to the recombination analysis unit to generate an information evaluation reference table, which is stored in the hospital information system in the form of an electronic document to assist doctors in analyzing early-stage lung cancer patients.

[0032] The present invention provides an early-stage lung cancer screening information management system, which has the following beneficial effects:

[0033] (1) By combining the random forest algorithm, the system can accurately calculate the probability that the feature value belongs to the rare lung cancer pattern, construct the probability vector P and complete feature extraction, just like accurately screening out the features highly related to rare lung cancer from a large amount of data, providing a key basis for subsequent diagnosis. The feature substitution subsystem locates the target patients according to the feature set, and through clustering similarity analysis and abnormal score YF calculation, sensitively discovers the clustering groups that may represent the rare lung cancer pattern. For example, patients with similar features are grouped together, and the clustering groups with high abnormal scores are very likely to be related to rare lung cancer. On this basis, the combined condition test is performed to accurately identify the probability corresponding to the rare lung cancer pattern, further improving the screening ability for rare lung cancer. The information summary subsystem integrates the tested information to generate an evaluation reference table and uploads it. Doctors can conveniently obtain comprehensive information. For example, when doctors open the system, they can clearly see the comprehensive evaluation of the patients' various test results, quickly make diagnosis and treatment decisions, improve the efficiency and quality of medical services, and help the efficient development of early-stage lung cancer screening work.

[0034] (2) In terms of prediction results, the majority voting method combined with the probability vector P is used to predict the results corresponding to the feature values, which can make full use of the information of multiple decision trees in the random forest. By calculating the average probability that all decision trees give for the feature x belonging to the rare lung cancer pattern, and then taking the rare lung cancer pattern corresponding to the maximum value as the prediction result, it effectively avoids the bias that may be generated by a single decision tree, making the prediction result more accurate and reliable. For example, when facing complex patient feature data, multiple decision trees analyze from different angles, and the comprehensive prediction result can better reflect the real situation and provide more valuable diagnostic references for doctors. In terms of feature importance analysis, by constructing the feature importance score ZP and comparing the classification errors of the decision tree before and after scrambling the feature values, the influence degree of each feature on the rare lung cancer screening can be accurately measured. If a certain feature has a high importance score, it means that it plays a key role in the screening result. For example, among many patient features, the importance scores of some gene features are very high, so these features can be focused on in subsequent diagnoses and studies to improve the pertinence of the screening.

[0035] (3) The clustering unit is guided by the feature set provided by the feature extraction unit, and can efficiently and accurately extract early-stage lung cancer patients with relevant features from many in-hospital patients, mark them as target early-stage lung cancer patients, and obtain their relevant data from the lung cancer database. After clustering these data using the DBSCAN algorithm, the patients are grouped according to feature similarity to generate comprehensive clustering information. This process helps medical staff quickly sort out the feature distribution of rare lung cancer patients and identify patient groups with similar features. For example, in practical applications, it may be found that a group of patients show similar features in aspects such as gene test results, imaging features, and living habits, providing a clear direction for subsequent in-depth analysis.

[0036] (4) The recombination analysis unit accurately locks the target clustering group corresponding to the maximum anomaly score by obtaining the anomaly score YF of each clustering group. This is like finding the key path to the clues of rare lung cancer in the maze of complex patient data. Randomly combine and test the features within the target clustering group, and calculate the probability of developing the rare lung cancer pattern under the corresponding conditions, which can deeply explore the potential relationship between the feature combination and rare lung cancer. For example, by diversely combining the patient's gene test data, imaging features, living habits and other features, it may be found that the association probability between some unique combinations and rare lung cancer far exceeds expectations, providing a very valuable reference direction for clinical diagnosis and further improving the pertinence and accuracy of rare lung cancer diagnosis. The information evaluation clearly presents key data such as the probability of rare lung cancer and the patient clustering situation under different feature combinations through the generation of a table. Whether it is for the preliminary diagnosis of patients in the outpatient clinic or for in-depth discussion of the condition in the ward, doctors can quickly make scientific and reasonable judgments based on this information evaluation table. Brief Description of the Drawings

[0037] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application.

[0038] Figure 1 It is a block diagram of an information management system for early screening of lung cancer according to the present invention. Specific embodiments

[0039] To make the objectives, technical solutions and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the present application in conjunction with the drawings in the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present application.

[0040] Embodiment 1

[0041] As Figure 1 shown, the present invention provides an information management system for early screening of lung cancer, including a prediction subsystem, a feature substitution subsystem and an information aggregation subsystem;

[0042] The prediction subsystem is used to collect relevant data of each early-stage lung cancer patient in the hospital by using the storage device in the hospital, generate a lung cancer database, and analyze the probability of the corresponding feature value belonging to the rare lung cancer pattern in combination with the random forest algorithm, and construct a probability vector P to complete feature extraction and construct a feature set;

[0043] The feature substitution subsystem is used to extract the target early-stage lung cancer patients in the hospital according to the feature set, extract the feature set of the target early-stage lung cancer patients, perform clustering similarity analysis, traverse to obtain the abnormal score YF, and based on the abnormal score YF, execute a combined condition test operation to identify the probability of the corresponding rare lung cancer pattern occurring under the corresponding conditions;

[0044] The information aggregation subsystem is used to perform information aggregation after the test to generate an information evaluation reference table and upload it to the hospital information system.

[0045] It should be noted that for the population not yet diagnosed with lung cancer, the present invention can also be used to identify the probability of the corresponding rare lung cancer pattern occurring under the corresponding conditions to evaluate the risk levels of different types of lung cancer in the population not yet diagnosed with lung cancer, perform hierarchical management according to the risk levels, strengthen monitoring and preventive measures for high-risk populations, such as increasing the screening frequency, providing more stringent lifestyle interventions, etc.; for low-risk populations, the screening intensity can be appropriately reduced to achieve reasonable allocation of medical resources.

[0046] In the embodiments of the present invention, in terms of data utilization, the prediction subsystem calls HIS to collect data and generate a lung cancer database, enabling the originally scattered and messy data to be centrally integrated. For example, it can converge information such as the imaging data, gene detection results, and medical history of patients, facilitating subsequent in-depth analysis and improving the availability and value of the data. By constructing a probability vector P and a feature set through the random forest algorithm, the accuracy and pertinence of feature extraction are further improved. For example, after analyzing a large amount of patient data, features closely related to rare lung cancer patterns are accurately screened out, providing strong support for subsequent diagnosis. The feature substitution subsystem locates target patients based on the feature set. Through clustering similarity analysis and abnormal score calculation, it can sensitively identify clustering groups that may represent rare lung cancer patterns. For example, patients with similar features are grouped together. If a clustering group has a high abnormal score, it means that the combination of features of the patients in this group may be related to rare lung cancer, and then a combined condition test is performed to accurately identify the probability of having rare lung cancer under the corresponding conditions, greatly improving the screening efficiency and accuracy of rare lung cancer. The information aggregation subsystem generates an information evaluation reference table and uploads it to the hospital information system, facilitating doctors to consult at any time. For example, when a doctor opens the system, they can see the comprehensive evaluation of the patient's various test results, quickly master the condition, make more reasonable diagnosis and treatment decisions, improve the efficiency and quality of medical services, and provide convenience for the early screening and treatment of lung cancer.

[0047] Embodiment 2

[0048] As Figure 1 shown, specifically: the prediction subsystem includes a data collection unit, a soft classification unit, and a feature extraction unit;

[0049] The data collection unit is used to collect the basic information, diagnosis information, examination information, and medication records of each early-stage lung cancer patient according to the in-hospital hospital information system, and perform data sharing and interaction with the Picture Archiving and Communication System (PACS), Laboratory Information System (LIS), and Electronic Medical Record System (EMR) respectively according to the call function of the hospital information system, so as to fill in the medical images, pathology, laboratory data, and electronic medical records of each early-stage lung cancer patient. After statistics, a lung cancer database is generated.

[0050] The soft classification unit is used to construct a random forest model according to the random forest algorithm, divide the data in the lung cancer database into a training set, a test set, and a validation set, and preset an initial range of the number of decision trees. The random forest model is trained using the training set, and then the accuracy A of the random forest model is calculated using the validation set to evaluate the performance of the random forest model under different numbers of decision trees. By comparing the accuracy A values under different numbers of decision trees, the number of decision trees corresponding to the maximum accuracy A value is selected as the optimal number of decision trees M in the random forest;

[0051] A decision tree is a model for making decisions based on a tree structure. Each internal node is a test on a feature, each branch is the output of the test, and each leaf node is a class. In a classification task, through a series of tests on the input features, the features are finally classified into a certain class.

[0052] According to the optimal number of decision trees, each feature in the lung cancer database is input to the root node of each decision tree, and through the feature judgment conditions of each internal node of the decision tree, the direction of each feature is determined. After traversal, it reaches the leaf node of the corresponding decision tree (the leaf node is the node in the decision tree without child nodes, which represents the final decision result). After statistics, the number of features N of the training set contained in the leaf node of the corresponding decision tree is obtained, and according to the number of features N of the training set contained in the leaf node of the corresponding decision tree, the number of features n of the training set contained in the leaf node of the corresponding decision tree belonging to each rare lung cancer pattern is determined. According to the number of features N of the training set contained in the leaf node of the corresponding decision tree and the number of features n of the training set contained in the leaf node of the corresponding decision tree belonging to each rare lung cancer pattern, the probability situation of the corresponding feature value belonging to the rare lung cancer pattern is analyzed, and a probability vector P is constructed, specifically:

[0053] ;

[0054] The meaning of this formula is that under the feature combination represented by the leaf node, the probability that the feature belongs to the rare lung cancer pattern is equal to the proportion of the number of features belonging to the rare lung cancer pattern in the total number of features in that leaf node.

[0055] ;

[0056] Specifically, combining the probabilities corresponding to each class gives the entire probability vector , and this probability vector reflects the decision tree 's likelihood estimate of the feature x belonging to each class.

[0057] In the formula, represents the probability estimate value of the decision tree for the feature x in the lung cancer database belonging to the rare lung cancer pattern ; represents the number of features belonging to the rare lung cancer pattern in the training set contained in the leaf node of the decision tree ; represents the probability vector of the decision tree for the feature x in the lung cancer database belonging to different rare lung cancer patterns; , and respectively represent decision trees the probability estimation value that feature x in the lung cancer database belongs to the rare lung cancer pattern of the decision tree the probability estimation value that feature x in the lung cancer database belongs to the rare lung cancer pattern of the decision tree the probability estimation value that feature x in the lung cancer database belongs to the rare lung cancer pattern of the decision tree; m represents the number of types of rare lung cancer patterns; i represents the decision tree number; j represents the rare lung cancer pattern number.

[0058] Among them, the rare lung cancer pattern refers to a pulmonary malignant tumor with a lower incidence and a special pathological type compared with common lung cancers (such as non-small cell lung cancer and small cell lung cancer). These rare types usually have unique histological features, gene mutations, treatment regimens and prognoses, are less common clinically, and may have a poor response to conventional treatments. Rare lung cancer patterns include, but are not limited to, adenoid cystic carcinoma, clear cell carcinoma of the lung, alveolar soft part sarcoma, and hemangiopericytoma of the lung, etc.

[0059] In the embodiment of the present invention, the data collection unit relies on the powerful calling function of the hospital information system to realize the integration of multi-source data. Obtain the patient's lung images from the picture archiving and communication system, which can intuitively display the lung lesions; collect pathological and laboratory data from the laboratory information system, such as the values of tumor markers, to provide a quantitative basis for diagnosis; integrate the past medical history and medication records from the electronic medical record system to construct a comprehensive patient portrait. Through data sharing and interaction, fill in the data gaps and generate a complete lung cancer database. For example, for a patient with early-stage lung cancer, the information such as the patient's lung CT images, blood tumor marker test results, and past treatment medications are integrated by the system, providing a rich data basis for subsequent analysis. The soft classification unit divides the training set, test set and validation set, pre-sets the range of the number of decision trees, and conducts multi-dimensional evaluations on the random forest model. After calculating the accuracy rate under different numbers of decision trees, the optimal number of decision trees M is selected to ensure the optimal performance of the model. Then, according to the optimal number of decision trees, the features in the lung cancer database are input into the decision tree, and through node feature judgment, it accurately traverses to the leaf node.

[0060] Taking the rare adenocarcinoma of the lung pattern as an example, in a certain decision tree leaf node, it is statistically found that the total number of features N is 100, and the number of features n belonging to the rare adenocarcinoma of the lung pattern is 30. According to the formula, the probability that the features of this leaf node belong to the rare adenocarcinoma of the lung pattern is calculated to be 30%. Combining the probabilities of each category into a probability vector P provides a quantitative basis for the accurate judgment of the rare lung cancer pattern, and further improves the recognition ability and accuracy of the rare lung cancer pattern in the early screening of lung cancer.

[0061] Example 3

[0062] As shown Figure 1 below, specifically: based on the probability vector P, the majority voting method is used to predict the prediction result corresponding to the corresponding feature value, specifically:

[0063] ;

[0064] In the formula, represents the prediction result corresponding to the feature x; represents the set of rare lung cancer patterns, = ; represents the average probability given by all decision trees that the feature x belongs to the rare lung cancer pattern ; argmax(*) represents the value of the independent variable that makes the function reach the maximum value, which is a mathematical term;

[0065] wherein, the average probability given by all decision trees that the feature x belongs to the rare lung cancer pattern is specifically obtained through the following formula:

[0066] ;

[0067] In the formula, M represents the number of optimal decision trees;

[0068] It should be noted that by averaging the probabilities given by all decision trees that the feature x belongs to each lung cancer type, the results of multiple decision trees can be comprehensively considered, and the possibility of the feature x corresponding to different lung cancer types can be analyzed more comprehensively.

[0069] According to the prediction result and the actual result corresponding to the corresponding feature value, and shuffling the features in the test set, the feature importance score ZP is constructed to analyze the importance of each feature for rare lung cancer screening, specifically:

[0070] ;

[0071] In the formula, represents the importance score of the feature x; represents the classification error of the decision tree without shuffling the value of the feature x; represents the classification error of the decision tree when shuffling the value of the feature x;

[0072] It should be noted that the steps for obtaining the classification error of the decision tree without shuffling the value of the feature x are as follows: traverse each decision tree in the random forest, use the test set data for prediction, and obtain the prediction result (the prediction result corresponding to the feature x ), calculate the error between the predicted result and the true label (true result), that is, the proportion of misclassified samples, and take the average of the classification errors of each decision tree as the classification error of the entire random forest without shuffling the values of feature x. .

[0073] The steps to obtain the classification error of the decision tree when shuffling the values of feature x are as follows: First, select the feature x to be evaluated, specified by the feature index. Then, traverse each decision tree in the random forest, randomly shuffle the values of feature x in the test set data to obtain a shuffled feature matrix, and use the shuffled feature matrix for prediction to get the predicted result. Calculate the error between the predicted result and the true label (true result), that is, the proportion of misclassified samples. Take the average of the classification errors of each decision tree after shuffling feature x as the classification error of the entire random forest when shuffling the values of feature x. . .

[0074] Among them, the specific process of the shuffled feature matrix is as follows: Determine the index position of the feature x to be shuffled in the feature matrix, extract the column vector corresponding to feature x from the feature matrix, and perform a random shuffling operation on the extracted column vector. Here, a random number generator can be used to generate a random permutation index sequence with the same length as the column vector, and then rearrange the elements in the column vector according to this index sequence. Assume the length of the column vector is Y, and the generated random permutation index sequence is , ,..., , then the a-th element of the shuffled column vector is the original column vector's -th element. Replace the shuffled column vector back to the column position corresponding to feature x in the original feature matrix to obtain the shuffled feature matrix.

[0075] Specifically, the higher the score of the feature importance, the greater the impact of the feature on rare lung cancer screening. During the model training process, the most important features can be selected according to the feature importance score to reduce the complexity of the model and improve the performance of the model.

[0076] Among them, by introducing sample data of rare lung cancer types for model training and optimization, it prepares for improving the sensitivity of the system to rare lung cancer types, so as to be able to more accurately identify rare lung cancer patients, reduce the missed diagnosis and misdiagnosis of rare lung cancer patients, avoid unnecessary further examinations and treatments, and improve the utilization efficiency of medical resources.

[0077] The feature extraction unit is used to set an importance threshold according to the feature importance score ZP of different features, compare the feature importance score ZP of the corresponding feature with the importance threshold. If the feature importance score ZP of the corresponding feature exceeds the importance threshold, the corresponding feature is included in the feature set. If the feature importance score ZP of the corresponding feature does not exceed the importance threshold, the corresponding feature is not included in the feature set.

[0078] In the embodiment of the present invention, the majority voting method is used to predict the result corresponding to the feature value based on the probability vector P, which can effectively integrate the judgments of multiple decision trees. Different decision trees classify and predict the feature x from their respective perspectives. By calculating the average probability, various judgment bases are synthesized. For example, in a set of feature data of lung cancer patients, multiple decision trees respectively give the probabilities that the feature x belongs to different rare lung cancer patterns. After average calculation, the argmax function is used to select the rare lung cancer pattern with the largest average probability as the prediction result, which makes the prediction result more reliable and reduces the deviation that may occur in a single decision tree. The process of constructing the feature importance score ZP can accurately evaluate the role of each feature in the screening of rare lung cancer. By comparing the classification errors of the decision tree before and after scrambling the feature values, the feature importance is quantified. If the classification error of a certain feature x increases significantly after scrambling, it indicates that this feature has a great impact on classification and its importance score is high. For example, after scrambling the feature "a certain specific gene mutation situation", the classification error of the decision tree for the rare lung cancer pattern increases from 5% to 30%, indicating that this feature is crucial in the screening of rare lung cancer. The feature extraction unit screens features according to the feature importance score ZP, sets an importance threshold, and includes the features that exceed the threshold in the feature set, which effectively reduces the number of features and removes the features with less impact on screening. When constructing the screening model, unnecessary interference factors are reduced, the model complexity is lowered, and the operation efficiency is improved. At the same time, retaining the key features helps to improve the accuracy of the model for screening rare lung cancer, enables the model to focus more on the core influencing factors, provides more targeted diagnostic basis for doctors, and facilitates the efficient development of early lung cancer screening work.

[0079] Example 4

[0080] As Figure 1 shown, specifically: the feature substitution subsystem includes a clustering unit, an anomaly analysis unit, and a recombination analysis unit;

[0081] The clustering unit is used to extract early-stage lung cancer patients in the hospital who correspond to the relevant features in the feature set according to the feature set in the feature extraction unit, and label the extracted early-stage lung cancer patients as target early-stage lung cancer patients. Based on the target early-stage lung cancer patients, relevant data of the target early-stage lung cancer patients are extracted from the lung cancer database, and according to the DBSCAN algorithm, the relevant data of the target early-stage lung cancer patients are clustered to obtain several clustering groups. According to the corresponding clustering groups, clustering information is generated. The clustering information includes the features in clustering group z, the number of features in clustering group z , the average number of features of all clusters , the average distance from clustering group z to its nearest neighbor cluster and the average distance from all clusters to their nearest neighbor clusters .

[0082] The anomaly analysis unit is used to traverse all clustering groups according to the clustering information to calculate and obtain the anomaly score YF. The anomaly score YF is specifically obtained through the following formula:

[0083] ;

[0084] In the formula, represents the anomaly score of clustering group z; represents the average number of features of all clusters; represents the number of features in clustering group z. Clusters with fewer features may be more prone to anomalies, reflecting their certain rarity in terms of quantity. In lung cancer screening, the features corresponding to rare lung cancer patterns are often in the minority in the overall population; represents the average distance from clustering group z to its nearest neighbor cluster. A large distance indicates that this cluster is significantly different from other clusters; represents the average distance from all clusters to their nearest neighbor clusters; and represents the weight coefficient, which is used to adjust the importance of the number of features and distance factors in the anomaly score. When has a large value, it indicates that the separation degree of this cluster from other clusters is relatively high, and its features are significantly different from those of other clusters. Rare lung cancer patterns usually have obvious differences in feature manifestations from common lung cancer patterns or normal conditions;

[0085] The anomaly score of clustering group z measures the anomaly degree of clustering group z relative to other clustering groups. The higher the anomaly score, the more likely this clustering group represents a rare lung cancer pattern.

[0086] In the embodiments of the present invention, the clustering unit accurately locates early-stage in-hospital target lung cancer patients based on the feature set obtained by the feature extraction unit. After extracting relevant data, the DBSCAN algorithm is used for clustering processing. In this way, patients with similar features can be grouped together to form different clustering groups, and comprehensive clustering information is generated. This helps doctors quickly identify patient groups with similar features and improve the understanding of the feature distribution of lung cancer patients. For example, through cluster analysis, it may be found that patients with certain specific gene features and similar living habits are grouped together, providing a basis for subsequent targeted research and treatment. The anomaly analysis unit calculates the anomaly score YF according to the clustering information. This score comprehensively considers the number of features in the clustering group and the average distance to the nearest neighbor clustering, and adjusts the importance of the two through weight coefficients. Clustering groups with fewer features and those far from other clusters will obtain higher anomaly scores, enabling the system to keenly capture clustering groups that may represent rare lung cancer patterns. For example, when the number of features in a certain clustering group is significantly less than the average level and it is far from other clusters, its anomaly score will be very high, indicating that the patients in this clustering group may have rare lung cancer. Doctors can conduct in-depth research on these clustering groups with high anomaly scores, carry out more detailed examinations and diagnoses, thereby increasing the early detection rate of rare lung cancer, striving for more timely treatment for patients, and improving the prognosis of patients. At the same time, this accurate anomaly screening also helps to reasonably allocate medical resources and invest more resources in patients who may have rare lung cancer.

[0087] Embodiment 5

[0088] As Figure 1 shown, specifically: The recombination analysis unit is used to obtain the anomaly score YF of the corresponding clustering group respectively according to the acquisition method of the anomaly score YF, and extract the clustering group corresponding to the maximum anomaly score YF as the target clustering group. Randomly combine the features within the target clustering group, and test the combined conditions after random combination to identify the probability of the occurrence of the corresponding rare lung cancer pattern under the corresponding conditions. Specifically:

[0089] ;

[0090] In the formula, represents the probability of suffering from the rare lung cancer pattern under the combined condition ;

[0091] represents the probability that it is the combined condition under the condition that it is known to be the rare lung cancer pattern ; This probability can be obtained through statistical analysis of existing lung cancer case data. For example, in patients with Among patients with type

[0092] Indicates a rare lung cancer pattern in historical data The proportion in the total cases;

[0093] Indicates the combined conditions The probability, for example, the probability that features x1, x2, x3, ..., xF take values V1, V2, V3, ..., VF respectively;

[0094] The information summary subsystem is used to sort out the test results according to the recombination analysis unit to generate an information evaluation reference table, which is stored in the hospital information system in the form of electronic documents (such as Excel tables, PDF reports, etc.) to assist doctors in analyzing early-stage lung cancer patients; the information evaluation reference table can be printed as needed or displayed on the system interface.

[0095] Among them, different types of lung cancer may require different treatment plans. Clarifying the probability that a patient has a certain type of lung cancer under a specific combination of characteristics helps doctors plan personalized treatment strategies in advance. For example, if the probability that a patient has a certain specific type of lung cancer is high, doctors can more specifically select a treatment method relatively suitable for this type of lung cancer among options such as surgery, radiotherapy, chemotherapy, or targeted therapy to improve the treatment effect. Moreover, the probability relationship between certain combinations of characteristics and specific types of lung cancer may also be related to the patient's response to specific treatment methods. By analyzing the probability values, doctors can, to a certain extent, predict the possible responses of patients to different treatment plans, adjust the treatment plan in advance, and avoid the harm and burden caused by ineffective treatment to patients.

[0096] In the embodiments of the present invention, the recombination analysis unit focuses on starting from the abnormal score YF. By extracting the clustering group corresponding to the maximum abnormal score as the target clustering group, it accurately locates the set of patient group characteristics most likely related to the rare lung cancer pattern. Randomly combine and test the characteristics within the target clustering group, and use the formula to calculate the probability of developing the rare lung cancer pattern under the corresponding conditions. This process can deeply explore the potential relationship between different characteristic combinations and rare lung cancer. For example, by conducting various combination tests on characteristics such as age, specific gene mutations, and tumor marker levels, it may be found that certain special characteristic combinations have a relatively high correlation probability with specific rare lung cancer patterns, providing a more targeted reference basis for clinical diagnosis. The information aggregation subsystem undertakes the test results of the recombination analysis unit, organizes them to generate an information evaluation reference table, and stores it in the hospital information system in the form of an electronic document. This enables doctors to conveniently obtain comprehensive and systematic screening information. Doctors do not need to search through a large amount of scattered data tediously. By opening the system or viewing the electronic document, they can clearly and intuitively see various test data and analysis results, such as the rare lung cancer probability under different characteristic combinations, patient clustering information, etc. The information evaluation reference table can be printed or displayed on the system interface, meeting the usage requirements of doctors in different working scenarios. When making ward rounds, doctors can print the table and analyze it in comparison with the actual situation of the patients; in the office, they can view and study various data in detail through the system interface. This efficient information presentation method greatly improves the efficiency and accuracy of doctors' analysis of early-stage lung cancer patients, helps doctors make more rapid scientific and reasonable diagnosis and treatment decisions, thereby improving the patients' medical experience and treatment effect, and also helps the hospital optimize the allocation of medical resources and improve the overall quality of medical services.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A lung cancer early screening information management system, characterized by: It includes prediction subsystem, feature substitution subsystem and information aggregation subsystem; The prediction subsystem is used to collect relevant data of early-stage lung cancer patients in the hospital using the storage devices in the hospital, generate a lung cancer database, and combine the random forest algorithm to analyze the probability that the corresponding feature values ​​belong to rare lung cancer patterns, so as to complete feature extraction and construct a feature set; The prediction subsystem includes a data collection unit, a soft classification unit, and a feature extraction unit; The soft classification unit is used to construct the probability vector P; Based on the probability vector P, the majority voting method is used to predict the prediction results corresponding to the corresponding feature values, specifically: ; In the formula, Represents the prediction result corresponding to feature x; represents a collection of rare lung cancer patterns, = ; Indicates that all decision trees give the conclusion that feature x belongs to a rare lung cancer pattern The average probability of According to the predicted results and actual results corresponding to the corresponding feature values, the features in the test set are shuffled to construct the feature importance score ZP and analyze the importance of each feature for rare lung cancer screening, specifically: ; In the formula, represents the importance score of feature x; Represents the classification error of the decision tree without disrupting the value of feature x; It represents the classification error of the decision tree when the value of feature x is shuffled; M is the optimal number of decision trees, and i is the decision tree number; The feature substitution subsystem is used to extract the target early-stage lung cancer patients in the hospital, and extract the feature set of the target early-stage lung cancer patients. After clustering similarity analysis, the abnormal score YF is traversed to obtain the abnormal score YF. Based on the abnormal score YF, the combined condition test operation is performed to identify the probability of the corresponding rare lung cancer pattern under the corresponding conditions. The information summary subsystem is used to summarize the information after the test to generate an information evaluation form and upload it to the hospital information system.

2. The early screening information management system for lung cancer according to claim 1, characterized in that: The data collection unit is used to collect the basic information, diagnosis information, examination information and medication records of each early lung cancer patient according to the hospital information system within the hospital, and share and interact with the image storage and transmission system, laboratory information system and electronic medical record system according to the calling function of the hospital information system, so as to fill in the medical images, pathology and laboratory data and electronic medical records of each early lung cancer patient, and generate a lung cancer database after statistics.

3. The early screening information management system for lung cancer according to claim 2, characterized in that: The soft classification unit is used to construct a random forest model based on the random forest algorithm, divide the data in the lung cancer database into a training set, a test set, and a validation set, and pre-set the initial range of decision trees. The training set is used to train the random forest model, and then the validation set is used to perform an accuracy A calculation on the random forest model to evaluate the performance of the random forest model under different decision tree quantities. By comparing the accuracy A values ​​under different decision tree quantities, the number of decision trees corresponding to the maximum accuracy A value is selected as the optimal number of decision trees M in the random forest. According to the optimal number of decision trees, each feature in the lung cancer database is input into the root node of each decision tree, and the direction of each feature is determined through the feature judgment condition of each internal node of the decision tree, and the leaf node of the corresponding decision tree is reached after traversal. After statistics, the number of features N of the training set contained in the leaf node of the corresponding decision tree is obtained, and according to the number of features N of the training set contained in the leaf node of the corresponding decision tree, the number of features n of the training set contained in the leaf node of the corresponding decision tree belonging to each rare lung cancer pattern is determined. According to the number of features N of the training set contained in the leaf node of the corresponding decision tree and the number of features n of the training set contained in the leaf node of the corresponding decision tree belonging to each rare lung cancer pattern, the probability of the corresponding feature value belonging to the rare lung cancer pattern is analyzed, and the probability vector P is constructed, which is specifically: ; ; In the formula, Represents a decision tree The feature x in the lung cancer database belongs to a rare lung cancer pattern The probability estimate of ; Represented in the decision tree The leaf nodes of The number of features; Represents a decision tree The probability vector of feature x in the lung cancer database belonging to different rare lung cancer patterns; , and Represents decision tree The feature x in the lung cancer database belongs to a rare lung cancer pattern Probability estimate, decision tree The feature x in the lung cancer database belongs to a rare lung cancer pattern Probability estimate and decision tree The feature x in the lung cancer database belongs to a rare lung cancer pattern The probability estimate of ; m represents the number of types of rare lung cancer patterns; i represents the decision tree number; j represents the rare lung cancer pattern number.

4. The early screening information management system for lung cancer according to claim 3 is characterized by: The feature extraction unit is used to set an importance threshold according to the feature importance score ZP of different features, and compare the feature importance score ZP of the corresponding feature with the importance threshold. If the feature importance score ZP of the corresponding feature exceeds the importance threshold, the corresponding feature is included in the feature set; if the feature importance score ZP of the corresponding feature does not exceed the importance threshold, the corresponding feature is not included in the feature set.

5. The early screening information management system for lung cancer according to claim 4, characterized in that: The feature substitution subsystem includes a clustering unit, an anomaly analysis unit, and a recombination analysis unit; The clustering unit is used to extract the early-stage lung cancer patients corresponding to the relevant features in the feature set in the hospital according to the feature set in the feature extraction unit, and mark the extracted early-stage lung cancer patients as target early-stage lung cancer patients. Based on the target early-stage lung cancer patients, the relevant data of the target early-stage lung cancer patients are extracted from the lung cancer database. According to the DBSCAN algorithm, the relevant data of the target early-stage lung cancer patients are clustered to obtain several cluster groups. According to the corresponding cluster groups, clustering information is generated, and the clustering information includes the features in cluster group z and the number of features in cluster group z. , the average number of features across all clusters , the average distance from cluster group z to its nearest neighbor cluster and the average distance of all clusters to the nearest neighbor cluster .

6. The early screening information management system for lung cancer according to claim 5, characterized in that: The anomaly analysis unit is used to traverse all cluster groups according to the cluster information to calculate the anomaly score YF. The anomaly score YF is obtained by the following formula: ; In the formula, represents the anomaly score of cluster group z; represents the average number of features across all clusters; represents the number of features in cluster group z; Represents the average distance from cluster group z to its nearest neighbor cluster; Represents the average distance of all clusters to the nearest neighbor cluster; and Represents the weight coefficient, which is used to adjust the importance of feature quantity and distance factors in anomaly scoring.

7. The early screening information management system for lung cancer according to claim 6, characterized in that: The recombination analysis unit is used to obtain the abnormality scores YF of the corresponding cluster groups according to the acquisition method of the abnormality scores YF, and extract the cluster group corresponding to the maximum abnormality score YF as the target cluster group, randomly combine the features in the target cluster group, and test the combined conditions after random combination to identify the probability of the corresponding rare lung cancer pattern under the corresponding conditions, specifically: ; In the formula, Indicates that in the combination condition Cases, rare lung cancer pattern probability; In a known rare lung cancer pattern In the case of probability; Indicates rare lung cancer patterns in historical data The proportion of total cases; Indicates combination conditions probability.

8. The early screening information management system for lung cancer according to claim 7, characterized in that: The information summary subsystem is used to organize the test results according to the recombinant analysis units to generate information evaluation aid tables, which are stored in the hospital information system in the form of electronic documents to assist doctors in analyzing early-stage lung cancer patients.

Citation Information

Patent Citations

  • Endometrial tumor classification marking method based on random forest

    CN111860576A