Method for establishing pulmonary nodule identification feature model and feature model thereof
By establishing a lung nodule identification feature model based on dielectric characteristic data, the problem of difficulty in real-time identification of lung malignant tumor tissue in the prior art is solved, and rapid and accurate identification of benign and malignant during surgery is achieved, and the accuracy of the surgery is improved.
Patent Information
- Application Number
- CN202510183758.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has not yet effectively detected and identified the dielectric properties of lung malignant tumor tissue, making it difficult to achieve real-time benign and malignant identification detection in lung surgery.
By establishing a lung nodule identification feature model, using dielectric characteristic data such as dielectric constant and conductivity, combined with gradient enhancement decision tree method, a characteristic model that can quickly identify benign and malignant lung nodules.
It realizes rapid and efficient identification of benign and malignant lung nodules in emergency application scenarios, such as surgery, reduces the risk of relying on intraoperative frozen pathological diagnosis, and improves the accuracy and safety of the surgery.
Smart Images

Figure CN120126795A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for establishing a lung nodule discrimination feature model and the feature model, belonging to the technical field of detection tools and research. Background Art
[0002] In the prior art, the research on the dielectric properties of human tissues in the radio frequency or microwave range has exceeded 100 years. The frequency spectrum of the dielectric properties of healthy tissues has been systematically studied, and a database of the dielectric properties of healthy human tissues has been constructed and publicly available worldwide, becoming a milestone achievement in the field of biomedical engineering. Basic research has long confirmed that when the physiological or pathological state of the basic structural unit cells of human tissues changes, and the microenvironment in which the cells are located changes, the dielectric properties of the tissues will also change. After tissue canceration, the dielectric property changes are relatively large, and some differences even reach 10 times. If the dielectric properties of living tissues can be detected or imaged, it can provide valuable information for tumor diagnosis. In particular, dielectric property imaging may be used for early cancer diagnosis and may even be used to track the entire change process of normal tissues evolving into cancer tissues, which may have groundbreaking value for cancer research. Previous studies have found that the dielectric properties of human gliomas and breast cancer tissues have changed significantly compared with normal tissues, but no relevant studies on detecting the dielectric properties of lung malignant tumor tissues have been seen.
[0003] The existing clinical test results confirm that the dielectric property differences of cancerous tissues in tumor patients are generally more than 30% or even several times. Using this functional prototype can fully realize the intraoperative real-time in-vivo discrimination detection of tissue benignity and malignancy, which is of great significance for guiding general surgery on tumors. And this new technology for real-time judging the benignity and malignancy of tumors is of even greater significance in the field of thoracic surgery, especially in the field of lung surgery.
[0004] In addition, the differences in the dielectric properties of different lung cancer tissues and their correlations with CT omics and genomics have not been deeply explored. Through the analysis of the three combined omics, it is expected to further explain and analyze the molecular mechanisms affecting the dielectric properties of lung cancer tissues, providing a research basis for the subsequent development of new diagnostic tools through dielectric omics. Summary of the Invention
[0005] In order to overcome the deficiencies of the prior art, the first object of the present invention is to provide a method for establishing a lung nodule discrimination feature model. The feature model of the present invention can analyze and quickly discriminate the benignity and malignancy of lung nodules, which is beneficial to obtaining accurate analysis results in emergency application scenarios, such as during surgery, and helps the precise progress of the surgery.
[0006] The second object of the present invention is to provide a feature model for the above-mentioned lung nodule discrimination.
[0007] The first object of the present invention can be achieved by adopting the following technical solutions:
[0008] A method for establishing a feature model for lung nodule identification, comprising:
[0009] Model output and prediction probability conversion step: is the prediction probability of the i-th sample; F(x i ) is the total score of the i-th sample after T rounds of iteration, that is, the total score of all T decision trees; F(x i ) is converted into a prediction probability through the Sigmoid function Objective function step:
[0010] The model objective function consists of a loss function and a regularization term: Taylor expansion approximation step of gradient boosting:
[0011] In the t-th round of iteration, when generating the t-th tree, the objective function is approximately expressed by the second-order Taylor expansion as: First-order gradient, the derivative of the loss with respect to the current prediction: Second-order gradient, the second derivative of the loss with respect to the current prediction: When When, Take 1; when When, take 0; y i ∈{0, 1} true label;
[0012] Tree construction and splitting gain step:
[0013] f t (x i ) is defined as where q(x i ) represents the leaf node where the sample i is located, and w represents the weight of the leaf node, that is, the predicted value of the decision tree. After replacing f t (x i ) and Ω(f t ), the objective function becomes:
[0014]
[0015] where: I j ={i|q(x i )=j}, represents the samples of the decision tree q located in the j-th leaf node. Let
[0016] At this time, the task is to minimize the objective function, take the derivative of the above formula and set it to 0;
[0017] That is, the value w of the leaf node j The optimal value is The optimal solution of the objective function is Also known as the structure score, similar to the information gain, it can score the structure of the tree; for each tree, the splitting point is selected by the greedy algorithm, and the benefit of the objective function after the next split is:
[0018]
[0019] To make the next objective function smaller, the splitting gain Gain is defined as: The sum of the first-order gradients of the left and right child nodes after splitting; The sum of the second-order gradients of the left and right child nodes after splitting; γ: the minimum gain threshold required for splitting, that is, the regularization parameter; judge the information gain after splitting the node, that is, whether the sum of the first part + the second part is greater than the unsplit situation, and considering the problem of whether the model will be too complex, if the increased score is less than the regularization term, the node will no longer split;
[0020] Derive the feature model:
[0021] The model is obtained by training on the training set. For each independent sample x i The best sub-leaf splitting principle for features. For a brand-new sample x different from the training set i+1 , through the feature judgment process of the model, the sample falls into a specific sub-leaf, and there is a corresponding node value w on the sub-leaf t,j , and the final score is obtained through the weighted sum of T rounds of iterations Substitute into Get That is, the predicted score of the new sample. Finally, according to the binary classification principle: when , Take 1; when Take 0 for classification.
[0022] Furthermore, the data input by this establishment method is the dielectric constant and the conductivity.
[0023] Furthermore, in the model output and probability conversion step, the probability that the sample belongs to the positive class and the class value is 1.
[0024] Furthermore, in the objective function step, the regularization term Ω: controls the model complexity,
[0025] Furthermore, after the tree construction and splitting gain steps, it also includes a model update rule step:
[0026] After each round of iteration, the prediction result of the model is gradually updated through the learning rate η (eta): Let \(T\) be the round number, and \(J\) t be the total number of leaf nodes in the \(t\)-th round; \(w\) t,j be the value of the \(j\)-th leaf node in the \(t\)-th round; \(I(x\in R\) t,j ) be all the samples that fall within the region of the \(j\)-th leaf node in the \(t\)-th round; \(j = 1, 2, 3,\cdots\).
[0027] Furthermore, in the tree construction and splitting gain steps, \(i\) is the \(i\)-th sample; \(x\) i be the \(i\)-th input sample, which contains all the features of the sample, \(x\) i = [x i1 , x i2 , \cdots, x ij , and \(j\) is the total number of features of the \(i\)-th sample. \(f\) k (x i ) be the prediction score of the \(k\)-th decision tree for the \(i\)-th sample, where \(K\) is the total number of trees; \(\eta\in(0, 1)\) is the learning rate, which restricts the contribution of each tree to the overall and prevents overfitting.
[0028] The second object of the present invention can be achieved by adopting the following technical solutions:
[0029] A feature model obtained by the method for establishing a feature model for lung nodule discrimination as described in claim 1, wherein the feature model is: Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] During the operation, surgeons usually need to decide the surgical resection range according to the intraoperative frozen pathological results. However, intraoperative frozen pathological diagnosis relatively depends on the experience of pathologists, is time-consuming and has a certain misdiagnosis probability. While this method is fast, efficient, objective and low-cost; through the feature model of the present invention, the benign and malignant nature of lung nodules can be quickly and accurately identified, which is beneficial to obtaining accurate analysis results in urgent application scenarios, such as during the operation, and helps the precise progress of the operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a diagram showing the comparison of the dielectric constants of benign and malignant lung nodules;
[0032] Figure 2 It is a diagram showing the conductivity of benign and malignant lung nodules;
[0033] Figure 3 It is the ROC curve of the training set;
[0034] Figure 4 It is the ROC curve of the test set. DETAILED DESCRIPTION OF THE INVENTION
[0035] Next, in combination with the accompanying drawings and specific embodiments, the present invention will be further described:
[0036] Method for establishing a characteristic model for lung nodule identification:
[0037] 1. Dielectric property detection
[0038] In the operating room, while performing wedge resection of lung tumors / nodules and radical resection of lung cancer, wait to obtain the specimen. After the specimen is removed from the body, the surgeon dissects the specimen, locates the exact tumor or nodule and marks it, avoiding tumor necrosis and bleeding sites, and guides the placement of the probe by the detection personnel.
[0039] 2. Obtaining the benign and malignant discrimination results
[0040] Compare the dielectric property-related values obtained from the specimen detection with the optimal threshold to obtain the benign and malignant discrimination results.
[0041] In the previous detection of the dielectric properties of lung nodules and the comparison and analysis with the gold standard of routine paraffin pathological diagnosis, an ROC curve was drawn, and the optimal diagnostic threshold has been determined:
[0042] Take α as 0.05, β as 0.8, and the pre-estimated diagnostic test sensitivity Se as 0.8, and conduct a retrospective diagnostic test. Include patients with appropriate thoracic surgical lung nodules, compare and analyze the dielectric property measurement values (dielectric constant and conductivity) with the results of routine pathological examinations after the gold standard surgery, clarify the diagnostic threshold and the corresponding specificity and sensitivity, and make an ROC curve.
[0043] Among them, the variance function
[0044] 3. Input the dielectric property-related values (dielectric constant and conductivity) obtained from the specimen detection to establish a characteristic model; including the following steps:
[0045] 1) Model output and prediction probability conversion:
[0046] The model predicts the probability that the sample belongs to the positive class (class 1). is the prediction probability of the i-th sample; F(x i ) is the total score of the i-th sample after T rounds of iteration, that is, the total score of all T decision trees; F(x i ) is converted into a prediction probability through the Sigmoid function
[0047] 2) Objective function:
[0048] The model objective function consists of a loss function and a regularization term: The regularization term Ω: controls the model complexity,
[0049] 3) Taylor expansion approximation of gradient boosting
[0050] In the t-th iteration, when generating the t-th tree, the objective function is approximated by the second-order Taylor expansion as follows: First-order gradient, the derivative of the loss with respect to the current prediction: Second-order gradient, the second derivative of the loss with respect to the current prediction: When , Take 1; when , take 0; y i ∈ {0, 1} true label.
[0051] 4) Tree construction and split gain
[0052] f t (x i ) is defined as where q(x i ) represents the leaf node where sample i is located, and w represents the weight of the leaf node, that is, the predicted value of the decision tree. After replacing f t (x i ) and Ω(f t ), the objective function becomes:
[0053]
[0054] where: I j ={i|q(x i ) = j}, represents the samples in the j-th leaf node of the decision tree q. Let
[0055] At this time, the task is to minimize the objective function, take the derivative of the above formula and set it to 0;
[0056] That is, the optimal value of the leaf node value w j is The optimal solution of the objective function is Also known as the structure score, similar to the information gain, it can score the structure of the tree; each tree selects the split point through the greedy algorithm, and the benefit of the objective function after the next split is:
[0057]
[0058] To make the objective function smaller next time, the split gain Gain is defined as the maximization of: The sum of the first-order gradients of the left and right child nodes after splitting; The sum of the second-order gradients of the left and right child nodes after splitting; γ: the minimum gain threshold required for splitting, i.e., the regularization parameter; judge the information gain after splitting the node, that is, whether the sum of the first part and the second part is greater than the case without splitting, and considering the problem of whether the model will be too complex, if the increased score is less than the regularization term, the node will no longer split; i is the i-th sample; x i is the i-th input sample, which contains all the features of the sample, x i = [x i1 , x i2 , …, x ij , j is the total number of features of the i-th sample. f k (x i ) is the prediction score of the k-th decision tree for the i-th sample, and K is the total number of trees; η ∈ (0, 1) is the learning rate, which restricts the contribution of each tree to the overall situation and prevents overfitting;
[0059] Derive the feature model:
[0060] The model is trained through the training set to obtain the best sub-leaf splitting principle for each independent sample x i features. For a brand-new sample x i+1 different from the training set, the model passes through the feature judgment process, and the sample falls into a specific sub-leaf, and there is a corresponding node value w t,j on the sub-leaf. The final score is obtained through the weighted sum of T rounds of iterations Substitute into to get That is, the prediction score of the new sample. Finally, according to the binary classification principle: when , take 1; when , take 0 for classification; 5) Model update rule
[0061] After each round of iteration, the prediction result of the model is gradually updated through the learning rate η (eta): T is the number of rounds, J t is the total number of leaf nodes in the t-th round; w t,j is the value of the j-th leaf node in the t-th round; I(x ∈ R t,j ) is all the samples that fall into the area of the j-th leaf node in the t-th round; j = 1, 2, 3….
[0062] Use the previously self-developed tissue dielectric property detector to detect the dielectric property values (dielectric constant and conductivity) of specimens during lung nodule surgery, including lung cancer (squamous cell carcinoma, adenocarcinoma, small cell lung cancer, etc.), lung benign tumors or inflammatory tissues, and compare them with the postoperative routine pathology to clarify its practical value for rapid intraoperative benign and malignant diagnosis.
[0063] Example 1:
[0064] In the early stage of this project, a functional prototype was successfully developed, and 190 patients were successfully tested with the functional prototype, including 46 cases of benign pulmonary nodules and 144 cases of malignant ones. The dielectric properties (dielectric constant and conductivity) were statistically analyzed, and it was found that the differences in dielectric constant and conductivity between the benign group and the malignant group were statistically significant (p values were all less than 0.05). Figure 1 For the comparison of the dielectric constant of benign and malignant pulmonary nodules, Figure 2 For the comparison of the conductivity of benign and malignant pulmonary nodules. The existing clinical test results have confirmed that the differences in dielectric properties of cancerous tissues in tumor patients are generally more than 30% or even several times. Using this functional prototype can fully realize the real-time in-vivo detection of the benign and malignant nature of tissues during surgery, which is of great significance in the field of thoracic surgery, especially in the field of pulmonary surgery.
[0065] Example 2:
[0066] 114 patients were successfully tested with the functional prototype, including 18 cases of benign pulmonary nodules and 96 cases of malignant ones. The dielectric properties (dielectric constant and conductivity) were statistically analyzed, and the data was segmented in an 8:2 ratio to obtain a training set and a test set of 92 cases and 22 cases respectively. Among them, the training set had 16 cases of benign nodules and 76 cases of malignant nodules. The training set was 2 cases of benign nodules and 76 cases of 20 cases of malignant nodules. The learning rate η (eta) was set to 0.01, max_depth was set to 2, both γ and λ were set to 1, and the number of iterations was 180 times. After training the model with the training set, the test set was tested. Table 1 and Table 2 are the tables of the training set and the test set respectively, Figures 3 - 4 They are the ROC curves of the training set and the test set respectively.
[0067] Table 1 Training Set
[0068]
[0069] Accuracy: 0.95
[0070] Table 2 Test Set
[0071]
[0072] Accuracy: 0.8636
[0073] For those skilled in the art, various corresponding changes and deformations can be made according to the technical solutions and concepts described above, and all these changes and deformations should fall within the protection scope of the claims of the present invention.
Claims
1. A method for establishing a characteristic model for identifying pulmonary nodules, characterized in that include: Steps for converting model output to predicted probability: is the predicted probability of the i-th sample; F(x i ) is the total score of the ith sample after T rounds of iterations, that is, the total score of all T decision trees; F(x i ) is converted into predicted probability through Sigmoid function Objective function steps: The model objective function consists of a loss function and a regularization term: Taylor expansion approximation steps for gradient boosting: In the tth iteration, when the tth tree is generated, the objective function is approximated by the second-order Taylor expansion: First-order gradient, the derivative of the loss with respect to the current prediction: Second-order gradient, the second-order derivative of the loss with respect to the current prediction: when hour, Take 1; when When y i ∈{0, 1} true label; Tree construction and split gain steps: f t (x i ) is defined as where q(x i ) represents the leaf node where sample i is located, w represents the weight of the leaf node, that is, the predicted value of the decision tree, replacing f t (x i ) and Ω(f t ) Then the objective function becomes: Where: I j ={i|q(x i )=j}, represents the sample of the decision tree q located at the jth leaf node, let At this time, the task is to minimize the objective function, and take the derivative of the above formula and set it to 0; That is, the value w of the leaf node j The optimal value is The optimal solution of the objective function is Also known as structural score, it is similar to information gain and can score the structure of the tree. Each tree selects a split point through a greedy algorithm, and the objective function gain after the next split is: In order to make the next objective function smaller, the maximum split gain Gain is defined as: The sum of the first-order gradients of the left and right child nodes after the split; The sum of the second-order gradients of the left and right child nodes after the split; γ: the minimum gain threshold required for splitting, that is, the regularization parameter; judge whether the information gain after splitting the node, that is, the first part + the second part, is greater than the case without splitting, and consider whether the model is too complicated. If the increased score is less than the regularization term, the node will not split; The characteristic model is obtained: The model is trained through the training set to obtain for each independent sample x i The optimal cotyledon splitting principle for features. For a new sample x different from the training set i+1 , the model passes through the feature judgment process, and the sample falls into a specific cotyledon, which has a corresponding node value w t,j , the final score is obtained by T rounds of iterative large weighted sum Bring in have to That is, the prediction score of the new sample, and finally according to the binary classification principle: when hour, Take 1; when When , take 0 and classify.
2. The method for establishing a pulmonary nodule identification feature model according to claim 1, characterized in that: The input data for this setup method are dielectric constant and conductivity.
3. The method for establishing a pulmonary nodule identification feature model according to claim 1, characterized in that: In the model output and probability conversion step, the sample belongs to the positive class and the probability of the class value is 1.
4. The method for establishing a characteristic model for identifying pulmonary nodules according to claim 1, characterized in that: In the objective function step, the regularization term Ω: controls the model complexity, 5. The method for establishing a characteristic model for identifying pulmonary nodules according to claim 1, characterized in that: After the tree construction and split gain step, it also includes a model update rule step: After each iteration, the model's prediction results are gradually updated by the learning rate η(eta): T is the round, J is t is the total number of leaf nodes in round t; w t,j is the value of the jth leaf node in round t; I(x∈R t,j ) are all samples that fall into the region of the jth leaf node in the tth round; j = 1, 2, 3… 6. The method for establishing a characteristic model for identifying pulmonary nodules according to claim 1, characterized in that: In the tree construction and split gain steps, i is the i-th sample; x i is the i-th input sample, which contains all the features of the sample, x i =[x i1 , x i2 , …, x ij ], j is the total number of sample features of the i-th sample. k (x i ) is the prediction score of the kth decision tree for the ith sample, K is the total number of trees; η∈(0,1) is the learning rate, which limits the contribution of each tree to the overall and prevents overfitting.
7. A feature model obtained by the method for establishing a feature model for distinguishing pulmonary nodules according to claim 1, characterized in that: The feature model is: