Lung nodule benign and malignant disease risk prediction system, electronic device and application

CN122658412APending Publication Date: 2026-08-28HUNAN HUISEN BIOTECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611152543.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0006]为此,本发明提供肺结节良恶性疾病风险预测系统、电子设备及应用,解决现有技术中TCR信号独立性不足、特征表征层次不完整、高维稀疏数据筛选不稳定以及单一模型泛化能力有限的问题

Benefits of technology

[0022] This invention employs an RF-LASSO-RFE stable screening strategy, ensuring that features entering the final model are simultaneously constrained by the nonlinear splitting contribution of random forest, the linear sparse discriminant contribution of LASSO, the intersection consensus constraint, and the stability of cross-validation recursive elimination. After obtaining the optimal TCR feature subset, multiple risk probability perspectives are formed using artificial neural networks, radial basis support vector machines, gradient boosting trees, and random forests. Then, a generalized linear model completes probability calibration and fusion weight learning, thereby reducing the impact of local biases of a single model on the final result.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122658412A_ABST
    Figure CN122658412A_ABST
Patent Text Reader

Abstract

The application discloses a lung nodule benign and malignant disease risk prediction system, an electronic device and application, takes TCR immune repertoire sequencing data of an individual peripheral blood ex vivo sample as the only input, forms an end-to-end prediction system from original TCR clone type data to lung nodule malignant risk probability output through sequencing data quality control, six-dimensional immune characterization, derived feature construction, RF-LASSO-RFE stable screening and two-level heterogeneous stacking probability fusion, wherein the six-dimensional immune characterization includes V / J gene use preference, clone amplification and length distribution, CDR3 core amino acid motif, sequence convergence, tumor specificity enrichment score and ecological diversity. The application can highlight the independent prediction value of the TCR molecular signal itself, improve the interpretability, robustness and generalization ability of small sample high-dimensional immune repertoire data, and can be used for non-invasive auxiliary evaluation, follow-up monitoring and lung cancer early screening of lung nodule benign and malignant diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of risk prediction and immune repertoire data analysis technology, specifically relating to a risk prediction system, electronic equipment and applications for benign and malignant pulmonary nodules. Background Technology

[0002] Early-stage lung cancer often presents primarily as pulmonary nodules on imaging. Accurate differentiation between benign and malignant pulmonary nodules remains a challenge in clinical diagnosis and treatment. Current diagnostic methods have significant limitations: CT imaging has a high false-positive rate, easily leading to excessive biopsies and surgeries, thus wasting medical resources; percutaneous lung biopsy is an invasive procedure with risks of complications such as pneumothorax and bleeding, and has a low success rate for small nodules; serum tumor markers such as CEA and CYFRA21-1 have insufficient sensitivity, failing to meet the needs of early lung cancer screening. Therefore, developing a non-invasive, accurate, and convenient technique for differentiating between benign and malignant pulmonary nodules is of significant clinical and social value for achieving early diagnosis and treatment of lung cancer and reducing lung cancer mortality.

[0003] T-cell receptors (TCRs) are core molecules for T-cell specific recognition of antigens. During tumorigenesis and development, tumor-specific antigens drive T-cell clonal expansion and convergent evolution, leading to characteristic molecular changes in the peripheral blood TCR immune repertoire. In recent years, the rapid development of high-throughput sequencing technology has made comprehensive analysis of peripheral blood TCR immune repertoires possible, providing a new technological direction for non-invasive early cancer screening. Existing studies have shown that the diversity of TCR immune repertoires, V / J gene usage preferences, and clonal expansion patterns are closely related to the occurrence and development of lung cancer, and possess the potential to serve as biomarkers for differentiating benign and malignant lung nodules.

[0004] However, existing TCR immune repertoire-based technologies for assessing the benign and malignant nature of pulmonary nodules still suffer from numerous technical shortcomings, severely hindering their clinical translation and application: First, some schemes input TCR immune repertoire features along with age, smoking history, nodule size, radiomics, or pathological information into the model, making the model performance susceptible to the influence of multi-source information and making it difficult to prove the independent predictive value of the TCR molecular signal itself; Second, some schemes focus on specific CDR3 database matching, single diversity indices, or preferences for a few V / J genes, failing to cover the genetic composition, clonal dynamics, core regions of antigen recognition, sequence convergence evolution, tumor-related enrichment, and changes in ecological diversity involved in tumor-induced immune remodeling; Third, the TCR immune repertoire feature matrix is ​​characterized by high dimensionality, sparsity, strong noise, and multicollinearity, making it difficult for single linear or nonlinear screening methods to simultaneously consider discriminative contribution and stability; Fourth, single classification models are prone to local bias under conditions of small samples and class imbalance, making it difficult to generate stable, interpretable, and generalizable continuous risk probability outputs.

[0005] Therefore, there is an urgent need to develop a technical solution that uses pure TCR immune repertoire data as the core and can complete the mapping of malignant risk of pulmonary nodules without relying on clinical / imaging variables. After standardized quality control, multidimensional immune characterization, derived feature organization, stable feature compression and heterogeneous probability fusion, the original TCR clonal data can be transformed into a non-invasive auxiliary assessment tool with interpretability and generalization stability. Summary of the Invention

[0006] To address these issues, this invention provides a risk prediction system, electronic device, and application for benign and malignant pulmonary nodules, solving the problems of insufficient TCR signal independence, incomplete feature representation hierarchy, unstable screening of high-dimensional sparse data, and limited generalization ability of a single model in the prior art.

[0007] To achieve the above objectives, in a first aspect, the present invention provides the following technical solution: a risk prediction system for benign and malignant pulmonary nodules, which is prepared by combining reagent materials and / or instruments using biomarker information from an individual ex vivo sample TCR immune repertoire. The prediction system uses peripheral blood TCR immune repertoire sequencing data as the sole input, without inputting clinical baseline information, CT imaging information, pathological information, or serum tumor marker information. The TCR immune repertoire biomarker information includes: V / J gene usage preference, clonal amplification and length distribution, CDR3 core amino acid motif composition, TCR sequence convergence, tumor-specific enrichment score, and ecological diversity indicators.

[0008] As a preferred application, the V / J gene usage preference is: the weighted relative frequencies of TRBV and TRBJ gene fragments in the TCR immune repertoire after logarithmic transformation of TCR clonoid counts; The frequency and length distribution of TCR clones include: the distribution frequency of TCR clones classified into five levels according to the clone proportion: extremely amplified, large, medium, small and rare; and the cumulative distribution frequency of CDR3 amino acid sequences divided into five intervals according to length.

[0009] As a preferred embodiment of the application, the CDR3 core amino acid motif is composed of: After quality control screening, the TCR sequences were subjected to the removal of 3 conserved amino acids at each end, and the weighted frequencies of 20 standard amino acids in the CDR3 core recognition region were determined. The quality control screening criteria were: sequence length 10-24, count ≥10, starting with C and ending with F, and no stop codon. The TCR sequence convergence is defined as the percentage of clones with different nucleotide sequences but translated into the same combination of "V gene + CDR3 amino acids".

[0010] As a preferred application, the tumor-specific enrichment score It is calculated using the following formula: ;

[0011] In the formula, This represents the total number of valid TCR sequences after sample quality control. Needleman-Wunsch local alignment similarity between the sequence to be tested and sequences in lung cancer-related TCR databases; This is the clone count of the sequence to be tested in the sample. The clone count of the matching sequence in the database is used; the criteria for screening and scoring sequences are: normalized edit distance ≤ 0.3 and similarity score ≥ 4.5; The ecological diversity index is: in a unified The Shannon entropy, Simpson index, and Chao1 index were calculated after resampling at the sequencing depth.

[0012] Secondly, a risk prediction system for benign and malignant pulmonary nodules is provided, including a data acquisition unit and a data analysis unit; The data acquisition unit is used to acquire TCR immune repertoire biomarker information from individual ex vivo samples; the TCR immune repertoire biomarker information includes: V / J gene usage preference, clonal frequency and length distribution, CDR3 region amino acid motif composition, TCR sequence convergence, tumor-specific enrichment score, and ecological diversity indicators. The data analysis unit is used to sequentially perform sequencing data quality control, six-dimensional immune characterization, derived feature construction, RF-LASSO-RFE stable feature screening, and two-level heterogeneous stacking probability fusion on the information taken by the data acquisition unit to obtain a continuous malignancy risk score for pulmonary nodules.

[0013] As the preferred option for the prediction system, the risk score of benign or malignant lung nodules obtained by the data analysis unit conforms to the value obtained according to the following model formula: ;

[0014] In the formula, This indicates that the lung nodule is malignant. For the intercept term, , , , These are the decision weight coefficients assigned to the meta-learner for artificial neural networks, radial basis support vector machines, gradient boosting trees, and random forests, respectively. , , , These represent the probability of a malign prediction output by the corresponding base learner.

[0015] As the preferred option for predicting the risk of benign and malignant lung nodules, the numerical value obtained according to the model formula is the probability score of malignant lung nodules. The score reflects the level of individual lung nodule malignancy risk; individuals with a lung nodule malignancy risk score ≥0.5 are judged to have a high risk of lung nodule malignancy.

[0016] Thirdly, an electronic device is provided, including a first memory, a first processor, and a computer program stored in the first memory and executable on the first processor, wherein the first processor executes the program to implement a scoring process including the following steps: The raw sequencing data of TCR immune repertoire from individual ex vivo samples were obtained, and TCR immune repertoire biomarker information with defined dimensions was extracted without introducing any non-TCR variables. The defined dimensions include V / J gene usage preference, clonal amplification and length distribution, CDR3 core amino acid motif composition, TCR sequence convergence, tumor-specific enrichment score, and ecological diversity indicators. Based on the obtained TCR immune repertoire biomarker information, proportional, offset, density, correction and co-variation class derived features were constructed, and the optimal TCR feature subset was obtained through the RF-LASSO-RFE stable screening process of random forest nonlinear initial screening, LASSO linear sparse initial screening, intersection consensus feature pool and recursive feature elimination. Based on the optimal TCR feature subset, a pre-trained two-level heterogeneous stacking fusion model is used to perform basic learner probability output, meta-learner probability calibration, and fusion weight learning to calculate the malignancy risk score of individual lung nodules.

[0017] As a preferred embodiment of the electronic device, when the first processor executes the program, the process of calculating the benign or malignant risk score of an individual's lung nodules includes: The optimal feature subset is input into the pre-trained base learner layer to obtain the malign prediction probabilities output by the artificial neural network, radial basis support vector machine, gradient boosting tree, and random forest, respectively. The predicted probabilities output by the base learner are input into the generalized linear model of the meta-learner layer to calculate the final risk score for malignancy and the risk score for benign or malignant lung nodules. These scores conform to the values ​​obtained according to the following model formula: ; In the formula, This indicates that the lung nodule is malignant. For the intercept term, These are the decision weight coefficients assigned to the meta-learner for artificial neural networks, radial basis support vector machines, gradient boosting trees, and random forests, respectively. These represent the probability of a malign prediction output by the corresponding base learner; The numerical value obtained according to the model formula is the probability score of malignant lesions of pulmonary nodules. The score reflects the individual's risk of malignant pulmonary nodules. Individuals with a malignant pulmonary nodule risk score ≥0.5 are judged to have a high risk of malignant pulmonary nodules.

[0018] As a preferred embodiment of the electronic device, when the first processor executes the program, it provides individualized health guidance and clinical reference suggestions based on the individual's risk score for benign or malignant pulmonary nodules.

[0019] As a preferred electronic device, it is used to assess the risk of benign or malignant pulmonary nodules.

[0020] The core inventiveness of this invention lies first in establishing a pure TCR independent prediction system. During model training, feature selection, and actual prediction, non-TCR variables such as age, gender, smoking history, nodule size, nodule density, ground-glass opacity morphology, spiculation, CT radiomics features, pathological classification, or serum tumor markers are not input. This transforms the model's learning objective from multimodal clinical auxiliary judgment to an independent mapping from peripheral blood TCR immune molecular signals to the malignant risk of lung nodules.

[0021] This invention further characterizes the TCR immune repertoire into a six-dimensional primary feature layer and a derived feature layer. The six-dimensional primary features collectively cover the genetic composition, clonal dynamics, antigen recognition core region, sequence convergence evolution, tumor-associated clonal enrichment, and overall ecological diversity of the TCR repertoire; the derived feature layer calculates proportion, offset, density, correction, and co-variation class variables within the pure TCR data, transforming local sequence signals, overall ecological signals, and clonal dynamics signals into a unified composite feature space.

[0022] This invention employs an RF-LASSO-RFE stable screening strategy, ensuring that features entering the final model are simultaneously constrained by the nonlinear splitting contribution of random forest, the linear sparse discriminant contribution of LASSO, the intersection consensus constraint, and the stability of cross-validation recursive elimination. After obtaining the optimal TCR feature subset, multiple risk probability perspectives are formed using artificial neural networks, radial basis support vector machines, gradient boosting trees, and random forests. Then, a generalized linear model completes probability calibration and fusion weight learning, thereby reducing the impact of local biases of a single model on the final result.

[0023] This invention can generate a continuous malignancy risk score for pulmonary nodules without relying on clinical baseline, imaging or pathological priors, and can more directly reflect the molecular changes in the TCR immune repertoire induced by tumor antigen stimulation. This low-dependency, non-invasive and standardized data processing workflow is suitable for primary screening, follow-up monitoring, people with unclear imaging conclusions, and patients with pulmonary nodules who cannot tolerate invasive puncture for auxiliary assessment.

[0024] The technical solution of this invention can be directly transformed into various product forms such as TCR immune repertoire sequencing detection kits, supporting data analysis software, clinical auxiliary decision-making systems, and medical testing workstations; all steps are in vitro data processing procedures, with a short industrialization cycle and low implementation difficulty, enabling rapid commercial application and generating significant economic and social benefits. Attached Figure Description

[0025] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0026] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.

[0027] Figure 1 This is a comparison of the multi-model receiver operating characteristic (ROC) curves and diagnostic efficacy based on TCR immune repertoire features provided in this embodiment of the invention.

[0028] Figure 2 The distribution chart shows the predicted contribution (importance) of the top 20 core TCR immune repertoire features automatically selected by the model in this embodiment of the invention. Detailed Implementation

[0029] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] 1.1 Sample and Sequencing Data Preprocessing: This study included samples from 134 patients with pulmonary nodules diagnosed by the gold standard of pathology. Among them, 104 cases were malignant nodules (covering major non-small cell lung cancer subtypes such as lung adenocarcinoma and lung squamous cell carcinoma) and 30 cases were benign nodules (including common benign lesions such as inflammatory nodules, tuberculomas, hamartomas, and sclerosing alveolar cell tumors). All samples were collected from peripheral venous blood of patients before their initial diagnosis and before they received any anti-tumor treatment. Peripheral blood mononuclear cells (PBMCs) were separated using density gradient centrifugation, and high-purity T cell genomic DNA was extracted using a genomic DNA extraction kit. After passing agarose gel electrophoresis and Qubit quantitative detection, TCRβ chain high-throughput sequencing was performed.

[0031] Raw sequencing data underwent standardized preprocessing using MiXCR software: the `mixcr align` command was executed to align V / D / J gene fragments and identify the CDR3 region, allowing a maximum of two mismatches; the `mixcr assemble` command was executed to assemble clonogenic sequences, removing chimeric sequences, non-functional sequences (including stop codons and frameshift mutations), and low-quality sequences (Q value < 20); the final output was a clonogenic file containing clonogenic nucleotide sequences, CDR3 amino acid sequences, V / J gene names, and clone counts. Further formatting and filtering were performed using the `vdjtools FilterNonFunctional` and `vdjtools DownSample` commands to obtain standardized TCR clonogenic raw data.

[0032] 1.2 Extraction of multidimensional molecular features from TCR immune repertoire: Based on a self-built Python computational analysis pipeline, the aforementioned clonal raw data were mapped into a six-dimensional quantitative feature matrix, and a pure TCR-derived feature layer was further constructed. The construction process strictly adhered to the principle of clinical data isolation: all imaging phenotypes (such as GGN, nodule size, density, and morphology) and clinical baseline data (such as age, sex, smoking history, TNM stage, and tumor marker levels) were forcibly removed to ensure that the feature matrix contained only TCR molecular features. The six dimensions correspond to gene usage, clonal amplification, CDR3 core recognition region, sequence convergence, tumor-related enrichment, and ecological diversity, respectively. The extraction and quantification methods for each dimension are as follows: (1) V / J gene usage preference: The clonality counts of each sample were transformed by natural logarithmic transformation to eliminate the influence of extreme values, and the weighted relative frequency of each TRBV and TRBJ gene fragment in the TCR repertoire of that sample was calculated. This feature reflects the clonal origin and developmental lineage of T cells, and tumor-specific antigens drive the dominant expansion of T cells with specific V / J gene combinations.

[0033] (2) Clonal frequency and length distribution: Based on the clonal percentage, TCR clonal types were strictly divided into five levels: extremely amplified (>1%), large (0.01%-1%), medium (0.001%-0.01%), small (0.0001%-0.001%), and rare (<0.0001%). The cumulative distribution frequency of each level of clonal type was calculated. At the same time, the CDR3 amino acid sequence length was divided into five intervals: ≤10, 11-13, 14-16, 17-19, and ≥20, and the cumulative distribution frequency of each interval was calculated. This feature reflects the overall pattern of T cell clonal expansion. Malignant tumor patients usually exhibit significant oligoclonal expansion.

[0034] (3) Amino acid motif composition of CDR3 region: Strict sequence quality control standards were set, and only standard sequences meeting the following conditions were retained: length 10-24 amino acids, clone count ≥10, starting with cysteine ​​(C) and ending with phenylalanine (F), and without a stop codon; after removing 3 conserved amino acids at each end of the CDR3 sequence (CXXX at the C-terminus and XXXF at the N-terminus), the weighted frequencies of 20 standard amino acids in the CDR3 core recognition region were calculated. The CDR3 region is the core region for TCR-specific recognition of antigen peptide-MHC complex, and its amino acid composition directly reflects the antigen recognition specificity of T cells.

[0035] (4) TCR sequence convergence: The number of clones with different nucleotide sequences but translated into the same combination of "V gene + CDR3 amino acid" is counted, and the proportion of them to the total number of clones is calculated. This feature is used to assess the degree of antigen-driven T cell convergent evolution. Tumor-specific antigens can induce different T cell clones to produce the same antigen recognition receptor, resulting in a significant increase in sequence convergence.

[0036] (5) Tumor-specific enrichment score: The quality-controlled TCR sequences were compared with known lung cancer-related and healthy human TCR databases (including KRAS mutation, EGFR mutation enrichment libraries and 1000 healthy human control libraries), and equal-length sequences were extracted. The normalized Levenshtein edit distance and Needleman-Wunsch local alignment similarity based on the BLOSUM62 matrix were calculated. Sequences with edit distance ≤0.3 and similarity score ≥4.5 were selected, and the tumor-specific enrichment score was calculated using the following formula: ;

[0037] In the formula, This represents the total number of valid TCR sequences after sample quality control. This represents the local similarity between the sequence to be tested and sequences in the database. This is the clone count of the sequence to be tested in the sample. This score represents the clone count of the matching sequence in the database. It quantifies the enrichment of tumor-associated TCR clones in the sample.

[0038] (6) Ecological diversity indicators: Using vdjtools software, under a unified... All samples were resampled without replacement at the desired sequencing depth to eliminate the impact of sequencing depth differences on diversity calculations. Shannon entropy, Simpson index, and Chao1 index were calculated to quantify the overall polymorphism of the TCR repertoire. Patients with malignant tumors typically exhibit reduced TCR repertoire diversity and increased clonality.

[0039] 1.3 Feature dimensionality reduction based on stable screening using RF-LASSO-RFE consensus: To address the characteristics of high dimensionality, sparsity, strong noise, and multicollinearity in TCR immune repertoire data, this invention employs a three-level stable screening strategy: RF-LASSO-RFE. This strategy does not rely on a single algorithm for importance ranking. Instead, it first captures nonlinear classification contributions using Random Forest, then extracts linear sparse discriminative contributions using LASSO, subsequently constrains the candidate range with an intersection consensus feature pool, and finally performs recursive feature elimination within a cross-validation framework to obtain a stable and biologically meaningful optimal subset of TCR features. The specific steps are as follows: (1) Initial screening of nonlinear features in random forest (RF): constructing a system containing 500 decision trees ( , An ensemble classifier is used, which measures the contribution of each feature to splitting sample heterogeneity using the average reduction in Gini impurity. The Gini impurity of node m... Defined as: ;

[0040] in, (Corresponding to benign and malignant samples) Let m be the proportion of samples belonging to class c. Extract the top 30 core features that contribute most to the reduction of Gini impurity, forming a non-linear feature subset. Random forests excel at capturing the non-linear relationship between features and labels, and are highly robust to outliers and noise.

[0041] (2) Initial screening of LASSO linear features: First, the original feature matrix is ​​standardized using the Standard Scaler to eliminate the influence of feature dimension differences on the LASSO model; then, a model based on... Regularized logistic regression model (hyperparameters) Penalty coefficient , By optimizing the cost function with an absolute value penalty term, redundant feature weights that contribute little to classification are forced to absolute zero. Cost function Defined as: ;

[0042] in, N For the total sample size, For the first i The true label of each sample To predict probabilities, For the first j The regression coefficients of each feature are calculated. Features corresponding to all non-zero coefficients are extracted to form a linear feature subset. LASSO excels at processing high-dimensional sparse data and can effectively remove multicollinear features to obtain sparse feature subsets.

[0043] (3) Consensus Feature Pool Construction: Extract the intersection of the nonlinear feature subset and the linear feature subset to obtain a consensus feature pool that simultaneously possesses nonlinear classification value and linear sparse discriminative evidence. This step transforms feature retention from a single algorithm selection to a consistency test of different discriminative mechanisms, which can reduce the probability of random noise, locally correlated or multicollinear features entering subsequent models.

[0044] (4) Recursive Feature Elimination (RFE) Deep Screening: Under the strict 5-fold cross-validation closed framework, multiple rounds of recursive elimination are performed on the consensus feature pool. In each round, variables with low contribution are eliminated based on the model's feature contribution in the training fold and the performance in the validation fold. The validation performance under different feature subset sizes is compared, and the optimal TCR feature subset that makes the cross-validation performance stable is finally output as the sole input variable of the downstream diagnostic model.

[0045] 1.4 Two-level heterogeneous stacking probabilistic fusion and training: This embodiment faces the severe challenge of extremely imbalanced sample classes (104 malignant cases: 30 benign cases, a ratio of approximately 3.5:1). If the original data is used directly to train the model, the classifier will exhibit a severe decision boundary skew towards the majority class (malignant), resulting in an extremely low recognition rate for benign samples. To address this issue, a lossless upsampling technique is introduced within each training fold of cross-validation: benign samples are randomly copied with replacement, forcing the inter-class distribution of benign and malignant samples in the training set to reach 1:1. This method is only executed within the training fold, while the validation fold maintains the original distribution, thus solving the class imbalance problem while avoiding data leakage and the introduction of artificial noise.

[0046] To overcome the performance bottleneck and generalization limitations of single algorithms in small-sample, high-dimensional TCR data, this invention constructs a two-level heterogeneous stacking probability fusion model. This model is not a simple voting or single classifier replacement, but rather feeds the same optimal subset of TCR features into multiple decision spaces with different mathematical assumptions, and then a meta-learner calibrates and fuses the malign prediction probabilities of each base learner. Specifically: (1) Basic Learner Layer (Level-0): Artificial neural network, radial basis support vector machine, gradient boosting tree, and random forest are selected as basic learners. Artificial neural network is used to capture complex nonlinear mappings, radial basis support vector machine is used to handle high-dimensional small sample boundary partitioning, gradient boosting tree is used to reduce bias and fit residual structure, and random forest is used to improve perturbation resistance and generalization stability. To avoid data crossing during the model fusion stage, the four basic learners are bound to the same cross-validation partition index for synchronous training, and the predicted probability of the current sample belonging to a malignant nodule is output on the corresponding validation fold.

[0047] (2) Meta-learner fusion layer (Level-1): The malignancy prediction probability matrices output by the four basic learners are used as new secondary features, and a generalized linear model (GLM, i.e., logistic regression) is selected as the meta-learner. GLM no longer directly reads the original TCR variables, but instead learns the calibration relationship and fusion weights between the probability outputs of different basic learners and the pathological labels, compressing the risk perspectives in multiple decision spaces into a continuous malignancy risk score. The final joint probability follows the following Sigmoid distribution: ;

[0048] In the formula, For the intercept term, Dynamically assign meta-learners to the first k The decision weight coefficients of each base classifier. If a base model exhibits excellent diagnostic performance on a specific subset of samples, its corresponding... The weights will be automatically amplified, thereby mitigating discrepancies between individual models and correcting conservative biases. As a meta-learner, GLM has the advantages of simple structure, stable training, and low overfitting, making it very suitable for fusing multiple probability outputs.

[0049] 1.5 Model Diagnostic Efficacy Validation: To objectively evaluate the clinical value of the golden TCR feature subset selected by RFE (Recursive Feature Elimination) in differentiating benign and malignant pulmonary nodules, this study conducted receiver operating characteristic (ROC) curve analysis on the prediction probabilities of four basic classifiers (NNet, SVM, GBM, RF) and the final heterogeneous stacking fusion model (Stacking_GLM) within a strict 5-fold cross-validation (5-FoldCV) framework.

[0050] In independent evaluations at the base learners level, all four algorithms with heterogeneous mathematical assumptions demonstrated good classification potential. Among them, Random Forest (RF), based on feature bootstrapping and perturbation-resistant ensemble, performed best, achieving an Area Under the Curve (AUC) of 0.815, establishing a high-level baseline for single-model approaches. Furthermore, Gradient Boosting Tree (GBM), Support Vector Machine (SVM), and Artificial Neural Network (NNet) also achieved considerable diagnostic efficiencies of 0.788, 0.780, and 0.776, respectively. Figure 1 (As shown by the gray dashed line). This result preliminarily confirms that the multidimensional TCR immune repertoire features extracted in this embodiment (such as sequence diversity, tumor-specific enrichment scores, and specific V / J gene preferences) contain extremely rich tumor identification information and can maintain robust interpretability under different nonlinear and spatial distance algorithms.

[0051] To further break through the performance ceiling of a single algorithm and resolve local decision discrepancies between models, this study introduces logistic regression (GLM) as a meta-learner to construct a Stacking fusion model. Results show that the Stacking_GLM model, which integrates the prediction probabilities of the four base learners (…),… Figure 1 (As shown by the solid red line in the middle) A comprehensive leap in diagnostic efficacy was achieved, with its final AUC climbing to 0.837, a further improvement of 2.2% compared to the best-performing monoclonal RF model.

[0052] Geometrically, the red solid line of Stacking_GLM encloses the gray dashed lines of all base models at most decision thresholds, exhibiting a particularly strong balance in the upper left corner of the curve. This result suggests that the risk information captured by different base learners from the same pure TCR feature subset is complementary. The GLM meta-learner, through probability calibration and weight learning, can integrate the outputs of multiple decision spaces, thereby achieving comprehensive diagnostic efficacy superior to a single model. It also indicates that, without relying on prior knowledge of traditional clinical phenotypes or CT imaging, pure TCR immune repertoire sequencing data combined with stable feature screening and a heterogeneous fusion architecture can construct a non-invasive auxiliary assessment tool for benign and malignant pulmonary nodules with clinical translational potential.

[0053] 1.6 Optimization Results and Feature Contribution Analysis of Core Biomarkers: To further reveal the microscopic molecular basis for the high-precision prediction achieved by the model of this invention, and to verify the effectiveness of the aforementioned "dual machine learning consensus and recursive elimination (RF-LASSO + RFE)" feature dimensionality reduction strategy, this embodiment performs feature importance (i.e., prediction contribution) analysis on the optimal feature subset that is ultimately selected and retained. The results are as follows: Figure 2 As shown.

[0054] Figure 2 This visually demonstrates the top 20 core TCR immune repertoire features contributing significantly to the prediction of the stacked fusion model. Analysis of this map reveals that the core biomarker combination extracted in this invention exhibits high multidimensional heterogeneity and synergistic complementarity, specifically manifested in: (1) V / J gene fragment usage preference plays an important role: such as Figure 2 As shown, the weighted relative frequencies of "J gene primer ratio" (J_primer_percent), "V gene primer ratio" (V_primer_percent), and specific gene families such as TRBV2, TRBV11-3, TRBV5-6, and TRBV7-8 ranked among the top in predictive contributions. This fully confirms that in the process of pulmonary nodule malignancy and tumorigenesis, tumor-specific antigens drive a significant shift in the use preference of specific T cell clonal receptor genes, and the abnormal expression of these specific genes constitutes an important molecular basis for differentiating between benign and malignant tumors.

[0055] (2) Multidimensional ecological and dynamic features participate in synergy: In addition to single gene expression features, features reflecting the ecological diversity of the immune repertoire (such as "inverse Simpson index standard deviation"), features reflecting clonal amplification dynamics (such as "clonal read ratio" and "functional clone ratio"), and features reflecting the amino acid physicochemical motifs of CDR3 region (such as L_score and F_score) all occupy an indispensable weight in the Top 20 core features.

[0056] (3) Prominent synergistic effect: From Figure 2 The distribution pattern shows that the predictive contribution of each feature exhibits a smoothly decreasing step-like distribution, rather than being monopolized by a single feature. This result indicates that the diagnostic system constructed in this invention does not rely on a single database matching score, a single diversity index, or isolated gene fragments, but rather utilizes a combination of composite biomarkers spanning multiple TCR dimensions, including gene preference, diversity, CDR3 physicochemical properties, and clonal amplification, for joint decision-making.

[0057] In summary, Figure 2The results further support the technical effects of this invention: This invention can screen molecular signal combinations that cover multiple layers of immunological information and have stable predictive contributions from high-dimensional, sparse TCR repertoire data. Relying entirely on this multidimensional composite TCR feature, and abandoning the traditional single-marker approach and without depending on prior features of CT imaging, non-invasive auxiliary assessment of the benign and malignant nature of pulmonary nodules can be achieved, demonstrating substantial advantages over single-indicator, single-model, or multimodal clinical imaging-dependent schemes.

[0058] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A risk prediction system for benign and malignant pulmonary nodules, characterized in that, A composition of reagents and / or instruments using biomarker information from an individual ex vivo TCR immune repertoire is used to prepare a risk prediction system for evaluating benign and malignant pulmonary nodules. This prediction system uses peripheral blood TCR immune repertoire sequencing data as the sole input, excluding clinical baseline information, CT imaging information, pathological information, or serum tumor marker information. The TCR immune repertoire biomarker information includes: V / J gene usage preference, clonal amplification and length distribution, CDR3 core amino acid motif composition, TCR sequence convergence, tumor-specific enrichment score, and ecological diversity indicators. The V / J gene usage preference is: the weighted relative frequency of TRBV and TRBJ gene fragments in the TCR immune repertoire after logarithmic transformation of TCR clonoid count; The frequency and length distribution of TCR clones include: the distribution frequency of TCR clones divided into five levels according to the clone proportion: extremely amplified, large, medium, small and rare; and the cumulative distribution frequency of CDR3 amino acid sequences divided into five intervals according to length. The core amino acid motif of CDR3 is composed of: After quality control screening, the TCR sequences were subjected to the removal of 3 conserved amino acids at each end, and the weighted frequencies of 20 standard amino acids in the CDR3 core recognition region were determined. The quality control screening criteria were: sequence length 10-24, count ≥10, starting with C and ending with F, and no stop codon. The TCR sequence convergence is defined as the percentage of clones with different nucleotide sequences but translated into the same "V gene + CDR3 amino acid" combination.

2. The risk prediction system for benign and malignant pulmonary nodules according to claim 1, characterized in that, The tumor-specific enrichment score It is calculated using the following formula: ; In the formula, This represents the total number of valid TCR sequences after sample quality control. Needleman-Wunsch local alignment similarity between the sequence to be tested and sequences in lung cancer-related TCR databases; This is the clone count of the sequence to be tested in the sample. The clone count of the matching sequence in the database is used; the criteria for screening and scoring sequences are: normalized edit distance ≤ 0.3 and similarity score ≥ 4.5; The ecological diversity index is: in a unified The Shannon entropy, Simpson index, and Chao1 index were calculated after resampling at the sequencing depth.

3. A risk prediction system for benign and malignant pulmonary nodules, characterized in that, Includes a data acquisition unit and a data analysis unit; The data acquisition unit is used to acquire TCR immune repertoire biomarker information from individual ex vivo samples; the TCR immune repertoire biomarker information includes: V / J gene usage preference, clonal frequency and length distribution, CDR3 region amino acid motif composition, TCR sequence convergence, tumor-specific enrichment score, and ecological diversity indicators. The data analysis unit is used to sequentially perform sequencing data quality control, six-dimensional immune characterization, derived feature construction, RF-LASSO-RFE stable feature screening, and two-level heterogeneous stacking probability fusion on the information taken by the data acquisition unit to obtain a continuous malignancy risk score for pulmonary nodules.

4. The risk prediction system for benign and malignant pulmonary nodules according to claim 3, characterized in that, The risk score for benign or malignant lung nodules obtained by the data analysis unit conforms to the value obtained according to the following model formula: ; In the formula, This indicates that the lung nodule is malignant. For the intercept term, , , , These are the decision weight coefficients assigned to the meta-learner for artificial neural networks, radial basis support vector machines, gradient boosting trees, and random forests, respectively. , , , These represent the probability of a malign prediction output by the corresponding base learner; The numerical value obtained according to the model formula is the probability score of malignant lesions of pulmonary nodules. The score reflects the individual's risk of malignant pulmonary nodules. Individuals with a malignant pulmonary nodule risk score ≥0.5 are judged to have a high risk of malignant pulmonary nodules.

5. An electronic device, comprising a first memory, a first processor, and a computer program stored in the first memory and executable on the first processor, characterized in that, When the first processor executes the program, it implements a scoring process including the following steps: The raw sequencing data of TCR immune repertoire from individual ex vivo samples were obtained, and TCR immune repertoire biomarker information with defined dimensions was extracted without introducing any non-TCR variables. The defined dimensions include V / J gene usage preference, clonal amplification and length distribution, CDR3 core amino acid motif composition, TCR sequence convergence, tumor-specific enrichment score, and ecological diversity indicators. Based on the obtained TCR immune repertoire biomarker information, proportional, offset, density, correction and co-variation class derived features were constructed, and the optimal TCR feature subset was obtained through the RF-LASSO-RFE stable screening process of random forest nonlinear initial screening, LASSO linear sparse initial screening, intersection consensus feature pool and recursive feature elimination. Based on the optimal TCR feature subset, a pre-trained two-level heterogeneous stacking fusion model is used to perform basic learner probability output, meta-learner probability calibration, and fusion weight learning to calculate the malignancy risk score of individual lung nodules.

6. The electronic device according to claim 5, characterized in that, When the first processor executes the program, the process of calculating the benign or malignant risk score of an individual's lung nodules includes: The optimal feature subset is input into the pre-trained base learner layer to obtain the malign prediction probabilities output by the artificial neural network, radial basis support vector machine, gradient boosting tree, and random forest, respectively. The predicted probabilities output by the base learner are input into the generalized linear model of the meta-learner layer to calculate the final risk score for malignancy and the risk score for benign or malignant lung nodules. These scores conform to the values ​​obtained according to the following model formula: ; In the formula, This indicates that the lung nodule is malignant. For the intercept term, These are the decision weight coefficients assigned to the meta-learner for artificial neural networks, radial basis support vector machines, gradient boosting trees, and random forests, respectively. These represent the probability of a malign prediction output by the corresponding base learner; The numerical value obtained according to the model formula is the probability score of malignant lesions of pulmonary nodules. The score reflects the individual's risk of malignant pulmonary nodules. Individuals with a malignant pulmonary nodule risk score ≥0.5 are judged to have a high risk of malignant pulmonary nodules.

7. The electronic device according to claim 6, characterized in that, When the first processor executes the program, it provides individualized health guidance and clinical reference suggestions based on the individual's risk score for benign or malignant pulmonary nodules.

8. The application of the electronic device according to any one of claims 5 to 7, characterized in that, Used to assess the risk of benign or malignant diseases in pulmonary nodules.