Tumor detection data analysis method based on peripheral blood immune standard
By systematically collecting and analyzing peripheral blood immune markers, and using multimodal network models to construct tumor immune microenvironment scores, it solves the problem that peripheral blood immune markers cannot accurately predict early tumor risks in the prior art, and achieves high sensitivity and high specificity early tumor detection.
Patent Information
- Application Number
- CN202510363292.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to accurately predict early tumor risk through peripheral blood immune markers. The sensitivity and specificity of traditional detection methods are insufficient, and the lack of systematic collection and integration analysis methods leads to an increase in false positive or false negative results.
Peripheral blood samples were collected, plasma and cell components were separated, immune cell phenotype was analyzed by flow cytometry, cytokine concentration and immune checkpoint molecular expression were measured, immune cell subpopulations ratio function was established, and immune lineage transcriptome analytical model of multimodal attention fusion network architecture was used to construct tumor immune microenvironment scores, and combined with clinical follow-up verification scoring system.
It realizes accurate assessment of tumor-related immune status, improves the sensitivity and specificity of early tumor detection, provides a reliable tumor risk prediction model, and provides an accurate decision-making basis for early clinical tumor screening and intervention.
Smart Images

Figure CN120253623A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of immunoassay, and specifically relates to a method for analyzing tumor detection data based on peripheral blood immune markers. Background Art
[0002] Early detection of tumors is crucial for improving the prognosis of patients. Traditional detection methods mainly rely on imaging examinations and determination of tumor markers. Liquid biopsy technology based on peripheral blood has received extensive attention in recent years due to its minimally invasive nature and potential for early detection. Currently, commonly used tumor markers in clinical practice, such as carcinoembryonic antigen (CEA), alpha-fetoprotein (AFP), etc., have limited specificity and sensitivity, resulting in a relatively high rate of missed diagnosis of early tumors; while detection technologies for new markers such as circulating tumor DNA (ctDNA) face problems such as high detection costs and low standardization.
[0003] The immune system plays a crucial role in the occurrence and development of tumors. Tumor-related immune responses can show significant changes in the early stage of tumor formation. However, existing detection methods based on single or a few immune indicators are difficult to comprehensively reflect the complexity of the tumor microenvironment, and lack effective data integration and analysis means, resulting in insufficient accuracy and reliability of detection results. False positive or false negative results often occur in clinical practice, increasing the psychological burden on patients and wasting medical resources.
[0004] Existing technologies are difficult to solve the problems of systematic collection and integration analysis of multi-dimensional immune indicators, lack a standardized method for comprehensively evaluating immune cell composition, cytokine network, and immune checkpoint expression, and cannot establish an accurate tumor risk prediction model, thus limiting the early tumor detection efficacy based on immune markers. That is to say, there is a technical problem in the prior art that peripheral blood immune markers cannot accurately predict the risk of early tumors. Summary of the Invention
[0005] In view of this, the present invention provides a method for analyzing tumor detection data based on peripheral blood immune markers, which can solve the technical problem in the prior art that peripheral blood immune markers cannot accurately predict the risk of early tumors.
[0006] The present invention is implemented as follows: The present invention provides a method for analyzing tumor detection data based on peripheral blood immune markers, including: collecting a peripheral blood sample from a patient and separating the plasma and cell components; performing phenotypic analysis on various immune cells in the blood sample using flow cytometry; measuring the concentrations of key cytokines in the peripheral blood; analyzing the expression levels of immune checkpoint molecules in peripheral blood mononuclear cells; establishing an empirical function for the proportion of immune cell subsets; applying a pre-trained immune lineage transcriptome analysis model to analyze the cytokine expression pattern, constructing a multi-level cytokine network association map based on a deep learning algorithm, and identifying key cytokine interaction pathways and regulatory nodes; combining the proportion of immune cell subsets with the output results of the immune lineage transcriptome analysis model to establish a tumor immune microenvironment score; verifying the accuracy of the scoring system through clinical follow-up and imaging examinations; and formulating a standardized report template to provide a basis for early tumor detection and monitoring for clinicians.
[0007] Among them, collecting a peripheral blood sample from a patient specifically includes: collecting 3 to 5 ml of whole blood using an anticoagulant tube, standing at room temperature for 30 to 60 minutes, and centrifuging at 1500 to 2000 revolutions per minute for 10 minutes at 4 to 10 °C to separate the plasma and cell components.
[0008] Among them, performing phenotypic analysis on various immune cells in the blood sample using flow cytometry specifically includes: analyzing the counts and proportions of CD3-positive T lymphocytes, CD4-positive helper T cells, CD8-positive cytotoxic T cells, CD19-positive B lymphocytes, and CD56-positive natural killer cells.
[0009] Among them, measuring the concentrations of key cytokines in the peripheral blood specifically includes: measuring interleukin-2, interleukin-6, interleukin-10, tumor necrosis factor-α, and interferon-γ, and quantitatively detecting their concentration levels using an enzyme-linked immunosorbent assay.
[0010] Among them, analyzing the expression levels of immune checkpoint molecules in peripheral blood mononuclear cells specifically includes: analyzing the expression densities of programmed death protein 1, programmed death ligand 1, cytotoxic T lymphocyte-associated antigen 4, and T cell immunoglobulin and mucin domain-containing protein 3.
[0011] Among them, establishing an empirical function for the proportion of immune cell subsets specifically includes: calculating the ratios of CD4 to CD8 cells, effector T cells to regulatory T cells, and neutrophils to lymphocytes, and determining the normal reference ranges for each ratio.
[0012] Among them, combining the proportion of immune cell subsets with the output results of the immune lineage transcriptome analysis model to establish a tumor immune microenvironment score specifically includes: classifying patients into high-risk, medium-risk, and low-risk groups according to the scoring results.
[0013] Among them, the specific structure of the immune lineage transcriptome analysis model is a multi-modal attention fusion network architecture, which includes three parallel encoder modules and an interactive decoder module. The encoder modules process the immune cell data stream, cytokine concentration data stream, and immune checkpoint expression data stream respectively.
[0014] Among them, the steps for establishing the training dataset of the immune lineage transcriptome analysis model include collecting peripheral blood samples from 10,000 confirmed tumor patients and 10,000 healthy control groups, and performing standardized flow cytometry analysis on all samples to obtain immune cell subset distribution data.
[0015] Among them, the steps for establishing the training dataset also include detecting the concentration levels of 40 cytokines using high-throughput enzyme-linked immunosorbent assay, and measuring the expression abundances of 20 immune checkpoint molecules by mass spectrometry analysis technology.
[0016] Compared with the prior art, the present invention provides a method for analyzing tumor detection data based on peripheral blood immune markers. The present invention proposes a method for analyzing tumor detection data based on peripheral blood immune markers. By systematically collecting and integratively analyzing the immune cell distribution, cytokine concentration, and immune checkpoint molecule expression in peripheral blood, a multi-level immune index evaluation system and a tumor risk prediction model are established. This method can not only comprehensively reflect the individual immune state, but also capture the early immune change characteristics related to tumors through advanced algorithm analysis.
[0017] Compared with the traditional single-index detection, the multi-dimensional immune marker analysis method of the present invention significantly improves the sensitivity and specificity of early tumor detection. By integrating flow cytometry, enzyme-linked immunosorbent assay, and deep learning algorithms, this method overcomes the defects of discrete data and low standardization degree in the prior art, and realizes the precise analysis and risk assessment of complex immune networks. Especially through the application of the immune lineage transcriptome analysis model, it is possible to identify key immune regulation nodes and pathway changes, providing a new theoretical basis and technical means for early tumor detection.
[0018] The present invention solves the technical problem that peripheral blood immune markers cannot accurately predict the early tumor risk. By establishing a systematic immune index evaluation system and a tumor risk prediction model, it realizes the precise evaluation of the tumor-related immune state, providing a more reliable and accurate decision-making basis for clinical early tumor screening and intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention.
[0021] As Figure 1 shown, it is a flowchart of a method for analyzing tumor detection data based on peripheral blood immune markers provided by the present invention. This method includes the following steps:
[0022] S01. Collect a peripheral blood sample from the patient, collect 3 to 5 milliliters of whole blood using an anticoagulant tube, let it stand at room temperature for 30 to 60 minutes, and centrifuge it at 1500 to 2000 revolutions per minute for 10 minutes at 4 to 10 °C to separate the plasma and cell components;
[0023] S02. Use flow cytometry to perform phenotypic analysis on various immune cells in the blood sample, including counting and proportion of CD3+ T lymphocytes, CD4+ helper T cells, CD8+ cytotoxic T cells, CD19+ B lymphocytes, and CD56+ natural killer cells;
[0024] S03. Measure the concentrations of key cytokines in peripheral blood, including interleukin 2, interleukin 6, interleukin 10, tumor necrosis factor α, and interferon γ, and quantitatively detect their concentration levels using enzyme-linked immunosorbent assay;
[0025] S04. Analyze the expression levels of immune checkpoint molecules in peripheral blood mononuclear cells, including the expression density of programmed death protein 1, programmed death ligand 1, cytotoxic T lymphocyte-associated antigen 4, and T cell immunoglobulin and mucin domain-containing protein 3;
[0026] S05. Establish an empirical function for the proportion of immune cell subsets, calculate the ratios of CD4 to CD8 cells, effector T cells to regulatory T cells, and neutrophils to lymphocytes, and determine the normal reference ranges for each ratio;
[0027] S06. Apply a pre-trained immune lineage transcriptome analysis model to analyze the cytokine expression pattern, construct a multi-level cytokine network association map based on a deep learning algorithm, and identify key cytokine interaction pathways and regulatory nodes;
[0028] S07. Combine the proportion of immune cell subsets with the output results of the immune lineage transcriptome analysis model to establish a tumor immune microenvironment score, and divide the patients into high-risk, medium-risk, and low-risk groups according to the score results;
[0029] S08. Verify the accuracy of the scoring system through clinical follow-up and imaging examinations, calculate sensitivity, specificity, positive predictive value, and negative predictive value, and determine the optimal cut-off value;
[0030] S09. Develop a standardized report template, including the analysis of the proportion of immune cell subsets, the output results of the immune lineage transcriptome analysis model, the tumor immune microenvironment scoring results and clinical interpretations, providing a basis for early detection and monitoring of tumors for clinicians.
[0031] Among them, programmed death protein 1 specifically refers to an immunosuppressive receptor located on the surface of T cells, and its high expression is closely related to tumor immune escape.
[0032] Among them, programmed death ligand 1 specifically refers to a ligand molecule present on the surface of tumor cells or immune cells, which inhibits T cell activation after binding to programmed death protein 1.
[0033] Among them, cytotoxic T lymphocyte-associated antigen 4 specifically refers to a negative regulatory molecule expressed on the surface of T cells, which weakens the T cell immune response by competitively inhibiting co-stimulatory signals.
[0034] Among them, T cell immunoglobulin and mucin domain-containing protein 3 specifically refers to a membrane protein involved in the process of T cell exhaustion, and its expression is increased in chronic infections and the tumor microenvironment.
[0035] Among them, the ratio of effector T cells to regulatory T cells specifically refers to the ratio of the number of effector T cells with immune killing function to the number of regulatory T cells with immunosuppressive function.
[0036] Among them, the tumor immune microenvironment score specifically refers to a numerical index obtained by comprehensively analyzing the composition of immune cell subsets and the characteristics of the cytokine network, and is used to evaluate the tumor-related immune status.
[0037] The specific structure of the immune lineage transcriptome analysis model is a multi-modal attention fusion network architecture, including three parallel encoder modules and an interactive decoder module. The encoder modules process the immune cell data stream, cytokine concentration data stream and immune checkpoint expression data stream respectively. Each encoder consists of three stacked self-attention layers, the number of attention heads is set to 8, the hidden layer dimension is 512, the feed-forward network dimension is 2048, and the scaling factor in the attention mechanism is calculated according to the square root of the input feature dimension. The decoder module uses the cross-attention mechanism to integrate the outputs of the three encoders, including two layers of multi-head cross-attention layers and one layer of self-attention layer. Finally, the tumor risk prediction score is output through a fully connected layer. Layer normalization in the model is applied to the input end of each sub-layer, and residual connections are applied to the output end of each sub-layer. The activation function uses the rectified linear unit function.
[0038] The steps for establishing the training dataset of the immune lineage transcriptome analysis model specifically include collecting peripheral blood samples from 10,000 confirmed cancer patients and 10,000 healthy control subjects, performing standardized flow cytometry analysis on all samples to obtain immune cell subset distribution data, detecting the concentration levels of 40 cytokines using high-throughput enzyme-linked immunosorbent assay, measuring the expression abundances of 20 immune checkpoint molecules by mass spectrometry analysis, constructing a structured data matrix in combination with the clinical information of the patients, dividing the dataset into a training set, a validation set, and a test set using the multi-center stratified sampling method, with the proportions being 70%, 15%, and 15%, performing standardization processing on all data to eliminate batch effects, reducing the dimension and removing redundant features through the principal component analysis method, establishing a complete set of feature labels for model training, and evaluating the generalization performance of the model using the five-fold cross-validation method.
[0039] The steps for training the immune lineage transcriptome analysis model specifically include initializing the model parameters using the normal distribution random initialization method, where the variance of weight initialization is inversely proportional to the input dimension, setting the learning rate to 0.0001 and dynamically adjusting it using the cosine annealing strategy. In each round of training, forward propagation calculation is performed using a mini-batch of size 64 to obtain the predicted values. The loss function adopts a combined form of weighted cross-entropy loss and contrastive learning loss, where the weight coefficient of the contrastive learning loss is directly determined by the ratio of neutrophils to lymphocytes. In the backpropagation stage, gradient clipping technology is used to prevent gradient explosion. The optimizer selects the adaptive moment estimation algorithm with momentum, setting the momentum coefficient to 0.9 and the weight decay coefficient to 0.00001. The early stopping strategy is used to avoid overfitting, and training stops when the performance of the validation set has not improved for 5 consecutive rounds. Finally, the model weights with the best performance on the validation set are selected as the training model parameters, and the performance of the model is evaluated by the harmonic mean of specificity and sensitivity, i.e., the F1 score. The training process is completed when the F1 score on the independent test set reaches above 0.92.
[0040] The following is a detailed description of the specific implementation manners of the above steps.
[0041] The specific implementation of step S01 is to collect the patient's peripheral blood using aseptic operation techniques and strictly implement the standard blood collection process. First, collect 3 to 5 ml of fresh whole blood using an anticoagulant tube containing sodium ethylenediaminetetraacetate or sodium heparin. After blood collection, gently invert and mix 8 to 10 times to ensure full contact between the blood and the anticoagulant. Then, let the blood sample stand at room temperature (18 to 25 °C) for 30 to 60 minutes to allow the blood to naturally stratify. Next, place the sample in a refrigerated centrifuge with a temperature controlled at 4 to 10 °C and centrifuge at 1500 to 2000 revolutions per minute (equivalent to a relative centrifugal force of 250 to 450 × g) for 10 minutes. Carefully aspirate the upper plasma into a sterile tube using a sterile pipette and collect the lower cell fraction separately. The separated plasma and cell fraction are stored in a -80 °C ultra-low temperature freezer or immediately subjected to subsequent analysis. This step aims to obtain high-quality blood sample components and provide the basic materials for subsequent immune cell and biomarker analysis.
[0042] The specific implementation of step S02 is to perform immunophenotyping analysis of the blood sample using multi-color flow cytometry technology. First, wash the cell fraction 2 times with phosphate buffer, centrifuging at 400 × g for 5 minutes each time to remove residual plasma. Subsequently, resuspend the cells in a buffer containing 2% fetal bovine serum and adjust the concentration to 1×10 6 cells / ml. Using the direct immunofluorescence labeling method, add a pre-optimized concentration of a fluorescent-labeled antibody mixture, including monoclonal antibodies such as anti-CD3-FITC (clone number UCHT1), anti-CD4-PE (clone number RPA-T4), anti-CD8-APC (clone number RPA-T8), anti-CD19-PE-Cy7 (clone number HIB19), anti-CD56-BV421 (clone number B159), etc., to 100 μl of the cell suspension and incubate in the dark at 4 °C for 30 minutes. After incubation, treat the sample with lysis buffer for 10 minutes to remove red blood cells, and then wash 2 times with phosphate buffer. Use a flow cytometer (such as BD FACSCanto II or an equivalent device) to collect at least 50,000 cell events, and analyze the absolute counts and relative percentages of CD3-positive T lymphocytes, CD3+CD4+ helper T cells, CD3+CD8+ cytotoxic T cells, CD19+B lymphocytes, CD3-CD56+ natural killer cells through fluorescence compensation and gating strategies. Data analysis is performed using dedicated flow cytometry software (such as FlowJo), and a hierarchical gating strategy is adopted to ensure the accuracy of subset identification. This step aims to comprehensively evaluate the distribution of immune cell subsets in the patient's peripheral blood and provide key data for immune function evaluation.
[0043] The specific implementation of step S03 is to detect the concentration of key cytokines in plasma by using a highly sensitive enzyme-linked immunosorbent assay. First, thaw the plasma sample and centrifuge it at 400×g for 10 minutes to remove any possible microparticles. Use a commercial enzyme-linked immunosorbent assay kit (sensitivity ≤ 1 picogram per milliliter) to detect the concentrations of interleukin 2 (IL-2), interleukin 6 (IL-6), interleukin 10 (IL-10), tumor necrosis factor α (TNF-α), and interferon γ (IFN-γ). The specific operations include: adding 100 microliters of diluted plasma sample into a 96-well microplate pre-coated with a specific capture antibody, and incubating at room temperature for 2 hours; after washing 5 times, add a biotin-labeled detection antibody and incubate at room temperature for 1 hour; after washing again, add a horseradish peroxidase-labeled avidin and incubate at room temperature for 30 minutes; finally, add the chromogenic substrate 3,3',5,5'-tetramethylbenzidine, and add a stop solution after an appropriate time, and measure the absorbance value at a wavelength of 450 nanometers using an enzyme-linked immunosorbent assay reader. Based on the standard curve (concentration range 1 to 1000 picograms per milliliter, R 2 value ≥ 0.98), calculate the cytokine concentration. Set two replicates for each sample and take the average as the final result. The lower limits of detection are: IL-2 (2 picograms per milliliter), IL-6 (0.7 picogram per milliliter), IL-10 (0.5 picogram per milliliter), TNF-α (1.6 picograms per milliliter), IFN-γ (4 picograms per milliliter). This step aims to accurately quantify the levels of key cytokines and reflect the state of the body's immune microenvironment and the characteristics of the inflammatory response.
[0044] The specific implementation of step S04 is to analyze the expression levels of immune checkpoint molecules in peripheral blood mononuclear cells by multi-parameter flow cytometry. First, use density gradient centrifugation to isolate peripheral blood mononuclear cells. Specifically, take 2 milliliters of anticoagulated whole blood, mix it with an equal volume of phosphate buffer, gently layer it on 3 milliliters of lymphocyte separation medium, centrifuge at 700×g for 20 minutes, carefully collect the mononuclear cells in the middle layer and wash them 2 times. Then adjust the mononuclear cells to 1×10 6cells / mL. Using a multi-color fluorescent antibody labeling strategy, fluorescently labeled antibodies such as anti-CD3-PerCP (clone number SK7), anti-PD-1-PE (clone number EH12.1), anti-PD-L1-PE-Cy7 (clone number MIH1), anti-CTLA-4-APC (clone number BNI3), anti-TIM-3-BV510 (clone number 7D3) were used and incubated for 45 minutes at 4°C in the dark. At the same time, corresponding isotype control antibody labeling groups were set as negative controls. After washing, flow cytometry was used to collect data, and the positive expression threshold was determined by isotype control calibration (usually set at the level of fluorescence intensity higher than 99.5% of the cells in the isotype control). The percentage of positive cells and the mean fluorescence intensity of each immune checkpoint molecule were calculated to reflect the expression density of programmed death protein 1 (PD-1), programmed death ligand 1 (PD-L1), cytotoxic T lymphocyte-associated antigen 4 (CTLA-4), and T cell immunoglobulin and mucin domain-containing protein 3 (TIM-3) on different immune cell subsets. This step aims to evaluate the activation status of the immune checkpoint pathway, predict the risk of immune escape, and the potential effect of immunotherapy.
[0045] The specific implementation of step S05 is to construct a series of ratio calculation models reflecting the balance state of immune cell subsets. First, the ratio of CD4 to CD8 cells was calculated based on flow cytometry data, and the normal reference range was set at 1.0 to 2.5. Values below 0.9 or above 3.0 were considered abnormal. The calculation method for the ratio of effector T cells to regulatory T cells was: dividing the count of CD3+CD4+CD25-CD127+ effector T cells by the count of CD3+CD4+CD25+CD127low / - regulatory T cells. The normal reference range was 4.0 to 12.0. Values below 3.0 indicated an immunosuppressive state, and values above 15.0 indicated an over-activated immune state. The ratio of neutrophils to lymphocytes was calculated as the absolute count of neutrophils divided by the absolute count of lymphocytes. The normal reference range was 1.2 to 2.5. Values above 3.0 indicated the formation of an inflammatory response or an immunosuppressive microenvironment. For each ratio index, a non-parametric quantile regression method was used to determine the reference range based on data from 1000 healthy controls, and the 5th percentile and the 95th percentile were set as the boundaries of the normal range. At the same time, the ratio deviation index calculation formula was introduced: deviation = (measured value - reference median) / reference range width × 100, which was used to quantify the degree of imbalance in the proportion of immune cell subsets. This step aims to provide a quantitative index reflecting the balance state of the immune system, making up for the deficiency that a single cell subset count cannot reflect the mutual relationship.
[0046] The specific implementation of step S06 is to construct and apply a complex deep learning model to analyze cytokine expression patterns. This model adopts a multi-modal attention fusion network architecture, which includes three parallel encoder modules and an interactive decoder module to process different types of immune data streams respectively. First, the measured cytokine concentration data is converted into normalized expression values to eliminate the influence of differences in the magnitude of different cytokine concentrations. Then, a cytokine interaction network is constructed based on a graph neural network, where the nodes represent cytokines and the edge weights are determined by the Pearson correlation coefficient (the threshold is set to 0.4). A gated graph convolutional network is applied to extract cytokine interaction features, the convolution kernel size is set to 3×3, the number of convolutional layers is 4 layers, and the number of output channels for each layer is 32, 64, 128, and 256 in sequence. Then, a self-attention mechanism is used to capture the differences in cytokine expression patterns in different immune microenvironments, the number of attention heads is set to 8, and the hidden layer dimension is 512. Then, cytokine expression subtypes are identified through a hierarchical clustering algorithm (using Euclidean distance and Ward linkage method), and the number of subtypes is determined by the silhouette coefficient (usually 4 to 6). Finally, an immune regulation network association map is constructed, the minimum spanning tree algorithm is used to determine the key regulatory pathways, and the core regulatory nodes are identified through the betweenness centrality index (the threshold is set to 0.1). This step aims to deeply analyze the cytokine network structure and functional characteristics in the immune microenvironment and reveal potential immune regulation mechanisms.
[0047] The specific implementation of step S07 is to integrate multi-dimensional immune data to construct a tumor immune microenvironment scoring system. First, an immune cell subset balance score is constructed, and the formula is: subset balance score = w1×(CD4 / CD8 ratio deviation) + w2×(effector T / regulatory T ratio deviation) + w3×(monocyte / lymphocyte ratio deviation), where w1, w2, and w3 are 0.3, 0.5, and 0.2 respectively, and the weight coefficients are optimized and determined through a stepwise multi-factor Cox regression model. Then, the cytokine network score is calculated. Based on the feature vectors output by the deep learning model in step S06, the first 3 principal components are extracted by principal component analysis, which explains more than 85% of the total variance, and the formula is constructed: cytokine network score = v1×PC1 + v2×PC2 + v3×PC3, where v1, v2, and v3 are the eigenvalue ratios corresponding to each principal component. Then, the immune checkpoint molecule expression score is calculated, and the formula is: checkpoint expression score = u1×PD-1 expression score + u2×PD-L1 expression score + u3×CTLA-4 expression score + u4×TIM-3 expression score, where u1 to u4 are all 0.25, and the expression scores are processed by Z-score normalization. Finally, the total tumor immune microenvironment score is calculated by integrating the three scores: microenvironment total score = 0.35×
[0048] Subgroup balance score + 0.4 × cytokine network score + 0.25 × checkpoint expression score. According to the total score, patients are divided into three groups: low-risk group (<25 points), medium-risk group (25 to 75 points), and high-risk group (>75 points). This step aims to comprehensively evaluate the immune status and tumor risk of patients and provide a basis for stratified management.
[0049] The specific implementation of step S08 is to evaluate the accuracy of the established scoring system through strict clinical verification. First, 300 newly diagnosed suspected tumor patients are selected for a prospective study. After completing the microenvironment score detection, a follow-up of 6 months to 2 years is arranged. The follow-up methods include routine imaging examinations (enhanced CT / MRI / PET-CT), tumor marker monitoring, and necessary pathological biopsies. According to the follow-up results, the true disease status of the patients (tumor positive or negative) is determined. Then, the diagnostic efficacy indicators of the scoring system are calculated: sensitivity = true positive / (true positive + false negative), specificity = true negative / (true negative + false positive), positive predictive value = true positive / (true positive + false positive), and negative predictive value = true negative / (true negative + false negative). At the same time, the receiver operating characteristic curve is plotted, and the area under the curve (AUC value) is calculated to evaluate the overall performance of the model. Generally, it is required that AUC ≥ 0.85 to have clinical application value. The optimal cut-off value is determined by the Youden index method (the maximum point of sensitivity + specificity - 1). The cut-off point between the low-risk and medium-risk groups is optimized to 30 points (sensitivity 92%, specificity 85%), and the cut-off point between the medium-risk and high-risk groups is optimized to 70 points (sensitivity 88%, specificity 90%). The Bootstrap resampling method (1000 repetitions) is used to evaluate the stability of the model to ensure that the width of the 95% confidence interval does not exceed 10%. This step aims to objectively evaluate the application performance of the scoring system in the real clinical environment and ensure its accuracy and reliability.
[0050] The specific implementation of step S09 is to design a comprehensive and standardized report template to ensure the effective communication of test results to clinicians. The report template is divided into four main parts: The first part is the basic information area, which includes patient demographics, sample collection time, test completion time, etc.; The second part is the analysis area of immune cell subset ratios, which presents the counts, percentages, and reference ranges of each immune cell in tabular form and uses color coding (green for normal, yellow for mildly abnormal, and red for significantly abnormal) to visually display the results; The third part is the output result area of the immune lineage transcriptome analysis model, which shows the cytokine network map and key regulatory nodes, and attaches the concentration values and reference ranges of each cytokine; The fourth part is the comprehensive evaluation area, which presents the total score and grouping of the tumor immune microenvironment, and at the same time provides a targeted clinical interpretation template, such as "Patients in the high-risk group are recommended to undergo in-depth tumor screening and short-term follow-up monitoring (interval not exceeding 3 months)". The report adopts a hierarchical design principle, with the primary information (total score and risk grouping) highlighted, the secondary information (specific indicators and reference ranges) clearly arranged, and the tertiary information (technical details and methodological explanations) attached at the end of the report. This step aims to standardize the result presentation method to ensure the effective transmission of test information and clinical decision-making support.
[0051] The specific implementation of the immune lineage transcriptome analysis model structure is an advanced data fusion analysis system designed based on a multi-modal deep learning architecture. The model adopts an encoder-decoder architecture, which includes three parallel encoder modules and an interactive decoder module. Each encoder is responsible for processing a specific type of data: The first encoder processes immune cell subset distribution data, with an input dimension of 15 (corresponding to 15 major immune cell subsets); The second encoder processes cytokine concentration data, with an input dimension of 40 (corresponding to 40 cytokines); The third encoder processes immune checkpoint expression data, with an input dimension of 20 (corresponding to 20 checkpoint molecules). Each encoder consists of three stacked self-attention layers, adopting a multi-head self-attention mechanism, with the number of attention heads set to 8 to capture the feature correlations in different subspaces. The hidden layer dimension is uniformly set to 512, the feed-forward network dimension is 2048, and the scaling factor calculation formula is (where d k(the input feature dimension of the attention layer) to ensure training stability. The decoder module adopts a two-stage fusion strategy: first, information exchange between the encoder outputs is achieved through two layers of multi-head cross-attention layers, with 8 attention heads set for each layer; then, the fused features are integrated through one self-attention layer. Finally, the fully connected layer consists of three layers of neural networks with dimensions of 256, 64, and 1 in sequence, and the last layer uses the Sigmoid activation function to output the tumor risk prediction score between 0 and 1. The model also adopts multiple techniques to improve performance: layer normalization is used before each sub-layer to stabilize training; residual connections are applied after each sub-layer to alleviate the problem of gradient disappearance; the rectified linear unit function (ReLU) is used as the activation function; an attention visualization mechanism is introduced to explain the model decision-making process. This model can effectively capture and integrate the complex associations between multi-omics data and provide a more accurate tumor risk assessment.
[0052] The specific implementation method for establishing the training data set of the immune lineage transcriptome analysis model is to systematically collect and organize large-scale multi-center clinical immune data. First, peripheral blood samples of 10,000 diagnosed cancer patients (including common cancer types such as lung cancer, liver cancer, colorectal cancer, breast cancer, etc.) and 10,000 healthy controls matched by age and gender were collected from the oncology centers of multiple hospitals. All samples were processed using standardized flow cytometry in the same central laboratory, and 15 major immune cell subsets were analyzed using a 16-color flow cytometer, including CD3+ T cells, CD4+ T cells, CD8+ T cells, memory T cells, effector T cells, regulatory T cells, B cell subsets, NK cell subsets, monocyte subsets, dendritic cell subsets, forming an immune cell atlas data set. At the same time, the concentrations of 40 cytokines were detected simultaneously using high-throughput enzyme-linked immunosorbent assay chip technology, including various interleukins, interferon families, chemokines, growth factors, etc., with a sensitivity of 0.1 picogram per milliliter. Subsequently, liquid chromatography-tandem mass spectrometry technology was used to quantitatively analyze the expression abundances of 20 immune checkpoint molecules, including key molecules such as PD-1, PD-L1, CTLA-4, TIM-3, LAG-3, etc. All data were integrated with the patient's clinical information (including tumor stage, histological type, survival prognosis, etc.) to construct a multi-dimensional structured data matrix. Using a multi-center stratified random sampling method, ensure that different centers and different tumor types are distributed in each subset according to the original proportion, and divide the entire data set into a training set, a validation set, and a test set at a ratio of 70%, 15%, and 15%. Perform batch effect correction on all data using the ComBat algorithm, and make different indicators comparable through Z-score standardization. Use principal component analysis for dimensionality reduction, retain the principal components with a cumulative contribution rate of 95%, and remove highly redundant features. Finally, a training data matrix containing 426 feature dimensions is constructed, and each sample corresponds to a binary label (tumor or healthy). Evaluate the performance stability of the model on different data subsets through five-fold cross-validation to ensure the generalization ability of the model.
[0053] The following details the mathematical models or calculation processes involved in the present invention.
[0054] The calculation process of the immune cell subset ratio empirical function in step S05 is specifically expressed as follows:
[0055]
[0056] In the formula, R CD4 / CD8 is the ratio of CD4 to CD8 cells; N CD4 is the CD4-positive T cell count (cells / μL); N CD8 is the CD8-positive T cell count (cells / μL).
[0057]
[0058] In the formula, R eff / reg is the ratio of effector T cells to regulatory T cells; N CD3+CD4+CD25-CD127+ is the count of effector T cells (cells / μL); N CD3+CD4+CD25+CD127low / - is the count of regulatory T cells (cells / μL).
[0059]
[0060] In the formula, R NLR is the ratio of neutrophils to lymphocytes; N Neu is the absolute count of neutrophils (cells / μL); N Lym is the absolute count of lymphocytes (cells / μL).
[0061]
[0062] In the formula, D i is the ratio deviation index; R i is the measured ratio; R median,i is the reference median of the ratio for the healthy population; R upper,i is the upper reference limit of the ratio for the healthy population (95th percentile); R lower,i is the lower reference limit of the ratio for the healthy population (5th percentile).
[0063] The methods for obtaining each parameter in the above formula are as follows: N CD4 , N CD8 , N CD3+CD4+CD25-CD127+ , N CD3+CD4+CD25+CD127low / - are obtained by flow cytometry analysis. For the specific operation, see step S02. After labeling cell surface molecules with fluorescent antibodies, data are collected using a flow cytometer, and the absolute count of each subset of cells is analyzed using professional software (such as FlowJo); N Neu and N Lym are obtained by completing a complete blood count test using an automated hematology analyzer; R median,i , R upper,i , R lower,i are calculated by analyzing the data of 1000 healthy controls using the non-parametric quantile regression method.
[0064] The reason why these ratio calculation formulas adopt a direct division relationship is mainly based on the principle of immunological balance, that is, the functional state of the immune system depends not only on the absolute number of individual cell subsets, but more on the relative balance between different functional subsets. R CD4 / CD8 reflects the balance between helper and cytotoxic T cells and is maintained within the normal range when the immune function is normal; R eff / reg reflects the balance between immune activation and inhibition, and this ratio is commonly abnormally decreased in cancer patients; R NLRReflects the balance between innate and acquired immunity, and its increase often indicates a chronic inflammatory state or the formation of a tumor microenvironment. Deviation index D i Through standardization, different ratios are converted into comparable relative deviation degrees, which is conducive to comprehensively evaluating the immune balance state.
[0065] The cytokine expression pattern analysis in step S06 involves the construction of a graph neural network, which is specifically represented as follows:
[0066] G = (V, E, W);
[0067] Wherein, G is the cytokine interaction network graph; V = {v1, v2,..., v n} is the node set, representing n cytokines; E = {e ij |v i , v j ∈V} is the edge set; W = {w ij |e ij ∈E} is the edge weight set.
[0068]
[0069] Wherein, w ij is the edge weight between cytokine v i and v j ; r ij is the Pearson correlation coefficient of the concentration values of cytokine v i and v j ; τ is the association threshold, set to 0.4.
[0070]
[0071] Wherein, H (l) is the node feature matrix output by the l-th layer of the graph convolutional network; A is the adjacency matrix, A ij = w ij ; I is the identity matrix; D is the degree matrix, D ii = ∑ j (A ij + I ij ); W (l) is the trainable weight matrix of the l-th layer; σ is the activation function, and the ReLU function σ(x) = max(0, x) is adopted.
[0072]
[0073] Wherein, BC(v i ) is the betweenness centrality of node v i ; σ st is the number of shortest paths from node s to node t in the network; σ st (vi ) is the number of the shortest paths from node s to node t passing through node v i .
[0074] The method for obtaining the above parameters is as follows: v i represents the cytokines to be analyzed, including 40 kinds in total, specifically IL-1β, IL-2, IL-4, IL-6, IL-8, IL-10, IL-12, IL-17, TNF-α, IFN-γ, etc., and their concentration values are obtained by the enzyme-linked immunosorbent assay method in step S03; r ij is obtained by calculating the Pearson correlation coefficient between the normalized concentration values of each cytokine; W (l) is obtained through the learning process of the model training, and the He initialization method is used for initialization; σ st and σ st (v i ) is obtained by calculating the number of paths in the network through the Dijkstra shortest path algorithm.
[0075] The graph neural network model for cytokine analysis is based on the principle of biological network. The cytokines are regarded as a complex regulatory network, in which there are interactions such as activation and inhibition between nodes. Constructing the edge weights using the Pearson correlation coefficient can reflect the co-variation pattern of cytokine expression levels, while the setting of the threshold τ filters out weak correlation relationships and highlights significant biological associations. The graph convolution operation captures the local network structure features by aggregating the information of neighbor nodes, introduces self-loops (identity matrix I) to retain the information of the nodes themselves, and the degree normalization ( term) balances the influence of nodes with different connection degrees. The betweenness centrality index is used to identify the "traffic hub" cytokines in the network. These cytokines often play key roles in immune regulation and have important influences on the formation of the tumor microenvironment.
[0076] The construction of the tumor immune microenvironment scoring system in step S07 involves multiple calculations, which are specifically expressed as follows:
[0077] S cell = w1×D CD4 / CD8 + w2×D eff / reg + w3×D NLR ;
[0078] In the formula, S cell is the immune cell subset balance score; D CD4 / CD8 is the CD4 / CD8 ratio deviation; D eff / reg is the effector T / regulatory T cell ratio deviation; D NLR is the neutrophil / lymphocyte ratio deviation; w1, w2, and w3 are weight coefficients, which are 0.3, 0.5, and 0.2 respectively.
[0079] Scyto = v1 × PC1 + v2 × PC2 + v3 × PC3;
[0080] In the formula, S cyto is the cytokine network score; PC1, PC2, and PC3 are the first three principal component values extracted by principal component analysis; v1, v2, and v3 are the weights corresponding to each principal component, equal to the ratio of their respective eigenvalues to the sum of the total eigenvalues.
[0081]
[0082] In the formula, Z i is the standardized score of immune checkpoint molecule expression; X i is the original expression level; μ i is the average value of the expression of this molecule in the healthy population; σ i is the standard deviation of the expression of this molecule in the healthy population.
[0083] S checkpoint = u1 × Z PD-1 + u2 × Z PD-L1 + u3 × Z CTLA-4 + u4 × Z TIM-3 ;
[0084] In the formula, S checkpoint is the immune checkpoint molecule expression score; Z PD-1 , Z PD-L1 , Z CTLA-4 , Z TIM-3 are the standardized expression scores of four key immune checkpoint molecules respectively; u1, u2, u3, and u4 are weight coefficients, all of which are 0.25.
[0085] S total = α × S cell + β × S cyto + γ × S checkpoint ;
[0086] In the formula, S total is the total score of the tumor immune microenvironment; S cell is the immune cell subset balance score; S cyto is the cytokine network score; S checkpoint is the immune checkpoint molecule expression score; α, β, and γ are comprehensive weight coefficients, which are 0.35, 0.4, and 0.25 respectively.
[0087] The method for obtaining the above parameters is: D CD4 / CD8 、D eff / reg 、D NLRDerived from the calculation results of step S05; w1, w2, and w3 are optimized and determined by applying a stepwise multivariate Cox regression model to the data of 1000 patients, with tumor diagnosis as the endpoint event in the model; PC1, PC2, and PC3 are obtained by performing principal component analysis on the 40-dimensional cytokine feature vectors output by the deep learning model in step S06; v1, v2, and v3 are determined by calculating the proportion of the eigenvalues of the corresponding principal components; X i is the mean fluorescence intensity of immune checkpoint molecules measured by flow cytometry, derived from step S04; μ i and σ i are obtained by analyzing the data of 500 healthy controls; α, β, and γ are determined by iterative optimization of the logistic regression model, with the goal of maximizing the AUC value of the model on the training set.
[0088] The tumor immune microenvironment scoring system adopts a linear weighted combination form, fully considering the characteristics of multiple levels of the immune system. The weight allocation reflects the relative importance of each component in predicting tumor risk: the cytokine network scoring weight is the highest (0.4), because the cytokine network state most directly reflects the tumor-related inflammation and immunosuppressive microenvironment; the immune cell subset balance scoring follows (0.35), reflecting the basic immune surveillance function; the immune checkpoint molecule expression scoring weight is the lowest (0.25), serving as a supplementary indicator of immune escape potential. Principal component analysis is used for dimensionality reduction of complex cytokine data, retaining the information in the direction of the largest variance and solving the problem of multicollinearity. Z-score normalization enables comparison of indicators of different magnitudes and eliminates the influence of measurement unit differences. The internal weight allocation of each sub-score is also based on biological significance. For example, in the cell subset score, the ratio of effector T / regulatory T cells has the largest weight (0.5), because it directly reflects the balance between anti-tumor immunity and immunosuppression.
[0089] The calculation of the diagnostic efficacy of the scoring system in step S08 is specifically expressed as follows:
[0090]
[0091] In the formula, Se is the sensitivity, indicating the ability of the scoring system to correctly identify tumor patients; TP is the number of true positive samples; FN is the number of false negative samples.
[0092]
[0093] In the formula, Sp is the specificity, indicating the ability of the scoring system to correctly exclude non-tumor patients; TN is the number of true negative samples; FP is the number of false positive samples.
[0094]
[0095] In the formula, PPV is the positive predictive value, indicating the probability of actually having a tumor when the score is positive; TP is the number of true positive samples; FP is the number of false positive samples.
[0096]
[0097] In the formula, NPV is the negative predictive value, indicating the probability of actually having no tumor when the score is negative; TN is the number of true negative samples; FN is the number of false negative samples.
[0098] J = Se + Sp - 1;
[0099] In the formula, J is the Youden index, used to determine the optimal cut-off value; Se is the sensitivity; Sp is the specificity.
[0100]
[0101] In the formula, AUC is the area under the receiver operating characteristic curve; ROC(t) is the curve function composed of sensitivity and 1 - specificity at different thresholds t.
[0102] CI 95% = [stat - 1.96×SE, stat + 1.96×SE];
[0103] In the formula, CI 95% is the 95% confidence interval; stat is the statistic to be evaluated (such as sensitivity, specificity, etc.); SE is the standard error estimated by Bootstrap resampling.
[0104] The methods for obtaining the above parameters are as follows: TP, FP, TN, and FN are obtained by comparing the immune microenvironment score results of 300 patients with newly diagnosed suspected tumors with the true disease status determined after 6 months to 2 years of follow-up; the ROC(t) curve is drawn by calculating the sensitivity and specificity at different cut-off values; SE is obtained by performing 1000 times of Bootstrap resampling on the original data, calculating the target statistic each time of resampling, and then finding its standard deviation.
[0105] These diagnostic efficacy indicators adopt conventional medical diagnostic test evaluation methods to construct a complete performance evaluation system. Sensitivity and specificity reflect the accuracy of the model in different dimensions, and their trade-off relationship is visually presented through the ROC curve. The Youden index is used to determine the optimal threshold point to maximize the sum of sensitivity and specificity, providing an objective basis for clinical decision-making. Positive predictive value and negative predictive value take into account the impact of disease prevalence on the diagnostic value, being more in line with the actual application requirements. As a comprehensive evaluation indicator, AUC quantifies the average performance of the model at all possible thresholds, and the standard above 0.85 ensures that the model has good clinical application prospects. Confidence interval estimation uses the Bootstrap resampling method to evaluate the stability and reliability of the statistical results, and the width controlled within 10% ensures the accuracy of the performance estimation of the scoring system.
[0106] Specifically, the principle of the present invention is as follows: The core technical principle of the present invention is based on the close association between the tumor immune microenvironment and the systemic immune response. The occurrence and development of tumors will cause a series of changes in the immune system, which are not only manifested in the local microenvironment but also reflected in the peripheral blood. The present invention constructs a multi-level and multi-dimensional immune index evaluation system by systematically collecting and analyzing the distribution of immune cell subsets, the characteristics of cytokine networks, and the expression of immune checkpoint molecules in peripheral blood.
[0107] The key innovation of this method lies in the immune lineage transcriptome analysis model constructed by using deep learning technology. This model is based on a multi-modal attention fusion network architecture and can simultaneously process and integrate three different types of immune data streams: immune cell data, cytokine concentration data, and immune checkpoint expression data. The self-attention mechanism in the model can capture the complex associations within the same data stream, while the cross-attention mechanism can effectively integrate the interactions between different data streams, thus comprehensively reflecting the changes in the tumor-related immune state. In addition, the introduction of contrastive learning loss enables the model to better distinguish the immune characteristics between the healthy state and the tumor state, and the design of the association between the weight coefficient in the loss function and the neutrophil-to-lymphocyte ratio reflects the special consideration of the tumor-related inflammatory response.
[0108] The present invention constructs a tumor immune microenvironment scoring system by establishing an empirical function to analyze the proportions of immune cell subsets and combining the output results of an immune lineage transcriptome analysis model, achieving a stratified assessment of the tumor risk of patients. The reason why this method can effectively solve the problem of predicting tumor risk using peripheral blood immune markers is that it not only considers the changes in individual immune indicators, but also pays more attention to the systematic changes in the entire immune network and the identification of key regulatory nodes, thereby capturing the subtle but specific immune response characteristics during early tumorigenesis. At the same time, the construction of a large-scale balanced dataset and a rigorous model training process ensure that this method has good generalization performance and stability, enabling it to provide reliable tumor risk assessment results in actual clinical applications.
[0109] A specific Example 1 of the present invention is provided below. The specific implementation of each step in this Example 1 is described in detail as follows.
[0110] The specific implementations of steps S01 - S04 are the same as those described above and will not be elaborated here.
[0111] The specific implementation of step S05 is to construct a series of ratio calculation models reflecting the balance state of immune cell subsets. Based on the cell count data obtained by flow cytometry, the following key ratios are calculated: the ratio of CD4 to CD8 cells, calculated according to the formula where N CD4 is the count of CD4+ T cells (cells / μL), N CD8 is the count of CD8+ T cells (cells / μL), and the normal reference range is set to 1.0 to 2.5. Values below 0.9 or above 3.0 are considered abnormal; the ratio of effector T cells to regulatory T cells, calculated according to the formula where N CD3+CD4+CD25-CD127+ is the count of effector T cells (cells / μL), N CD3+CD4+CD25+CD127low / - is the count of regulatory T cells (cells / μL), and the normal reference range is 4.0 to 12.0. Values below 3.0 indicate an immunosuppressive state, and values above 15.0 indicate an over-activated immune state; the ratio of neutrophils to lymphocytes, calculated according to the formula where N Neu is the absolute count of neutrophils (cells / μL), N Lym is the absolute count of lymphocytes (cells / μL), and the normal reference range is 1.2 to 2.5. Values above 3.0 indicate the formation of an inflammatory response or an immunosuppressive microenvironment. For each ratio index, its deviation index is also calculated, according to the formula where D i is the ratio deviation index, R i is the measured ratio, R median,i is the median reference value of the ratio in the healthy population, R upper,iis the upper limit of the ratio reference for the healthy population (95th percentile), R lower,i is the lower limit of the ratio reference for the healthy population (5th percentile). These reference values are determined based on non-parametric quantile regression analysis of the data of 1000 healthy controls. This step aims to provide a quantitative index reflecting the balance state of the immune system and make up for the deficiency that the counting of a single cell subset cannot reflect the mutual relationship.
[0112] The specific implementation of step S06 is to construct and apply a complex deep learning model to analyze the cytokine expression pattern. First, construct a cytokine interaction network graph G=(V, E, W), where V={v1, v2,..., v n} is the node set, representing n cytokines; E={e ij |v i , v j ∈V} is the edge set; W={w ij |e ij ∈E} is the edge weight set. The edge weight w ij is determined by calculating the Pearson correlation coefficient of the cytokine concentration, expressed as where r ij is the Pearson correlation coefficient of the concentration values of cytokines v i and v j , τ is the association threshold, set to 0.4. Then apply the graph convolutional network to extract cytokine interaction features, and the calculation of each layer is expressed as where H (l) is the node feature matrix output by the l-th layer of the graph convolutional network, A is the adjacency matrix, A ij =w ij , I is the identity matrix, D is the degree matrix, D ii =∑ j (A ij +I ij ), W (l) is the trainable weight matrix of the l-th layer, σ is the activation function, and the ReLU function σ(x)=max(0, x) is adopted. The network has 4 layers, and the number of output channels of each layer is 32, 64, 128, and 256 in sequence. Then use the self-attention mechanism to capture the expression pattern differences of cytokines in different immune microenvironments, and the attention calculation is expressed as where Q is the query matrix, K is the key matrix, V is the value matrix, d k is the input feature dimension of the attention layer, and the number of attention heads is set to 8. Identify the key regulatory nodes in the network by calculating the node betweenness centrality where σ st is the number of the shortest paths from node s to node t in the network, σ st (v i ) is the number of times passing through node vi The number of shortest paths from node s to node t, with the betweenness centrality threshold set to 0.1. This step aims to deeply analyze the structural and functional characteristics of the cytokine network in the immune microenvironment and reveal potential immune regulation mechanisms.
[0113] The specific implementation of step S07 is to integrate multi-dimensional immune data to construct a tumor immune microenvironment scoring system. First, construct the immune cell subset balance score, calculated according to the formula S cell = w1×D CD4 / CD8 + w2×D eff / reg + w3×D NLR where S cell is the immune cell subset balance score, D CD4 / CD8 is the deviation degree of the CD4 / CD8 ratio, D eff / reg is the deviation degree of the effector T / regulatory T cell ratio, D NLR is the deviation degree of the neutrophil / lymphocyte ratio, and w1, w2, w3 are weight coefficients, which are 0.3, 0.5, and 0.2 respectively, determined by optimizing the stepwise multi-factor Cox regression model. Then calculate the cytokine network score, calculated according to the formula S cyto = v1×PC1 + v2×PC2 + v3×PC3, where S cyto is the cytokine network score, PC1, PC2, PC3 are the first three principal component values extracted by principal component analysis, and v1, v2, v3 are the weights corresponding to each principal component, equal to the proportion of their respective eigenvalues in the sum of the total eigenvalues. Next, calculate the immune checkpoint molecule expression score. First, standardize the original expression level, where Z i is the standardized score of the immune checkpoint molecule expression, X i is the original expression level, μ i is the average value of the expression of this molecule in healthy people, and σ i is the standard deviation of the expression of this molecule in healthy people; then calculate according to the formula S checkpoint = u1×Z PD-1 + u2×Z PD-L1 + u3×Z CTLA-4 + u4×Z TIM-3 where S checkpoint is the immune checkpoint molecule expression score, Z PD-1 , Z PD-L1 , Z CTLA-4 , Z TIM-3 are the standardized expression scores of four key immune checkpoint molecules respectively, and u1, u2, u3, u4 are weight coefficients, all of which are 0.25. Finally, calculate the total tumor immune microenvironment score by integrating the three scores, according to the formula S total = α×S cell + β×Scyto + γ × S checkpoint Calculation, where S total is the total score of the tumor immune microenvironment, and α, β, and γ are comprehensive weight coefficients, which are 0.35, 0.4, and 0.25 respectively, and are determined by iterative optimization of the logistic regression model. According to the total score, patients are divided into three groups: low-risk group (<25 points), medium-risk group (25 to 75 points), and high-risk group (>75 points). This step aims to comprehensively evaluate the immune status and tumor risk of patients and provide a basis for stratified management.
[0114] The specific implementation of step S08 is to evaluate the accuracy of the established scoring system through strict clinical verification. Select 300 newly diagnosed suspected tumor patients for a prospective study, and arrange a follow-up of 6 months to 2 years after completing the microenvironment score detection. Determine the true disease status (tumor positive or negative) of the patients according to the follow-up results, and calculate the diagnostic efficacy indicators of the scoring system: sensitivity where Se is the sensitivity, TP is the number of true positive samples, and FN is the number of false negative samples; specificity where Sp is the specificity, TN is the number of true negative samples, and FP is the number of false positive samples; positive predictive value where PPV is the positive predictive value; negative predictive value where NPV is the negative predictive value. Draw the receiver operating characteristic curve and calculate the area under the curve where AUC is the area under the receiver operating characteristic curve, and ROC(t) is the curve function composed of sensitivity and 1 - specificity at different thresholds t. Determine the optimal cut-off value through the Youden index J = Se + Sp - 1, where J is the Youden index. Use the Bootstrap resampling method to evaluate the stability of the model and calculate the 95% confidence interval CI 95% = [stat - 1.96 × SE, stat + 1.96 × SE], where CI 95% is the 95% confidence interval, stat is the statistic to be evaluated, and SE is the standard error estimated by Bootstrap resampling. This step aims to objectively evaluate the application performance of the scoring system in the real clinical environment and ensure its accuracy and reliability.
[0115] The specific implementation of step S09 is the same as the foregoing and will not be elaborated in detail here.
[0116] To better understand and implement the present invention, Example 2 of a specific application scenario of the present invention is provided below: Researchers selected 120 suspected lung cancer patients and 60 healthy physical examination subjects who visited the respiratory department of a certain tertiary hospital in the north from January 2023 to December 2023 for a prospective study, aiming to evaluate the clinical application value of a tumor detection method based on peripheral blood immune markers in early lung cancer screening.
[0117] First, the researchers strictly collected peripheral blood samples from all participants according to step S01. Each subject had 4 ml of peripheral venous blood collected on an empty stomach in the early morning, which was collected using an EDTA-K2 anticoagulant tube. After standing at room temperature for 45 minutes, it was centrifuged at 1800 revolutions per minute for 10 minutes at 8°C to separate the plasma and cell components, which were numbered and stored separately. The plasma was stored in an ultra-low temperature freezer at -80°C, and the cell components were stored at 4°C for no more than 6 hours to ensure the sample quality.
[0118] According to step S02, the researchers used a BD FACSCanto II flow cytometer to perform standardized 10-color flow cytometry analysis on all samples to measure the distribution of each immune cell subset. The main detection indicators included the absolute count and percentage of CD3+ T lymphocytes, CD3+CD4+ helper T cells, CD3+CD8+ cytotoxic T cells, CD19+ B lymphocytes, and CD3-CD56+ NK cells. The test results of some samples are shown in Table 1:
[0119] Table 1 Analysis results of peripheral blood immune cell subsets of some subjects
[0120]
[0121] According to step S03, the researchers used a highly sensitive enzyme-linked immunosorbent assay to detect the levels of key cytokines in peripheral plasma. The Bio-Plex Pro TM Human Cytokine Multiplex Assay System (Bio-Rad) was used to simultaneously measure the concentrations of five core cytokines, namely IL-2, IL-6, IL-10, TNF-α, and IFN-γ. The test results of some samples are shown in Table 2:
[0122] Table 2 Detection results of cytokine concentrations in peripheral plasma of some subjects
[0123]
[0124]
[0125] According to step S04, the researchers analyzed the expression levels of immune checkpoint molecules in peripheral blood mononuclear cells. First, mononuclear cells were separated by density gradient centrifugation, and then flow cytometry was used to detect the expression density of four key checkpoint molecules, namely PD-1, PD-L1, CTLA-4, and TIM-3, on different immune cell subsets. The test results of some samples are shown in Table 3:
[0126] Table 3 Expression of immune checkpoint molecules in peripheral blood of some subjects
[0127]
[0128] According to step S05, the researchers calculated the ratios of various immune cell subsets. Using the formula to calculate the CD4 / CD8 ratio, using the formula to calculate the ratio of effector T cells to regulatory T cells, using the formula to calculate the ratio of neutrophils to lymphocytes. Then using the formula to calculate the deviation index of each ratio. Partial calculation results are shown in Table 4:
[0129] Table 4 Calculation results of immune cell subset ratios of some subjects
[0130]
[0131] According to step S06, the researchers analyzed the cytokine expression patterns using a pre-trained immune lineage transcriptome parsing model. Hierarchical clustering analysis was performed on the model output results, and a cytokine network association heatmap was drawn. There were significant differences in the cytokine network characteristics between the healthy control group and the suspected lung cancer group. The suspected lung cancer group showed obvious characteristics of inflammatory network activation and enhanced immunosuppressive network.
[0132] According to step S07, the researchers integrated all the above data to construct a tumor immune microenvironment scoring system. Using the formula S cell = w1×D CD4 / CD8 + w2×D eff / reg + w3×D NLR to calculate the immune cell subset balance score, where w1 = 0.3, w2 = 0.5, w3 = 0.2. Using the formula S cyto = v1×PC1 + v2×PC2 + v3×PC3 to calculate the cytokine network score, where v1, v2, and v3 are 0.65, 0.25, and 0.1 respectively. Using the formula to standardize the expression of immune checkpoint molecules, and then using the formula S checkpoint = u1×Z PD-1 + u2×Z PD-L1 + u3×Z CTLA-4 + u4×Z TIM-3 to calculate the immune checkpoint molecule expression score, where u1 to u4 are all 0.25. Finally, using the formula S total = α×S cell + β×S cyto + γ×S checkpoint to calculate the total tumor immune microenvironment score, where α = 0.35, β = 0.4, γ = 0.25. The scoring results of some patients are shown in Table 5:
[0133] Table 5 Scoring results of the tumor immune microenvironment of some subjects
[0134]
[0135]
[0136] According to step S08, the researchers completed a 12-month follow-up for all subjects. Among the 120 suspected lung cancer patients, 47 cases of lung cancer (including 32 cases of early lung cancer) were finally diagnosed by chest CT, bronchoscopy and pathological biopsy, and 73 cases were excluded from the diagnosis of lung cancer. Using the formula to calculate the sensitivity, using the formula to calculate the specificity, using the formula to calculate the positive predictive value, using the formula to calculate the negative predictive value. The performance evaluation results of the scoring system in lung cancer diagnosis are shown in Table 6:
[0137] Table 6 Performance evaluation of the tumor immune microenvironment scoring system in lung cancer diagnosis
[0138] Evaluation index Overall lung cancer Early-stage lung cancer 95% confidence interval Sensitivity (%) 91.5 87.5 83.2-95.7 Specificity (%) 88.3 85.6 81.4-92.6 Positive predictive value (%) 81.1 75.7 72.3-86.8 Negative predictive value (%) 94.7 92.8 89.5-97.2 AUC value 0.924 0.886 0.867-0.942
[0139] The optimal cut-off values were determined by the Youden index method: the cut-off point between low risk and medium risk was 30 points (sensitivity 93.6%, specificity 87.2%), and the cut-off point between medium risk and high risk was 72 points (sensitivity 89.4%, specificity 91.7%). In addition, the study also found that the total score of the tumor immune microenvironment was significantly positively correlated with tumor size and stage (r = 0.786, p < 0.001), suggesting that this scoring system may have the potential to reflect the degree of tumor progression.
[0140] According to step S09, the researchers generated a standardized report for each subject, which included four parts: basic information area, analysis area of immune cell subset proportions, output result area of the immune lineage transcriptome analysis model, and comprehensive evaluation area. The report detailed the numerical values of each index, the reference range, and the clinical interpretation. For patients with a high-risk score, the report recommended further tumor screenings such as enhanced chest CT and bronchoscopy, and follow-up every 3 months; for patients with a medium-risk score, routine chest CT examinations were recommended, and follow-up every 6 months; for patients with a low-risk score, routine physical examinations were recommended, and follow-up once a year.
[0141] Traditional early screening for lung cancer mainly relies on low-dose computed tomography (LDCT), which has problems such as a high false positive rate, radiation exposure, and high costs. Although blood tumor markers (such as CEA, CYFRA21-1, NSE, etc.) have less trauma and lower costs, their sensitivity in early lung cancer is poor, and they often increase significantly only in the advanced stage of the tumor. The method of the present invention is based on tumor detection technology with peripheral blood immune markers and has the following improvements compared with traditional methods: (1) comprehensively evaluates the immune microenvironment characteristics in three dimensions of immune cell subsets, cytokine networks, and immune checkpoint expressions, constructs a multi-level and multi-parameter tumor risk assessment system, which is more comprehensive than single marker detection; (2) applies deep learning technology to analyze complex immune regulation networks, reveals the early cytokine change patterns in the formation of the tumor microenvironment, and captures subclinical immune abnormalities that are difficult to detect by traditional methods; (3) innovatively introduces a variety of immune cell ratios and deviation indices, quantitatively evaluates the immune balance state, and makes up for the deficiency of traditional methods that only focus on absolute numbers while ignoring relative balance; (4) in this study, the sensitivity of this method for detecting early lung cancer reaches 87.5%, and the specificity is 85.6%, which is significantly better than traditional blood tumor markers (the sensitivity is usually <50%) and is close to the performance of LDCT, but there is no radiation risk; (5) through standardized reports and risk stratification-based management strategies, it provides a more accurate guiding basis for clinical decision-making and can effectively reduce the risks of overexamination and missed diagnosis.
[0142] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 7 and 8 below.
[0143] Table 7 Variable Explanation Table (First Part)
[0144]
[0145]
[0146] Table 8 Variable Explanation Table (Second Part)
[0147]
[0148]
[0149] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A method for analyzing tumor detection data based on peripheral blood immune markers, characterized in that, Including: Collecting a peripheral blood sample from a patient and separating plasma and cellular components; Performing phenotypic analysis of multiple immune cells in the blood sample using flow cytometry; measuring the concentrations of key cytokines in peripheral blood; Analyzing the expression levels of immune checkpoint molecules in peripheral blood mononuclear cells; Establishing an empirical function for the proportion of immune cell subsets; applying a pre-trained immune lineage transcriptome analysis model to analyze cytokine expression patterns, constructing a multi-level cytokine network association map based on a deep learning algorithm, and identifying key cytokine interaction pathways and regulatory nodes; Combining the proportion of immune cell subsets with the output results of the immune lineage transcriptome analysis model to establish a tumor immune microenvironment score; verifying the accuracy of the scoring system through clinical follow-up and imaging examinations; formulating a standardized report template to provide a basis for early tumor detection and monitoring for clinicians.
2. The method for analyzing tumor detection data based on peripheral blood immune markers according to claim 1, wherein Collecting a peripheral blood sample from a patient specifically includes: collecting 3 to 5 ml of whole blood using an anticoagulant tube, allowing it to stand at room temperature for 30 to 60 minutes, and centrifuging at 1500 to 2000 revolutions per minute at 4 to 10 °C for 10 minutes to separate plasma and cellular components.
3. The tumor detection data analysis method based on peripheral blood immune markers according to claim 2, wherein Performing phenotypic analysis of multiple immune cells in the blood sample using flow cytometry specifically includes: analyzing the counts and proportions of CD3+ T lymphocytes, CD4+ helper T cells, CD8+ cytotoxic T cells, CD19+ B lymphocytes, and CD56+ natural killer cells.
4. The method for analyzing tumor detection data based on peripheral blood immune markers according to claim 3, wherein Measuring the concentrations of key cytokines in peripheral blood specifically includes: measuring interleukin-2, interleukin-6, interleukin-10, tumor necrosis factor-α, and interferon-γ, and quantitatively detecting their concentration levels using enzyme-linked immunosorbent assay.
5. The method for analyzing tumor detection data based on peripheral blood immune markers according to claim 4, wherein Analyzing the expression levels of immune checkpoint molecules in peripheral blood mononuclear cells specifically includes: analyzing the expression densities of programmed death protein 1, programmed death ligand 1, cytotoxic T lymphocyte-associated antigen 4, and T cell immunoglobulin and mucin domain-containing protein 3.
6. The method for analyzing tumor detection data based on peripheral blood immune markers according to claim 5, wherein Establishing an empirical function for the proportion of immune cell subsets specifically includes: calculating the ratios of CD4 to CD8 cells, effector T cells to regulatory T cells, and neutrophils to lymphocytes, and determining the normal reference ranges for each ratio.
7. The method for analyzing tumor detection data based on peripheral blood immune markers according to claim 6, characterized in that, Combining the proportion of immune cell subsets with the output results of the immune lineage transcriptome analysis model to establish a tumor immune microenvironment score specifically includes: classifying patients into high-risk, medium-risk, and low-risk groups according to the scoring results.
8. The method for analyzing tumor detection data based on peripheral blood immune markers according to claim 7, wherein The specific structure of the immune lineage transcriptome analysis model is a multi-modal attention fusion network architecture, which includes three parallel encoder modules and an interactive decoder module. The encoder modules process immune cell data streams, cytokine concentration data streams, and immune checkpoint expression data streams respectively.
9. The method for analyzing tumor detection data based on peripheral blood immune markers according to claim 8, wherein The steps for establishing the training dataset of the immune lineage transcriptome analysis model include collecting peripheral blood samples from 10,000 confirmed tumor patients and 10,000 healthy control groups, and performing standardized flow cytometry analysis on all samples to obtain immune cell subset distribution data.
10. The method for analyzing tumor detection data based on peripheral blood immune markers according to claim 9, wherein The steps for establishing the training dataset also include detecting the concentration levels of 40 cytokines using high-throughput enzyme-linked immunosorbent assay, and measuring the expression abundances of 20 immune checkpoint molecules through mass spectrometry analysis technology.
Citation Information
Cited By
Method and device for establishing diagnosis and treatment system of digestive system disease multi-modal information
CN120998466A
Immune detection method of immune checkpoint detection kit
CN121703411A
Peripheral blood immune function scoring method based on immune tertiary structure
CN122337645A
Method for evaluating immune function in peripheral blood based on tertiary structure of immune
CN122337645B