A prognostic analysis marker for diffuse large b-cell lymphoma patient, a prognostic analysis system and application thereof

CN119673292BActive Publication Date: 2026-09-18SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411790343.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2026-09-18
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

尽管R-CHOP(利妥昔单抗、环磷酰胺、阿霉素及其脂质体、长春新碱、泼尼松)等一线治疗改善了DLBCL患者的生存,但耐药和复发患者的临床预后仍然不理想

Benefits of technology

[0023] The aforementioned technical solution establishes, for the first time, a prognostic analysis system based on key amino acids and clinically relevant indicators in DLBCL patients. It can effectively predict the prognosis of DLBCL patients and is closely related to their clinical characteristics. Furthermore, among the key amino acids related to prognosis, tryptophan and glutamine are key molecules, confirming their good prognostic value in DLBCL patients. Simultaneously, this technical solution integrates the prognostic model into an autonomous AI system, enabling autonomous operation, detection, and adjustment, thus improving efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119673292B_ABST
    Figure CN119673292B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of disease prognosis and molecular biology, and particularly relates to a diffuse large B-cell lymphoma (DLBCL) patient prognosis analysis marker, a prognosis analysis system and application thereof. The application first establishes a prognosis analysis system based on key amino acids and clinical indicators in DLBCL patients. The system can effectively predict the progression-free survival of DLBCL patients and is closely related to the clinical characteristics of the patients. Further, among the key amino acids related to prognosis, tryptophan and glutamine are key molecules, which confirms that tryptophan and glutamine have good prognostic value in DLBCL patients. At the same time, the above technical solution imports the prognosis model into an autonomous AI system for operation, realizes autonomous operation, autonomous detection and autonomous adjustment, and improves efficiency. The application provides a new DLBCL patient prognosis evaluation strategy, lays a theoretical foundation for individualized management of DLBCL patients, and therefore has good practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of disease prognosis and molecular biology technology, specifically relating to a prognostic biomarker, a prognostic analysis system, and its application for patients with diffuse large B-cell lymphoma. Background Technology

[0002] The information disclosed in this background section is intended only to enhance understanding of the overall background of the invention and is not necessarily to be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.

[0003] Diffuse large B-cell lymphoma (DLBCL) is the most common malignant tumor of the B-cell lineage, exhibiting high molecular heterogeneity and accounting for up to 40% of newly diagnosed non-Hodgkin lymphomas. Although first-line treatments such as R-CHOP (rituximab, cyclophosphamide, doxorubicin and its liposomes, vincristine, and prednisone) have improved the survival of DLBCL patients, the clinical prognosis of drug-resistant and relapsed patients remains unsatisfactory.

[0004] Metabolomics is a research approach that quantitatively analyzes all metabolites in an organism and seeks the relative relationships between metabolites and physiological and pathological changes. Using metabolomics methods, a series of metabolites can be detected in a single experiment to study low-molecular-weight metabolites present in biological systems. Bioinformatics is a discipline that uses computers to analyze, collect, and organize biological information such as DNA, proteins, amino acids, etc. It provides a foundation for the treatment of some major diseases, such as cardiovascular diseases, tumors, and infectious diseases.

[0005] Basic amino acids and amino acid-like substances are the building blocks of proteins and the most fundamental substances for life activities, playing crucial roles in specific physiological functions such as metabolism, neurotransmission, and liposome transport. They generate energy through intermediate metabolites, providing fuel for other biosynthetic pathways. Research has found that hematologic malignancies often rely on specific amino acids for survival, and insufficient production of these amino acids can lead to metabolic vulnerability and limited treatment opportunities. Some amino acid-targeting enzymes already approved for the treatment of hematologic malignancies have shown promising results. Quantitative analysis of amino acids in biological samples using metabolomics methods has played an important role in the diagnosis of hematologic diseases and the study of growth and development mechanisms.

[0006] Autonomous artificial intelligence (AI) refers to AI systems capable of performing tasks without direct human intervention. Such systems possess fundamental attributes such as autonomy, responsiveness, social capabilities, and initiative. Autonomous AI systems can perceive their environment, take actions to maximize their chances of success, and learn and improve their performance through techniques such as machine learning, deep learning, and reinforcement learning. Autonomous AI has broad application prospects in the healthcare field, gradually changing traditional healthcare service models and improving the quality and efficiency of medical services. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a prognostic biomarker, a prognostic analysis system, and its application for patients with diffuse large B-cell lymphoma (DLBCL). Specifically, this invention utilizes liquid chromatography-tandem mass spectrometry (LC-MS / MS) to analyze targeted metabolomics data of plasma amino acid profiles from DLBCL patients. Principal component analysis (PCA) and sparse partial least squares discriminant analysis (sPLS-DA) were used to perform bioinformatics analysis on the amino acid sequencing screening results. Finally, partial least squares regression (PLSR) and support vector machine (SVR) were applied to establish a prognostic model, which was then imported into an autonomous AI system for autonomous operation. This further validated the good prognostic value of serum tryptophan and glutamine in DLBCL patients. Based on the above research results, this invention was completed.

[0008] To achieve the above technical objectives, the present invention adopts the following technical solution:

[0009] In a first aspect, the present invention provides a prognostic biomarker for patients with diffuse large B-cell lymphoma, the prognostic biomarker comprising key amino acids and clinically relevant indicators of DLBCL;

[0010] The amino acids include any one or more of alanine, arginine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, glycine, asparagine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.

[0011] Furthermore, the key amino acids are tryptophan and glutamine.

[0012] The clinically relevant indicators for DLBCL include any one or more of the following: IPI score at initial admission, IPI risk stratification, age, number of extranodal involvement sites, and lactate dehydrogenase (LDH).

[0013] The prognostic analysis included an analysis and assessment of progression-free survival (PFS) in patients with DLBCL.

[0014] A second aspect of the invention provides the use of reagents for detecting the above-mentioned prognostic biomarkers in the preparation of prognostic analysis products for patients with diffuse large B-cell lymphoma.

[0015] A third aspect of the present invention provides a prognostic analysis system for diffuse large B-cell lymphoma, the prognostic analysis system comprising:

[0016] The acquisition unit is configured to acquire the aforementioned prognostic analysis biomarkers of the subjects;

[0017] An assessment unit is configured to predict progression-free survival of the patient with diffuse large B-cell lymphoma based on prognostic assessment biomarkers obtained by the acquisition unit.

[0018] The output unit is configured to output the progression-free survival prediction results based on the evaluation unit.

[0019] The prognostic analysis system can be run autonomously using Python.

[0020] In a fourth aspect, the present invention provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, performs the functions of the diffuse large B-cell lymphoma prognostic analysis system as described in the third aspect of the present invention.

[0021] In a fifth aspect, the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor, when executing the program, performs the functions of the diffuse large B-cell lymphoma prognostic analysis system as described in the third aspect of the present invention.

[0022] Compared with existing technical solutions, one or more of the above technical solutions have the following beneficial effects:

[0023] The aforementioned technical solution establishes, for the first time, a prognostic analysis system based on key amino acids and clinically relevant indicators in DLBCL patients. It can effectively predict the prognosis of DLBCL patients and is closely related to their clinical characteristics. Furthermore, among the key amino acids related to prognosis, tryptophan and glutamine are key molecules, confirming their good prognostic value in DLBCL patients. Simultaneously, this technical solution integrates the prognostic model into an autonomous AI system, enabling autonomous operation, detection, and adjustment, thus improving efficiency.

[0024] In summary, the above-mentioned technical solutions provide a new prognostic evaluation strategy for DLBCL patients, laying a theoretical foundation for the individualized management of DLBCL patients, and therefore have good practical application value. Attached Figure Description

[0025] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0026] Figure 1 This is the critical roadmap in Embodiment 1 of the present invention.

[0027] Figure 2 A supervised sPL-SDA model was established based on the mass spectrometry data and clinical data obtained from the application experiment in Example 1 of this invention.

[0028] Figure 3 This is to further cross-validate the sPL-SDA model that uses permutation testing for data association in Embodiment 1 of the present invention.

[0029] Figure 4 This is a box plot of the quantitative analysis of 20 amino acids in a patient in Example 1 of the present invention.

[0030] Figure 5 The variable is the importance projection (VIP) value of the differentially expressed amino acids screened from plasma in Example 1 of this invention. VIP > 1 indicates that the variable is of great significance for the classification of categories in the model.

[0031] Figure 6 This is a heatmap showing the expression of amino acids in the plasma of DLBCL patients and normal controls in Example 1 of the present invention.

[0032] Figure 7 This is a comparison of plasma amino acid expression in DLBCL patients and normal controls in Example 1 of the present invention; wherein, (a) shows an increasing trend of plasma amino acid expression in DLBCL patients and normal controls, and (b) shows a decreasing trend of plasma amino acid expression in DLBCL patients and normal controls.

[0033] Figure 8 The hierarchical clustering analysis performed in Example 1 of this invention was used to observe the significant variation characteristics of different metabolites of the same type among samples.

[0034] Figure 9 The pathway analysis results for two amino acids that significantly affect prognostic grouping in Example 1 of this invention are shown in the pathway enrichment bubble diagram and the enrichment analysis diagram using the KEGG method.

[0035] Figure 10 This is a comparison chart of the actual and predicted values ​​of the PLSR model established using MATLAB in Embodiment 1 of the present invention. It shows the relationship between the standardized actual and predicted values, and the red dashed line represents the ideal fitting curve.

[0036] Figure 11This is a graph showing the difference (residual) between the predicted and actual values ​​of the PLSR model established using MATLAB in Embodiment 1 of the present invention. The blue dashed line represents the zero residual line.

[0037] Figure 12 The SVR machine learning model established using MATLAB in Embodiment 1 of this invention has a training set sample ratio of 75% (n=15), and the remaining samples are used as the test set, resulting in good model training parameters. Detailed Implementation

[0038] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0039] It should be noted that the terminology used herein is for descriptive purposes only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. Additionally, the molecular biology methods not detailed in the embodiments are conventional methods in the art; specific operations can be found in molecular biology guides or product manuals.

[0040] In a typical embodiment of the present invention, a prognostic biomarker for patients with diffuse large B-cell lymphoma is provided, the prognostic biomarker including key amino acids and clinically relevant indicators of DLBCL;

[0041] The amino acids include any one or more of alanine, arginine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, glycine, asparagine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.

[0042] Furthermore, the key amino acids are tryptophan and glutamine.

[0043] The clinically relevant indicators for DLBCL include any one or more of the following: IPI score at initial admission, IPI risk stratification, age, number of extranodal involvement sites, and lactate dehydrogenase (LDH).

[0044] The prognostic analysis included an analysis and assessment of progression-free survival (PFS) in patients with DLBCL.

[0045] In another specific embodiment of the present invention, the use of reagents for detecting the above-mentioned prognostic biomarkers in the preparation of prognostic analysis products for patients with diffuse large B-cell lymphoma is provided.

[0046] The reagents for detecting the aforementioned prognostic biomarkers include reagents for detecting key amino acids and reagents for detecting clinically relevant indicators of DLBCL, without specific limitations. In one specific embodiment of the present invention, the reagents for detecting key amino acids include, but are not limited to, reagents required for detecting amino acids based on selective reaction / multiple reaction monitoring (SRM / MRM) methods and / or LC-MS / MS mass spectrometry methods.

[0047] The prognostic analysis includes the analysis and assessment of PFS in DLBCL patients.

[0048] The specific product for prognostic analysis of patients with diffuse large B-cell lymphoma can be a test kit, a test device, or an equipment.

[0049] In another specific embodiment of the present invention, a prognostic analysis system for diffuse large B-cell lymphoma is provided, the prognostic analysis system comprising:

[0050] The acquisition unit is configured to acquire the aforementioned prognostic analysis biomarkers of the subjects;

[0051] An assessment unit is configured to predict progression-free survival of the patient with diffuse large B-cell lymphoma based on prognostic assessment biomarkers obtained by the acquisition unit.

[0052] The output unit is configured to output the progression-free survival prediction results based on the evaluation unit.

[0053] The assessment unit includes at least one prognostic assessment model for diffuse large B-cell lymphoma, specifically a PLSR model and / or an SVR model.

[0054] The PLSR model can be a MATLAB model. Specifically, using MATLAB, the dataset is randomly divided into training and test sets. The PLSR model is fitted using the training data, and the model is used to make predictions using the test dataset. The correlation coefficient of the predictions is calculated. Based on the validation results, the number of latent variables or other parameters are adjusted, and the training and validation steps are repeated to improve the model's predictive performance, thus obtaining the PLSR model. This model can be used to predict progression-free survival, and its formula is: PFS = 110.2092 + (-5.5576 * Gln) + (-6.7974 * Trp) + (0.1395 * age) + (3.1042 * number of exonucleates) + (0.0047 * LDH) + (-7.2143 * IPI hazard stratification) + (-2.8840 * IPI score).

[0055] In this invention, Gln, Trp, age, number of extranodal sites, and LDH are all raw data. Gln and Trp were analyzed using targeted metabolomics to directly detect their concentrations in serum, measured in μmol / L. Age refers to the patient's actual age. The number of extranodal sites is based on raw data; specifically, it is calculated directly from imaging examinations such as CT, ultrasound, and PET / CT. LDH is the actual value of the patient's lactate dehydrogenase in serum, measured in U / L. PFS is measured in months.

[0056] The International Prognostic Index (IPI) is a classic prognostic evaluation system for DLBCL patients. In this invention, the corresponding IPI score of the patient is obtained according to the IPI scoring criteria specified in the "Chinese Guidelines for the Treatment of Lymphoma (2021 Edition)".

[0057] Similarly, according to the "Chinese Guidelines for the Treatment of Lymphoma (2021 Edition)," the calculation method for IPI risk stratification is as follows:

[0058] Low-risk group 0~1 Low-to-medium risk group 2 Medium- and high-risk groups 3 High-risk group 4~5

[0059] Based on this, the present invention calculates the risk stratification of each patient and assigns them scores: low-risk group 1 point; low-to-intermediate-risk group 2 points; intermediate-to-high-risk group 3 points; high-risk group 4 points.

[0060] In this invention, the SVR model is obtained by training the model using an algorithm based on pre-collected results of prognostic biomarkers from patients with diffuse large B-cell lymphoma. Furthermore, this invention uses the `fitrsvm` function in MATLAB to construct the SVR model. Cross-validation is used to automatically adjust hyperparameters, thereby improving model performance. The SVR model is trained using the training dataset and used to make predictions on the test dataset. Predictive performance metrics are calculated, model performance is analyzed, parameters are adjusted, and repeated training and evaluation are performed to obtain optimal performance.

[0061] Furthermore, the prognostic analysis system can be run autonomously via Python.

[0062] In another specific embodiment of the present invention, a computer-readable storage medium is provided, on which a program is stored, which, when executed by a processor, implements the function of the diffuse large B-cell lymphoma prognostic analysis system.

[0063] In another specific embodiment of the present invention, an electronic device is provided, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to realize the function of the diffuse large B-cell lymphoma prognostic analysis system.

[0064] The present invention will be further illustrated below with specific examples. These examples are for illustrative purposes only and do not limit the scope of the invention. Experimental conditions not specifically specified in the examples are generally performed under conventional conditions or as recommended by the sales company; unless otherwise specified in the present invention, these conditions are commercially available.

[0065] Example

[0066] Materials and methods

[0067] 1.1 Research Subjects and Sample Collection Methods

[0068] Peripheral blood samples were collected from 20 newly diagnosed DLBCL patients and 10 healthy volunteers. All patients were newly diagnosed and had never received any treatment. The control group consisted of healthy volunteers with normal laboratory test results during the same period. Informed consent was obtained from the patients, and basic information such as gender, age, and general condition was recorded. Tumor immunohistochemical results, tumor stage, and the International Prognostic Index (IPI) were also recorded. Risk stratification was performed based on the collected clinical information. 2 ml of fasting venous blood was collected from all subjects in the morning and anticoagulated with EDTA. The blood samples were allowed to clot for approximately 10-20 minutes, then centrifuged at 2000-3000 rpm for 10 minutes. The supernatant (plasma) was stored at -80°C until analysis.

[0069] 1.2 Sample Information

[0070] Information on the samples to be tested: Two groups of blood samples were collected. One group consisted of a healthy control group with 10 biological replicates, and the other group consisted of a DLBCL group with 20 biological replicates (details are shown in Table 1).

[0071] Table 1 Sample Information

[0072]

[0073] Note: DLBCL represents the diseased group, and H represents the healthy control group.

[0074] 1.3 Sequencing methods for 20 amino acids in DLBCL

[0075] LC-MS / MS was used to perform targeted metabolomics analysis of plasma amino acid profiles in DLBCL patients and controls, and 20 common amino acids in plasma were quantified.

[0076] 1.3.1 Standard Curve

[0077] Take the standard, dilute it with water to prepare a series of standard working solutions, prepare the standard curve solution according to 1.3.2, and establish the standard curve.

[0078] 1.3.2 Metabolite Extraction

[0079] The plasma samples were thawed at room temperature and analyzed by LC-MS / MS. 100 μL of blood was taken and mixed with pre-cooled methanol (Merck, 144282) / acetonitrile (Merck, 1499230-935) (1:1, v / v), vortexed for 30 s, allowed to stand for 10 min, centrifuged at 12000 rpm at 4℃ for 10 min, and the supernatant was collected and dried under nitrogen. The supernatant was reconstituted in 1000 μL of water, filtered, and injected. All samples were mixed in equal volumes to prepare the quality control (QC) samples for LC-MS / MS analysis.

[0080] 1.3.3 Chromatography-Mass Spectrometry Analysis

[0081] 1) High Performance Liquid Chromatography Conditions

[0082] Samples were separated using a Nexera X2 LC-30AD high-performance liquid chromatography system (Shimadzu). A QC sample was placed at regular intervals within the sample cohort to test and evaluate the system's stability and repeatability. A mixture of amino acid metabolite standards was also included in the sample cohort for chromatographic retention time correction.

[0083] 2) LC-MS / MS mass spectrometry analysis based on MRM method

[0084] Mass spectrometry analysis was performed using a 5500QTRAP mass spectrometer (AB SCIEX) in positive ion mode. The 5500QTRAP ESI source conditions were as follows: source temperature 500℃; ion source gas 1 (GS1): 50; ion source gas 2 (GS2): 50; curtain gas (CUR): 35; ion spray voltage (IS): 5500V; analyte ion pairs were detected using MRM mode.

[0085] 1.3.4 Data Analysis

[0086] Analyst and SCIEX OS software were used to extract chromatographic peak areas and retention times. Retention times were corrected using amino acid standards for metabolite identification. All samples were mixed in equal volumes to prepare QC samples, and the stability and repeatability of the data were evaluated.

[0087] 1.4 Bioinformatics Analysis Methods for Amino Acid Sequencing Screening Results

[0088] Further bioinformatics analysis was performed on the amino acid sequencing results and collected clinically relevant information. Multivariate statistical analysis was employed to perform dimensionality reduction and regression analysis on the collected multidimensional data while preserving the original information to the greatest extent possible. Then, differentially expressed metabolites were screened and further analyzed. This mainly included principal component analysis (PCA) and sparse partial least squares discriminant analysis (sPLS-DA). The metabolites responsible for distinguishing the severe and mild DLBCL groups were those that showed statistical significance in multivariate analysis (VIP > 1.0 and P < 0.05).

[0089] 1.4.1 Sample Selection

[0090] Sample selection: 17 samples were used, including DLBCL2, 4-11, and 13-20.

[0091] The classification is based on IPI scores: 1-2 indicates mild cases, and 3-5 indicates severe cases.

[0092] 1.4.2 Spectrum-Effect Correlation Analysis

[0093] 1) Principal component analysis + sparse partial least squares regression analysis

[0094] Variable selection: Clinical + mass spectrometry variables

[0095] Method of application: This study used PCA+sPLS-DA method to perform spectrum-effect correlation analysis. First, PCA was used to preview the distribution of data for all sample points. Then, sPLS-DA was used to model and analyze the data to explore the mathematical relationship between amino acid profiles and clinical prognostic values.

[0096] 2) Variable Importance Plot

[0097] The VIP value, also known as the variable projection importance, represents the contribution of each metabolite ion to distinguishing between groups. The higher the VIP value, the stronger the correlation.

[0098] 1.4.3 Displacement Test

[0099] The sample labels (mild to severe) were randomly shuffled 100 times. The average Euclidean distance between the two groups of samples after clustering was calculated, and the distance was standardized using the distance from the real data as the denominator. The p-value is the number of times the average distance is greater than the real data under random conditions / 100.

[0100] 1.4.4 Cluster Heatmap Analysis

[0101] Hierarchical clustering was used to create a cluster heatmap, with colors ranging from blue to red representing the increasing levels of related amino acids.

[0102] 1.4.5 Pathway Enrichment Bubble Chart Analysis

[0103] The selected differential metabolites were annotated with KEGG and subjected to pathway analysis. Enrichment analysis was used to retrieve the key metabolic pathways mapped by the differential metabolites.

[0104] 1.4.6 Data Analysis

[0105] SPSS 18.0 statistical analysis software was used to analyze the experimental data. Quantitative data were expressed as mean ± standard deviation (Mean ± SD). Univariate analysis was performed using t-tests; multivariate analysis was performed using linear regression analysis with multiple amino acid groups exhibiting specific characteristics, followed by stepwise regression using the backward regression method. P < 0.05 was considered statistically significant.

[0106] 1.5 Establishment and Operation of Prognostic Model

[0107] 1.5.1 Variable selection for inclusion in the prognostic model

[0108] Incorporating tumor molecular characteristics and microenvironment into IPI can better guide prognostic analysis. Key clinical information and key amino acids from 20 patients were included as input variables in the prognostic model. Progression-free survival (PFS) was calculated as the output variable by reviewing patients' treatment records and conducting telephone follow-ups.

[0109] 1.5.2 Establishment and Operation of PLSR Prognostic Model

[0110] Using MATLAB, the dataset is randomly divided into training and test sets. A PLSR model is fitted using the training data, and predictions are made using the test dataset. The correlation coefficients of the predictions are calculated. Based on the validation results, the number of latent variables or other parameters are adjusted, and the training and validation steps are repeated to improve the model's predictive performance, thus obtaining the PLSR mathematical model. The trained PLSR model is saved in MATLAB and exported to a Python-readable format. The PLSR model is trained and saved using the scikit-learn library in Python and saved to a file. Finally, the model is integrated into a Flask API to host and use the PLSR model, enabling autonomous operation.

[0111] 1.5.3 Construction and Running of SVR Machine Learning Models

[0112] This paper uses MATLAB to divide the dataset into training and test sets, selects an appropriate kernel function, and constructs an SVR model using MATLAB's `fitrsvm` function. Cross-validation is used to automatically adjust hyperparameters, thereby improving model performance. The SVR model is trained using the training dataset and applied to the test dataset for prediction. Prediction performance metrics are calculated, model performance is analyzed, parameters are adjusted, and repeated training and evaluation are performed to obtain optimal performance. The SVR model is exported to a portable MATLAB file format. In the Python environment, the MATLAB Engine API for Python is used for integration, allowing the MATLAB model code to be run directly in Python. Scripts are written for data input processing, model invocation, and result output. Finally, logging and monitoring mechanisms are set up to monitor the model's operation and performance in real time.

[0113] Experimental results

[0114] 2.1 Quantitative results of plasma amino acids in patients with diffuse large B-cell lymphoma

[0115] Twenty amino acids were detected in the QC samples. The quantitative results of amino acids are shown in Table 2. The RSD of the analytes in the QC samples was <30%, indicating that our data are reliable and stable. We plotted box plots for patients and healthy controls based on the detailed quantitative data of amino acids in the samples. Figure 4 Clustering heatmaps of metabolites under positive ion mode are shown below. Figure 6 This helps us compare the distribution characteristics of quantitative amino acid data between patients and healthy controls. We can clearly see that the expression of eight amino acids in the plasma of DLBCL patients differs significantly from that of normal healthy controls. Specifically, Aspartic acid, Glutamic acid, Phe (phenylalanine), and Threonine show an increasing trend in DLBCL, while Glycine, Histidine, Serine, and Tryptophan show a decreasing trend. The trend graph is shown below. Figure 7 As shown, all differences were statistically significant (P < 0.05).

[0116] Table 2 Quantitative results of 20 amino acids

[0117]

[0118] 2.2 Using bioinformatics analysis to explore the potential relationship and network between amino acid expression and patient prognosis

[0119] 2.2.1 Establishing a model to analyze the relationship between amino acid expression and clinical prognosis

[0120] Using patients' amino acid expression and clinical characteristics as variables, a supervised sPLSDA model was established based on PCA, combining experimental mass spectrometry data with clinical data. The model showed that patients with differential amino acid expression and poor clinical prognostic factors also had poor prognoses. sPLSDA can force the regression coefficients of variables with minor effects to zero, which performed well in this dataset, indirectly indicating that many variables do not play a role in classification. The basic information and clinical data of the 20 patients are shown in Table 3.

[0121] Table 3 Basic Information of 20 Patients

[0122]

[0123] 2) The model using permutation tests for data association was further cross-validated. The results of the permutation tests showed a significant correlation between amino acid expression and patient prognosis (P<0.05), further illustrating that differences in amino acid expression in patients can affect their prognosis. We also found that the permuted R^2 and Q^2 values ​​were lower than the original left-hand points, indicating that the model did not overfit and had good stability and predictive ability.

[0124] 2.2.2 Amino acid analysis related to patient prognosis

[0125] We set variable importance projection (VIP) values, where VIP > 1 indicates that the variable is of significant importance to the classification of categories in the model. The differentially expressed amino acids screened from plasma were tryptophan and glutamine, which are significantly correlated with patient prognosis. We also performed hierarchical cluster analysis to observe the significant changes in the same type of differentially expressed metabolites among samples. The hierarchical clustering results showed that the differences in amino acid expression were more significant in intermediate- and high-risk patients (severe cases) than in intermediate- and low-risk patients (mild cases).

[0126] 2.2.3 Analysis of differential amino acid metabolic pathways related to patient prognosis

[0127] Pathway analysis was performed on two amino acids that significantly affected prognostic grouping. The resulting pathway enrichment bubble diagram is shown below. Figure 9Analysis revealed that three metabolic pathways were most likely associated with patient prognosis, and these three pathways had the highest number of enriched metabolites: the D-Glutamine and D-Glutamate Metabolism pathway, the Phenylalanine, tyrosine, and tryptophan biosynthesis pathway, and the Alanine, aspartate, and glutamate metabolism pathway. These pathways correspond to D-Glu and D-Glu metabolism, phenylalanine, tyrosine, and tryptophan biosynthesis, and alanine, aspartate, and glutamate metabolism, respectively. KEGG enrichment analysis was performed using this method.

[0128] 2.3 Construction and operation of the prognostic model

[0129] 2.3.1 Selection of prognostic model variables

[0130] Differentially identified amino acids in plasma, specifically tryptophan and glutamine, were shown to be significantly associated with prognosis and were therefore included as important independent variables (input variables) in the prognostic model. In the clinical data, IPI score, group, age, number of extranodal sites, and lactate dehydrogenase (LDH) were also included as independent variables (input variables) in the prognostic model. At the end of follow-up, we obtained the progression-free survival (PFS) of 20 patients. The final variable dataset is shown in Table 4.

[0131] Table 4 Prognostic model dataset

[0132]

[0133]

[0134] 2.3.2 PLSR Prognostic Model

[0135] A PLSR prognostic model was established using MATLAB. Through repeated training and validation steps, we obtained the prognostic expression for PFS: PFS = 110.2092 ± 5.5576 * Gln ± 6.7974 * Trp + 0.1395 * Age + 3.1042 * Number of extranodal sites + 0.0047 * LDH ± 7.2143 * IPI hazard stratification ± 2.8840 * IPI score. We validated the predictive ability of the PLSR prognostic model by comparing the actual and predicted values ​​(see the graph). Figure 10 The graph shows the relationship between standardized actual and predicted values, with the red dashed line representing the ideal fit curve. The points are relatively concentrated, mainly located near the red dashed line (ideal fit curve), indicating that the model generally predicts the output values ​​well. The residual plot further illustrates this. Figure 11The graph shows the difference between predicted and actual values ​​(residuals), with the blue dashed line representing the zero residual line. Most of the residuals are distributed close to zero, indicating that the overall error of the model is small. The residuals do not exhibit any obvious systematic bias or trend, suggesting that the model fit is not problematic. Individualized predictions for PFS were achieved through an attempt at autonomous running.

[0136] 2.3.3 SVR Machine Learning Prognostic Model

[0137] An SVR machine learning model was built using MATLAB, with the training set comprising 75% of the samples (n=15) and the remaining samples serving as the test set. This resulted in well-trained model parameters. After integration with an AI system, inputting the parameter features allows for prediction of PFS, providing prediction accuracy and confidence intervals. The results are as follows... Figure 12 As shown.

[0138] Example 2

[0139] This embodiment provides a prognostic analysis system for diffuse large B-cell lymphoma, the prognostic analysis system comprising:

[0140] The acquisition unit is configured to acquire the aforementioned prognostic analysis biomarkers of the subjects;

[0141] An assessment unit is configured to predict progression-free survival of the patient with diffuse large B-cell lymphoma based on prognostic assessment biomarkers obtained by the acquisition unit.

[0142] The output unit is configured to output the progression-free survival prediction results based on the evaluation unit.

[0143] The assessment unit includes at least one prognostic assessment model for diffuse large B-cell lymphoma, specifically the PLSR model and / or the SVR model.

[0144] The system can be operated in accordance with the method for prognostic analysis (PFS) of patients with diffuse large B-cell lymphoma as described in Embodiment 1 of the present invention.

[0145] Example 3

[0146] This embodiment provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, performs the steps of prognostic analysis (PFS) for patients with diffuse large B-cell lymphoma as described in Embodiment 1 of the present invention.

[0147] Example 4

[0148] This embodiment provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the steps of prognostic analysis (PFS) for patients with diffuse large B-cell lymphoma as described in Embodiment 1 of this invention.

[0149] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0150] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0151] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0152] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0153] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0154] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A prognostic biomarker for patients with diffuse large B-cell lymphoma, characterized in that, The prognostic biomarkers include key amino acids and clinically relevant indicators of DLBCL. The key amino acids are tryptophan and glutamine; The clinically relevant indicators for DLBCL are the patient's IPI score at first admission, IPI risk stratification, age, number of extranodal involvement sites, and LDH. Prognostic analysis is an analysis and assessment of PFS in DLBCL patients; Bioinformatics analysis of amino acid sequencing screening results was performed using principal component analysis and sparse partial least squares discriminant analysis, and prognostic models were established using PLSR and / or SVR models.

2. The use of the reagent for detecting the prognostic biomarker of claim 1 in the preparation of a prognostic analysis product for patients with diffuse large B-cell lymphoma.

3. The application as described in claim 2, characterized in that, The prognostic analysis includes the analysis and assessment of PFS in DLBCL patients; The product for prognostic analysis of patients with diffuse large B-cell lymphoma is a test kit, test device, or equipment.

4. A prognostic analysis system for diffuse large B-cell lymphoma, characterized in that, The prognostic analysis system includes: An acquisition unit, configured to: acquire the prognostic analysis biomarker of the subject as described in claim 1; An assessment unit is configured to predict progression-free survival of the patient with diffuse large B-cell lymphoma based on prognostic assessment biomarkers obtained by the acquisition unit. The output unit is configured to output the progression-free survival prediction result based on the evaluation unit. The assessment unit includes at least one prognostic assessment model for diffuse large B-cell lymphoma, specifically the PLSR model and / or the SVR model; The PLSR model is a MATLAB model, and its formula is PFS=110.2092 +(-5.5576*Gln) + (-6.7974*Trp) + (0.1395*age) + (3.1042*number of exonucleates) + (0.0047*LDH) + (-7.2143*IPI hazard stratification)+ (-2.8840*IPI score); The SVR model is obtained by training the model using an algorithm based on the pre-collected results of the prognostic biomarkers of the aforementioned diffuse large B-cell lymphoma patients. The SVR model is constructed using the fitrsvm function in MATLAB.

5. The diffuse large B-cell lymphoma prognosis analysis system according to claim 4, wherein, The prognostic analysis system operates autonomously using Python.

6. A computer-readable storage medium having a program stored thereon that, when executed by a processor, performs the functions of the diffuse large B-cell lymphoma prognostic analysis system as described in any one of claims 4-5.

7. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor, when executing the program, performs the functions of the diffuse large B-cell lymphoma prognostic analysis system as described in any one of claims 4-5.

Citation Information

Patent Citations

  • Molecule mark for diffuse large b-cell lymphoma treating guide and prognosis judgment

    CN101470112A