A method and system for constructing a vascular depression recognition model

By screening out plasma protein markers from plasma samples and using machine learning algorithms to establish a vascular depression recognition model, the misdiagnosis, misdiagnosis and subjective factors in identifying vascular depression in the prior art are solved, and efficient and quantitative recognition effects are achieved.

CN118016271BActive Publication Date: 2025-06-24ZHONGNAN HOSPITAL OF WUHAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410052222.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-06-24
Estimated Expiration
2044-01-12

AI Technical Summary

Technical Problem

The prior art relies on clinical scales and neuroimaging when identifying vascular depression, and there is a risk of misdiagnosis and misdiagnosis. The results are greatly affected by subjective factors and are difficult to quantitatively evaluate.

Method used

Plasma samples were collected, pretreated and LC-MS/MS detection were performed to screen differentially expressed proteins of plasma. The plasma protein markers were isolated from them using XGBoost machine learning algorithm and LASSO analysis, and finally, a vascular depression recognition model was established using Logistic regression analysis.

Benefits of technology

Effective identification of vascular depression is achieved, the risks of misdiagnosis and misdiagnosis are reduced, and the results are highly quantitative and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118016271B_ABST
    Figure CN118016271B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for constructing a vascular depression recognition model, including: collecting plasma samples of a person to be tested, preprocessing the plasma samples to obtain the results of the processed plasma samples; performing LC-MS / MS detection on the results of the processed plasma samples to obtain plasma protein quantitative information; screening the plasma protein quantitative information through fold change and T-test P value to obtain plasma differentially expressed proteins; using the XGBoost machine learning algorithm and LASSO analysis to isolate plasma protein markers from the plasma differentially expressed proteins; and adopting Logistic regression analysis to establish a vascular depression recognition model using the plasma protein markers. The vascular depression model construction method proposed by the present invention, which does not rely on clinical scales and neuroimaging, can effectively identify vascular depression according to a specific combination of plasma protein markers, and has good clinical application value and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence medical technology, and particularly relates to a method and system for constructing a vascular depression recognition model. Background Art

[0002] Vascular depression is a depressive syndrome in old age caused by cerebrovascular diseases or vascular risk factors. The main cause of vascular depression is the lesion of cerebral blood vessels, such as arteriosclerosis, thrombosis or cerebrovascular lesions. These lesions can lead to poor blood circulation in the brain and insufficient blood supply to the brain, thus affecting the normal transmission of neurotransmitters and the function of brain cells. The symptoms of vascular depression are similar to those of other types of depression, mainly including: decline in executive function, lack of insight, apathy, and psychomotor retardation. However, vascular depression may also be accompanied by other brain symptoms, such as headache, memory problems, and decline in cognitive function.

[0003] Currently, the methods for identifying vascular depression mainly include the following aspects: 1. Detailed medical history inquiry: Doctors will ask the patient about their medical history, including whether there have been previous cerebrovascular lesions, vascular risk factors (such as hypertension, diabetes, hyperlipidemia, smoking, etc.), and the occurrence and duration of depressive symptoms. 2. Physical examination: Doctors will conduct a physical examination, paying particular attention to the signs of cerebrovascular diseases, such as blood pressure, pulse, and heart auscultation. 3. Psychological assessment: Doctors may use some psychological assessment tools, such as the Self-Rating Depression Scale (SDS) or the Hamilton Depression Rating Scale (HAMD), etc., to evaluate the severity of the patient's depressive symptoms. 4. Neuroimaging examination: Brain neuroimaging examination is one of the important means for diagnosing vascular depression. Commonly used techniques include magnetic resonance imaging (MRI) and cerebral hemodynamics examination, etc., which can observe the changes in brain structure and function. 5. Blood tests: Doctors may conduct some blood tests, such as blood lipid, blood glucose, and coagulation function, etc., to exclude the possibility of other diseases.

[0004] The diagnostic criteria for vascular depression are still being continuously improved and updated. The current identification methods highly rely on psychological assessment scales and neuroimaging examinations. Neuropsychological evaluations must be conducted by well-trained professional doctors, otherwise there is a risk of misdiagnosis and missed diagnosis due to improper assessment. On the other hand, due to the subjective bias of the assessors, there are significant differences in the results of neuropsychological evaluations. Currently, it is difficult to accurately identify neuroimages visually and quantitatively. Therefore, so far, determining whether a cerebrovascular disease patient has vascular depression still depends on experienced clinicians to comprehensively evaluate through medical history collection, combined with clinical examinations, neuropsychological evaluations, laboratory test indicators, and auxiliary examination results such as structural images. The consistency is poor, greatly affected by subjective factors, and difficult to quantitatively evaluate. Summary of the Invention

[0005] The present invention provides a method and system for constructing a vascular depression identification model to solve the defects existing in the prior art.

[0006] In a first aspect, the present invention provides a method for constructing a vascular depression identification model, including:

[0007] Collect a plasma sample of a subject to be tested, preprocess the plasma sample to obtain the result of the processed plasma sample;

[0008] Perform LC-MS / MS detection on the result of the processed plasma sample to obtain plasma protein quantitative information;

[0009] Screen the plasma protein quantitative information through fold change and T-test P value to obtain plasma differentially expressed proteins;

[0010] Use the XGBoost machine learning algorithm and LASSO analysis to isolate plasma protein markers from the plasma differentially expressed proteins;

[0011] Adopt Logistic regression analysis and establish a vascular depression identification model using the plasma protein markers.

[0012] According to the method for constructing a vascular depression identification model provided by the present invention, collecting a plasma sample of a subject to be tested and preprocessing the plasma sample to obtain the result of the processed plasma sample includes:

[0013] Collect the plasma sample using a standard blood collection tube;

[0014] Successively perform protein denaturation, reduction, alkylation, enzymatic digestion, and desalting on the plasma sample to obtain the result of the processed plasma sample.

[0015] According to a method for constructing a vascular depression recognition model provided by the present invention, the processed plasma sample results are detected by LC-MS / MS to obtain plasma protein quantification information, including:

[0016] The data-independent acquisition (DIA) raw data file of the processed plasma sample results is separated and extracted by an LC-MS / MS detector;

[0017] The spectral library is predicted by a deep learning algorithm in DIA-NN software, and the DIA raw data file is extracted by the spectral library to obtain the plasma protein quantification information.

[0018] According to a method for constructing a vascular depression recognition model provided by the present invention, the plasma protein quantification information is screened by fold change and T-test P-value to obtain plasma differentially expressed proteins, including:

[0019] Taking the protein expression abundance greater than a preset fold change and the T-test P-value less than a preset value as the criteria, the plasma protein quantification information is screened to obtain the plasma differentially expressed proteins.

[0020] According to a method for constructing a vascular depression recognition model provided by the present invention, plasma protein markers are separated from the plasma differentially expressed proteins by using the XGBoost machine learning algorithm and LASSO analysis, including:

[0021] The XGBoost machine learning algorithm is used to calculate the importance scores and rankings of the plasma differentially expressed proteins in different grouped samples, and important rank-differentiated protein markers are screened and obtained;

[0022] The LASSO regression model is used to screen the important rank-differentiated protein markers to obtain the plasma protein markers.

[0023] According to a method for constructing a vascular depression recognition model provided by the present invention, logistic regression analysis is adopted, and a vascular depression recognition model is established by using the plasma protein markers, including:

[0024] The plasma protein markers are randomly combined, and a combined model is constructed by logistic regression analysis;

[0025] The optimal combination of the combined model is determined by the AUC value, and the vascular depression recognition model is output.

[0026] According to a method for constructing a vascular depression recognition model provided by the present invention, it further includes:

[0027] The vascular depression recognition model is verified and evaluated by receiver operating characteristic (ROC) curve analysis and confusion matrix.

[0028] In a second aspect, the present invention further provides a system for constructing a vascular depression recognition model, including:

[0029] A collection and preprocessing module, configured to collect plasma samples of a subject to be tested, preprocess the plasma samples, and obtain the results of the preprocessed plasma samples;

[0030] A detection module, configured to perform triple tandem liquid chromatography-mass spectrometry (LC-MS / MS) detection on the results of the preprocessed plasma samples to obtain plasma protein quantification information;

[0031] A screening module, configured to screen the plasma protein quantification information through the fold change and the T-test P-value to obtain differentially expressed plasma proteins;

[0032] A separation module, configured to use the XGBoost machine learning algorithm and the least absolute shrinkage and selection operator (LASSO) analysis to separate plasma protein markers from the differentially expressed plasma proteins;

[0033] An establishment module, configured to perform Logistic regression analysis and establish a vascular depression recognition model by using the plasma protein markers.

[0034] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for constructing a vascular depression recognition model as described in any one of the above is implemented.

[0035] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for constructing a vascular depression recognition model as described in any one of the above is implemented.

[0036] The method and system for constructing a vascular depression recognition model provided by the present invention can effectively identify vascular depression according to a specific combination of plasma protein markers through the proposed method for constructing a vascular depression model that does not rely on clinical scales and neuroimaging, and has good clinical application value and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a schematic flowchart of the method for constructing a vascular depression recognition model provided by the present invention;

[0039] Figure 2 It is a logical schematic diagram of the method for constructing a vascular depression identification model provided by the present invention;

[0040] Figure 3 It is a statistical bar graph of the number of proteins quantified by LC-MS / MS detection of all samples provided by the present invention;

[0041] Figure 4 The present invention provides a volcano plot for screening differentially expressed proteins between the diseased group and the non-disease group based on protein quantitative data according to the difference multiple and T test P value;

[0042] Figure 5 The present invention provides a bar chart of important differential proteins screened out after importance scores are calculated and sorted by the XGBoost algorithm;

[0043] Figure 6 It is a nomogram of the application mode of the protein combination diagnostic model provided by the present invention;

[0044] Figure 7 It is a schematic diagram of the receiver operating characteristic curve (ROC) analysis results of the protein combination diagnostic model provided by the present invention in the training set and the validation set;

[0045] Figure 8 It is a schematic diagram of the confusion matrix of the test protein combination diagnostic model in the validation set provided by the present invention;

[0046] Figure 9 It is a structural schematic diagram of a system for constructing a vascular depression identification model provided by the present invention;

[0047] Figure 10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] Figure 1 is a flow chart of a method for constructing a vascular depression identification model provided by an embodiment of the present invention, such as Figure 1 As shown, including:

[0050] Step 100: collecting a plasma sample from a subject to be tested, pre-treating the plasma sample, and obtaining a result of the treated plasma sample;

[0051] Step 200: Perform LC-MS / MS detection on the results of the processed plasma sample to obtain plasma protein quantification information;

[0052] Step 300: Screen the plasma protein quantification information through the fold change and T-test P-value to obtain plasma differentially expressed proteins;

[0053] Step 400: Use the XGBoost machine learning algorithm and LASSO analysis to isolate plasma protein markers from the plasma differentially expressed proteins;

[0054] Step 500: Perform Logistic regression analysis and establish a vascular depression recognition model using the plasma protein markers.

[0055] Specifically, as Figure 2 shown, in the embodiments of the present invention, first, plasma samples of subjects to be tested (including vascular depression patients and non-patients) are collected, and the plasma samples are preprocessed to obtain the results of the processed plasma samples; then, LC-MS / MS detection is performed on the results of the processed plasma samples to obtain plasma protein quantification information; further, the plasma protein quantification information is screened through the fold change and T-test P-value, and plasma differentially expressed proteins are output; then, the XGBoost machine learning algorithm and LASSO analysis are applied to isolate plasma protein markers from the plasma differentially expressed proteins; finally, Logistic regression analysis is performed, and a vascular depression recognition model is established through the plasma protein markers.

[0056] Through the proposed method for constructing a vascular depression model that does not rely on clinical scales and neuroimaging, the present invention can effectively identify vascular depression according to a specific combination of plasma protein markers, and has good clinical application value and promotion value.

[0057] On the basis of the above embodiments, step 100 includes:

[0058] Collect the plasma samples using a standard blood collection tube;

[0059] Perform protein denaturation, reduction, alkylation, enzymatic digestion, and desalting treatments on the plasma samples in sequence to obtain the results of the processed plasma samples.

[0060] Specifically, in order to make the experimental data true and effective, the blood samples used in the embodiments of the present invention are all from 71 patients in a large tertiary hospital in a certain city, and the patient information is shown in Table 1.

[0061] Table 1

[0062]

[0063]

[0064] The above samples need to be pretreated. Add the reaction solution (1% SDC / 100 mM Tris-HCL pH = 8.5 / 10 mM TCEP / 40 mM CAA) to the samples, and incubate at 60 °C for 1 h to perform protein denaturation, reduction and alkylation in one step. After dilution with an equal volume of ultrapure water, add trypsin according to the mass ratio of enzyme to protein of 1:50, and incubate and shake overnight at 37 °C for digestion. The next day, add TFA to terminate the digestion, centrifuge at 16000 g and take the supernatant for desalting with a self-made SDB desalting column. After drying by suction, store at -20 °C for later use, and the treated plasma samples are obtained.

[0065] Based on the above embodiments, step 200 includes:

[0066] Separate and extract the data-independent acquisition DIA raw data file of the results of the treated plasma samples by an LC-MS / MS detector;

[0067] Predict the spectral library through the deep learning algorithm in the DIA-NN software, and extract the DIA raw data file by the spectral library to obtain the plasma protein quantitative information.

[0068] It should be noted that the LC-MS / MS triple tandem liquid chromatography-mass spectrometry instrument is an analytical instrument used in the fields of chemistry, biology and pharmacy. The mass spectrometry data of the embodiments of the present invention are collected using a liquid chromatography-mass spectrometry system with a Q Exactive Plus mass spectrometer in series with an EASY-nLC 1200 liquid phase. The peptide segment sample is dissolved with the loading buffer, inhaled by the auto sampler and then bound to the analytical column (50 μm * 15 cm, C18, 2 μm, ) for separation. An analytical gradient is established using two mobile phases (mobile phase A: 0.1% formic acid and mobile phase B: 0.1% formic acid, 80% ACN). The flow rate of the liquid phase is set to 300 nL / min. The mass spectrometry collects data in DIA mode. Each cycle of scanning includes one MS1 scan (R = 70K, AGC = 3e6, MaxIT = 30 ms, scan range = 350 - 1250 m / z) and 30 variable window MS2 scans (R = 17.5K, AGC = 1e6, Max IT = 50 ms). The collision energy is set to 28.

[0069] The DIA raw data file was analyzed using DIA-NN software (v 1.8). The database used for retrieval was the Human proteome reference database in Uniprot (February 9, 2022, containing 20,375 protein sequences). A spectral library was predicted through the deep learning algorithm in DIA-NN, and the DIA raw data was extracted using the predicted spectral library and the spectral library obtained by the MBR function to obtain protein quantification information. The final results were screened at a 1% FDR for precursor ions and protein levels. The proteome quantification information after screening was used for subsequent analysis. As Figure 3 shown, 1632 - 2076 different proteins were quantified in each sample respectively.

[0070] Based on the above embodiments, step 300 includes:

[0071] Screening the plasma protein quantification information with the protein expression abundance greater than the preset fold change difference and the T-test P value less than the preset value to obtain the plasma differentially expressed proteins.

[0072] Specifically, in the embodiments of the present invention, the differentially expressed proteins between the diseased group and the non-diseased group were screened with the protein expression abundance change greater than 1.2-fold and the T-test P value less than 0.05 as the criteria. As Figure 4 shown, compared with the non-diseased group population, the expressions of 60 proteins in the diseased group were significantly different, 29 proteins were up-regulated in the diseased group, and 31 proteins were down-regulated in the diseased group.

[0073] Based on the above embodiments, step 400 includes:

[0074] Calculating the importance scores and rankings of the plasma differentially expressed proteins in different grouped samples using the XGBoost machine learning algorithm, and screening to obtain the differentially expressed protein markers of different importance levels;

[0075] Screening the differentially expressed protein markers of different importance levels using the LASSO regression model to obtain the plasma protein markers.

[0076] Specifically, in the embodiments of the present invention, the XGBoost machine learning model was used to calculate the importance scores and rankings of the above-mentioned differentially expressed proteins in distinguishing different grouped samples. As Figure 5 shown, according to the analysis results, 26 important differentially expressed proteins were screened out from the differentially expressed proteins.

[0077] Furthermore, the LASSO regression model was used to further screen out 20 most useful diagnostic protein features as plasma protein markers that can be used to distinguish the diseased group from the non-diseased group. The quantitative data of each protein marker between the diseased group and the non-diseased group had significant differences. Information of each protein marker is shown in Table 2.

[0078] Table 2

[0079] Protein ID Gene Name Protein Name P01023 A2M Alpha-2-macroglobulin Q8NBJ4 GOLM1 Golgi membrane protein 1 Q96CV9 OPTN Optineurin P02763 ORM1 Alpha-1-acid glycoprotein 1 Q14764 MVP Major vault protein Q6B0K9 HBM Hemoglobin subunit mu P09466 PAEP Glycodelin Q5T447 HECTD3 E3 ubiquitin-protein ligase HECTD3 O95969 SCGB1D2 Secretoglobin family 1D member 2 Q9C0B1 FTO Alpha-ketoglutarate-dependent dioxygenase FTO Q9GZM8 NDEL1 Nuclear distribution protein nudE-like 1 A0A0C4DH43 IGHV2-70D Immunoglobulin heavy variable 2-70D Q9NR34 MAN1C1 Mannosyl-oligosaccharide 1,2-alpha-mannosidase IC P05198 EIF2S1 Eukaryotic translation initiation factor 2subunit 1 Q9H3P7 ACBD3 Golgi resident protein GCP60 Q4G0F5 VPS26B Vacuolar protein sorting-associated protein 26B P24821 TNC Tenascin Q6IAA8 LAMTOR1 Ragulator complex protein LAMTOR1 P47914 RPL29 60S ribosomal protein L29 P98164 LRP2 Low-density lipoprotein receptor-related protein 2

[0080] Based on the above embodiments, step 500 includes:

[0081] Randomly combine the plasma protein markers and construct a combined model through Logistic regression analysis;

[0082] Use the AUC value to determine the optimal combination of the combined model and output the vascular depression recognition model.

[0083] Specifically, randomly combine the obtained protein markers, construct a combined model through Logistic regression, select the optimal protein combination (Q14764, P09466, Q9NR34, P05198, P24821) by the area under the curve (AUC) of the Receiver Operating Characteristic Curve (ROC) curve, as shown in Table 3, output the vascular depression recognition model, and establish a nomogram according to the model to show the application of the model in actual application recognition, such as Figure 6 shown.

[0084] Table 3

[0085]

[0086]

[0087] Based on the above embodiments, it further includes:

[0088] Verify and evaluate the vascular depression recognition model through Receiver Operating Characteristic (ROC) curve analysis and confusion matrix.

[0089] Optionally, the embodiments of the present invention also verify and evaluate the vascular depression recognition model, and verify and evaluate the constructed diagnostic model through Receiver Operating Characteristic (ROC) curve analysis, confusion matrix, etc., such as Figure 7 shown in the ROC analysis schematic diagram provided by the embodiments of the present invention, Figure 8 is the schematic diagram of the confusion matrix.

[0090] It can be understood that in actual applications, more samples and more protein combinations can be selected for modeling according to the modeling method of the embodiments of the present invention to increase the accuracy of the model.

[0091] The embodiment of the present invention still uses 18 patients from a large three-patient hospital in a certain city as a test set (9 patients in the disease group and 9 patients in the disease group), a combination of 5 plasma proteins (Q14764, P09466, Q9NR34, P05198, P24821) as a diagnostic model feature, and the patients in the aforementioned embodiment as a training set to examine the performance of the vascular depression recognition model.

[0092] The results are as follows Figure 7 As shown in the figure, the model has good discrimination, the AUC (95% CI) of the training set is 0.9145 (0.842-0.9871), and the AUC (95% CI) of the test set is 0.9136 (0.7412-1). The results show that the diseased group and the non-disease group can be distinguished well. The confusion matrix analysis results are shown in the figure. Figure 8 As shown in the figure, in the test set, the sensitivity of the diagnostic model reached 88.9% and the specificity was 66.7%, indicating that the model has good predictive ability.

[0093] The following is a description of a system for constructing a vascular depression recognition model provided by the present invention. The system for constructing a vascular depression recognition model described below and the method for constructing a vascular depression recognition model described above can refer to each other.

[0094] Figure 9 is a schematic diagram of the structure of a system for constructing a vascular depression recognition model provided by an embodiment of the present invention, such as Figure 9 As shown, it includes: a collection preprocessing module 91, a detection module 92, a screening module 93, a separation module 94 and a building module 95, wherein:

[0095] The collection and pretreatment module 91 is used to collect the plasma sample of the subject to be tested, pre-treat the plasma sample, and obtain the treated plasma sample result; the detection module 92 is used to perform triple tandem liquid mass spectrometer LC-MS / MS detection on the treated plasma sample result to obtain plasma protein quantitative information; the screening module 93 is used to screen the plasma protein quantitative information by difference multiple and T test P value to obtain plasma differentially expressed protein; the separation module 94 is used to separate the plasma differentially expressed protein from the plasma differentially expressed protein by using XGBoost machine learning algorithm and least absolute shrinkage and selection algorithm LASSO analysis to obtain plasma protein markers; the establishment module 95 is used to use Logistic regression analysis to establish a vascular depression recognition model using the plasma protein markers.

[0096] Figure 10 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 10As shown in the figure, the electronic device may include: a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040. Among them, the processor 1010, the communications interface 1020, and the memory 1030 complete communication with each other through the communication bus 1040. The processor 1010 may call the logical instructions in the memory 1030 to execute the method for constructing a vascular depression recognition model. The method includes: collecting a plasma sample of a person to be tested, preprocessing the plasma sample to obtain the result of the processed plasma sample; performing LC-MS / MS detection on the result of the processed plasma sample to obtain plasma protein quantitative information; screening the plasma protein quantitative information through the fold change and the T-test P value to obtain plasma differentially expressed proteins; using the XGBoost machine learning algorithm and LASSO analysis to isolate plasma protein markers from the plasma differentially expressed proteins; and performing Logistic regression analysis to establish a vascular depression recognition model using the plasma protein markers.

[0097] In addition, when the logical instructions in the foregoing memory 1030 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0098] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the method for constructing a vascular depression recognition model provided by the above-mentioned various methods. The method includes: collecting a plasma sample of a person to be tested, preprocessing the plasma sample to obtain a processed plasma sample result; performing LC-MS / MS detection on the processed plasma sample result to obtain plasma protein quantification information; screening the plasma protein quantification information through fold change and T-test P value to obtain plasma differentially expressed proteins; using the XGBoost machine learning algorithm and LASSO analysis to isolate plasma protein markers from the plasma differentially expressed proteins; and performing Logistic regression analysis to establish a vascular depression recognition model using the plasma protein markers.

[0099] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0100] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a vascular depression identification model, characterized in that: include: Collecting a plasma sample from a subject to be tested, pre-treating the plasma sample, and obtaining a result of the treated plasma sample; The treated plasma sample results are subjected to triple tandem liquid spectrometry LC-MS / MS detection to obtain quantitative information of plasma protein; The plasma protein quantitative information is screened by difference fold and T test P value to obtain plasma differentially expressed proteins; Separating plasma protein markers from the differentially expressed plasma proteins using XGBoost machine learning algorithm and least absolute shrinkage and selection algorithm LASSO analysis; Logistic regression analysis was used to establish a vascular depression recognition model using the plasma protein markers; Collecting a plasma sample from a subject to be tested, pre-treating the plasma sample, and obtaining a result of the treated plasma sample, including: The plasma sample is collected using a standard blood collection tube; The plasma sample is subjected to protein denaturation, reduction, alkylation, enzymatic hydrolysis and desalting treatment in sequence to obtain the result of the treated plasma sample; The treated plasma sample results are subjected to LC-MS / MS detection to obtain plasma protein quantitative information, including: The data of the treated plasma sample results are separated and extracted by the LC-MS / MS detector and the DIA raw data file is collected independently; The spectrum library is predicted by the deep learning algorithm in the DIA-NN software, and the DIA raw data file is extracted from the spectrum library to obtain the plasma protein quantitative information.

2. The method for constructing a vascular depression identification model according to claim 1, characterized in that: The plasma protein quantitative information is screened by difference multiples and T test P value to obtain plasma differentially expressed proteins, including: The plasma protein quantitative information is screened to obtain the plasma differentially expressed proteins based on the protein expression abundance being greater than a preset difference multiple and the T test P value being less than a preset value.

3. The method for constructing a vascular depression identification model according to claim 1, characterized in that: Plasma protein markers are separated from the differentially expressed plasma proteins using XGBoost machine learning algorithm and LASSO analysis, including: The XGBoost machine learning algorithm is used to calculate the importance scores and rankings of the plasma differentially expressed proteins in different group samples, and to screen and obtain important level differential protein markers; The LASSO regression model is used to screen the important grade difference protein markers to obtain the plasma protein markers.

4. The method for constructing a vascular depression identification model according to claim 1, characterized in that: Logistic regression analysis was performed to establish a vascular depression recognition model using the plasma protein markers, including: The plasma protein markers are randomly combined, and a combination model is constructed by Logistic regression analysis; The optimal combination of the combined models is determined using the AUC value, and the vascular depression recognition model is output.

5. The method for constructing a vascular depression identification model according to claim 1, characterized in that: Also includes: The vascular depression recognition model was verified and evaluated by receiver operating characteristic curve (ROC) analysis and confusion matrix.

6. A system for constructing a vascular depression identification model, characterized in that: include: The collection and preprocessing module is used to collect the plasma sample of the subject to be tested, preprocess the plasma sample, and obtain the result of the processed plasma sample; A detection module, used to perform triple tandem liquid spectrometer LC-MS / MS detection on the treated plasma sample results to obtain quantitative information of plasma protein; A screening module, used to screen the plasma protein quantitative information by difference fold and T test P value to obtain plasma differentially expressed proteins; A separation module, used for separating plasma protein markers from the plasma differentially expressed proteins using an XGBoost machine learning algorithm and a least absolute shrinkage and selection algorithm LASSO analysis; Establishing a module for establishing a vascular depression recognition model using the plasma protein markers by using Logistic regression analysis; The acquisition preprocessing module is specifically used for: The plasma sample is collected using a standard blood collection tube; The plasma sample is subjected to protein denaturation, reduction, alkylation, enzymatic hydrolysis and desalting treatment in sequence to obtain the result of the treated plasma sample; The detection module is specifically used for: The data of the treated plasma sample results are separated and extracted by the LC-MS / MS detector and the DIA raw data file is collected independently; The spectrum library is predicted by the deep learning algorithm in the DIA-NN software, and the DIA raw data file is extracted from the spectrum library to obtain the plasma protein quantitative information.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for constructing a vascular depression recognition model according to any one of claims 1 to 5 is implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a vascular depression recognition model according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method for researching proteome change of rat suffering from depression or anxiety based on proteomics

    CN111584009A

  • Biomarker related to depression as well as diagnostic product and application thereof

    CN113528643A

  • EEG signal depression identification system and method based on XGBoost algorithm

    CN117064389A