Serum protein fingerprint detection analysis method and system based on lasso regression algorithm for abdominal aortic aneurysm
By combining LASSO regression algorithm with mass spectrometry, differentially expressed proteins associated with abdominal aortic aneurysms were screened, and a risk assessment model was constructed. This solved the problem of the lack of effective biomarkers in existing technologies and enabled early detection and risk assessment with high sensitivity and specificity.
Patent Information
- Application Number
- CN202511687211.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-11-18
AI Technical Summary
Current technologies lack effective blood biomarkers for early screening, risk assessment, and rupture prediction of abdominal aortic aneurysms. They are highly dependent on imaging examinations and it is difficult to screen out highly sensitive and specific biomarker combinations from the complex serum proteome.
By employing the LASSO regression algorithm combined with mass spectrometry, serum samples were collected, preprocessed, detected by mass spectrometry, and analyzed to screen differentially expressed proteins associated with abdominal aortic aneurysms. A risk assessment model was then constructed for early detection and risk stratification.
It achieves highly sensitive and specific detection of abdominal aortic aneurysms, can accurately screen key protein combinations, construct risk assessment models, improve the accuracy of early detection and rupture risk prediction, and provide target resources for targeted therapy.
Smart Images

Figure CN121142060B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical detection technology, specifically to a method and system for detecting and analyzing serum protein fingerprints of abdominal aortic aneurysms based on the LASSO regression algorithm. Background Technology
[0002] Abdominal aortic aneurysm is a common aortic disease with an extremely high mortality rate upon rupture. Currently, clinical diagnosis and monitoring mainly rely on imaging examinations (such as ultrasound and CT), but effective blood biomarkers are lacking for early screening, risk assessment, and rupture prediction. Proteomics technology can detect changes in serum protein expression in high throughput, offering the possibility of discovering disease-related biomarkers. However, how to screen for highly sensitive and specific biomarker combinations from the complex serum proteome and establish effective risk assessment models remains a challenge. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for detecting and analyzing serum protein fingerprints of abdominal aortic aneurysms based on the LASSO regression algorithm. This method and system has high detection sensitivity and specificity and can be used for early detection, risk stratification, and rupture risk prediction of abdominal aortic aneurysms.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: a method for detecting and analyzing serum protein fingerprints of abdominal aortic aneurysms based on the LASSO regression algorithm, comprising the following steps:
[0005] a) Collect blood samples from the subjects and prepare serum;
[0006] b) Pretreatment of serum samples using lysis buffer and buffer solution;
[0007] c) Use a weak cation exchange chip to capture proteins in the pretreated sample, followed by rinsing and matrix crystallization.
[0008] d) The processed chip was detected using a MALDI-TOF mass spectrometer to obtain mass spectrometry data of serum proteins;
[0009] e) Analyze the mass spectrometry data: perform preprocessing, protein identification, outlier removal and differential expression analysis on the mass spectrometry data to screen out differentially expressed proteins associated with abdominal aortic aneurysm;
[0010] f) The LASSO regression algorithm was used to screen the differentially expressed proteins to obtain key protein combinations and their regression coefficients.
[0011] g) Detection and analysis were performed using an abdominal aortic aneurysm risk assessment model constructed based on the aforementioned key protein combination and its regression coefficients.
[0012] In one embodiment of the present invention, the lysis buffer in step b) is U9 lysis buffer, and the buffer solution is WCX2 buffer with a pH of approximately 4.0, with a dilution ratio of 1:40.
[0013] In one embodiment of the present invention, the weak cation exchange chip in step c) is a WCX2 chip, the rinsing solution is a 0.1% Triton X-100 solution, and the matrix solution is an aqueous solution of acetonitrile in α-cyano-4-hydroxycinnamic acid.
[0014] In one embodiment of the present invention, the mass spectrometry detection in step d) has a mass range of m / z 1000-20000, adopts linear mode and positive ion reflection mode, and the number of laser bombardments is not less than 2000 times / point.
[0015] In one embodiment of the present invention, step e) of preprocessing the mass spectrometry data, identifying proteins, removing outliers, and performing differential expression analysis to screen for differentially expressed proteins associated with abdominal aortic aneurysms further includes:
[0016] Data preprocessing: Convert the raw data into a peak list, perform peak filtering, handle missing values, and Z-score standardization;
[0017] Protein identification: Mass spectrometry data were identified using library search software. Parameters included: maximum allowable number of missed cleavage sites for trypsin digestion (2), peptide length range of 8-50 amino acids, fixed modification urea methylation, variable modification oxidation and acetylation, precursor ion mass tolerance ±20 ppm, daughter ion mass tolerance ±20 ppm, and false discovery rate (FDR) ≤1%.
[0018] Outlier detection: First, Z-score is used to standardize the original protein expression matrix data. Principal component analysis is performed on proteins with a detection rate >80%. Mahalanobis distance is used to identify and remove outlier samples.
[0019] Differential expression analysis: The Kruskal-Wallis test (KW test) was used to compare the protein peak intensity between groups, and the Benjamini-Hochberg method was used to control the false discovery rate. The screening criteria were |Log2FC|>1 and FDR<0.05.
[0020] In one embodiment of the present invention, the preprocessing of the mass spectrometry data in step e) further includes: before performing differential expression analysis, an outlier removal step based on principal component analysis and Mahalanobis distance is included.
[0021] For the initial mass spectrometry data in .raw file format, the msconvert tool of Proteo Wizard was used to perform format conversion and peak extraction by peak centering, and the output .mzML file and corresponding mass spectrometry peak list results were obtained; the peak list was generated with a threshold of SNR>5 and peak width of 2–20 Da, and low-quality peaks were removed.
[0022] The rolling ball algorithm is used to roll a "virtual ball" of radius r at the bottom of the spectral line, and the trajectory of the ball's center is recorded as the baseline function. The corrected signal is then obtained by subtracting the original signal from the baseline function. The window size of the rolling ball algorithm is set in the range of 100–500 m / z to effectively remove most of the slowly drifting baseline while maintaining the true peak shape without distortion.
[0023] Z-score standardization is used to normalize the data and eliminate systematic errors between samples, ensuring that the data conforms to a standard normal distribution. The formula is as follows:
[0024] ;
[0025] Where is the original peak intensity, is the sample error, and is the standard deviation.
[0026] In one embodiment of the present invention, step f) uses the LASSO regression algorithm to perform feature screening on the differentially expressed proteins to obtain key protein combinations and their regression coefficients, including the following steps:
[0027] K-fold cross-validation was used to optimize the regularization parameter λ to determine the optimal λ value in the LASSO regression algorithm. All samples were randomly divided into 10 similarly sized subsets. Ten iterations were performed. In each iteration, one fold subset was used as the validation set, and the 10-1 fold subset as the training set. On the training set, a series of different λ values were used to fit the LASSO model. Each fitted model was then applied to the validation set, and the mean squared error (MSE) between the predicted and actual values was calculated. For each λ value, there were 10 MSE values from the 10 different validation sets. The average MSE corresponding to this λ value was calculated. An outlier curve was obtained, showing the average MSE as a function of the λ value.
[0028] Feature selection using the LASSO regression algorithm was employed to retain key proteins in constructing a disease risk prediction model. The formula is as follows: ,in Here, x represents the expression level of the protein, and b is the intercept of the LASSO fitted curve.
[0029] In one embodiment of the present invention, the risk assessment model in step g) is a weighted protein risk score (PRS), which is calculated as follows: PRS=Σ(β_i*E_i), where β_i is the LASSO regression coefficient of the i-th key protein and E_i is the expression level of the protein.
[0030] A second aspect of the present invention provides an abdominal aortic aneurysm risk assessment system configured to perform the above-described method.
[0031] The third aspect of the present invention provides the application of the above method in the preparation of reagents or systems for assisting in the diagnosis of abdominal aortic aneurysms, assessing the risk of abdominal aortic aneurysm rupture, or guiding the stratified management of abdominal aortic aneurysm patients.
[0032] The beneficial effects of this invention are as follows:
[0033] 1. The detection and analysis method provided in this embodiment of the invention combines differential expression analysis, bioinformatics functional annotation and machine learning algorithm (LASSO regression), which can accurately screen out the key protein combinations most related to the occurrence, development and breakdown of AAA from massive proteome data.
[0034] 2. The detection and analysis method provided in this embodiment of the invention, based on the risk scoring model (PRS) constructed from key proteins, shows good discrimination, calibration and clinical applicability in both the training and validation sets, which is superior to the traditional assessment method that relies solely on aortic diameter.
[0035] 3. The key proteins screened by the detection and analysis method provided in the embodiments of the present invention also provide important target resources and theoretical basis for understanding the pathological mechanism of AAA, developing targeted therapeutic drugs and companion diagnostic reagents, and are particularly helpful for risk stratification and intervention decisions for patients with early small aneurysms (diameter <55mm).
[0036] 4. The detection and analysis method provided in this embodiment of the invention not only relies on conventional differential analysis (Kruskal-Wallis test + FDR correction), but also adds steps such as outlier removal, missing value stratification imputation, and Z-score standardization, making the results more robust. The detection and analysis method provided in this embodiment of the invention can more accurately extract key information highly correlated with the occurrence, development, and rupture of abdominal aortic aneurysms from massive mass spectrometry peak data, and reveal disease mechanism pathways (inflammatory chemotaxis, vascular remodeling, smooth muscle function, etc.), which is significantly different from existing single biomarker screening methods. Based on differentially expressed proteins, this invention further refines nine key protein biomarkers through LASSO feature screening, rather than remaining at the level of single or a small number of differentially expressed proteins.
[0037] 5. The serum protein fingerprinting detection and analysis method for abdominal aortic aneurysms based on the LASSO regression algorithm provided in this embodiment of the invention has high sensitivity and specificity. Through optimized sample preprocessing (U9 lysis buffer + WCX2 chip) and high-precision mass spectrometry detection (MALDI-TOFMS), it can efficiently capture and quantify low-abundance proteins with high detection sensitivity. Among them, the combination of U9 + WCX2 can improve the visibility of low-abundance signals without sacrificing the stability of the spectrum and reduce the masking of characteristic peaks by high-abundance carrier proteins, ultimately improving the effectiveness of differential peak screening and downstream LASSO feature selection. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of PCA outlier detection in an embodiment of the present invention;
[0039] Figure 2 This is a volcano diagram of differentially expressed proteins in an embodiment of the present invention;
[0040] Figure 3 The PCA score map based on differentially expressed proteins in this embodiment of the invention shows the separation of the normal group (NC), the unbroken AAA group (AAA), and the broken AAA group (rAAA).
[0041] Figure 4 This is a bar chart showing the GO-BP enrichment analysis of differentially expressed proteins in this embodiment of the invention.
[0042] Figure 5 This is a bubble diagram of KEGG pathway enrichment analysis of differentially expressed proteins in this embodiment of the invention.
[0043] Figure 6 This is a network diagram of protein-protein interaction (PPI) of differentially expressed proteins in an embodiment of the present invention;
[0044] Figure 7 This is a coefficient path diagram of LASSO regression screening of key proteins in an embodiment of the present invention (or a list / schematic diagram of the finally screened proteins and their β values).
[0045] Figure 8 The flowchart for detecting serum protein fingerprinting of aortic aneurysms;
[0046] Figure 9 A graph showing the λ selection for the LASSO-Cox model;
[0047] Figure 10 ROC plot;
[0048] Figure 11 For calibration curves;
[0049] Figure 12DCA curve;
[0050] Figure 13 Heatmap of differentially expressed proteins between abdominal aortic aneurysms (AAA, including ruptured and unruptured) and healthy controls (NC);
[0051] Figure 14 Volcano plot of differentially expressed proteins in abdominal aortic aneurysms (AAA, including ruptured and unruptured) relative to healthy controls (NC);
[0052] Figure 15 A heatmap showing the differentially expressed proteins between ruptured abdominal aortic aneurysm (rAAA) and "unruptured AAA + healthy control (AAA∪NC)";
[0053] Figure 16 Volcano plot of differentially expressed proteins in ruptured abdominal aortic aneurysm (rAAA) compared to "unruptured AAA + healthy control (AAA∪NC)". Detailed Implementation
[0054] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0055] Example
[0056] This embodiment provides a method for detecting and analyzing serum protein fingerprints of abdominal aortic aneurysms based on the LASSO regression algorithm, including the following steps:
[0057] 1. Experimental population and grouping
[0058] ①Inclusion and exclusion criteria:
[0059] The study participants were selected from clinical visitors and stratified according to the size of their abdominal aortic aneurysms: abdominal aortic diameter <30mm, abdominal aortic aneurysms 30-55mm, 55-75mm, and >75mm. The study population was followed up annually. In accordance with the Declaration of Helsinki, written consent and informed consent forms were obtained from each participant before the baseline survey and each follow-up survey.
[0060] Inclusion criteria: ① Patients aged 65 years and older; ② Patients with abdominal aortic diameter <30mm; ③ Patients with abdominal aortic aneurysm 30-55mm; ④ Patients with abdominal aortic aneurysm 55-75mm; ⑤ Patients with abdominal aortic aneurysm >75mm.
[0061] Exclusion criteria: ① Incomplete research data ② Failure to sign informed consent form ③ Suffering from other serious illnesses.
[0062] Subject grouping: AAA cohort (n=312; rAAA group (ruptured group) 48 cases, AAA group (unruptured group) 132 cases, normal control group 132 cases).
[0063] 2. Sample collection and testing
[0064] Collection: 5 ml of fresh blood was drawn from the patient on an empty stomach and allowed to coagulate naturally at room temperature for 3 hours. The blood was then centrifuged at 3000 rpm for 30 minutes at 4°C. The supernatant serum was immediately stored at -70°C. The serum used in the experiment underwent only one freeze-thaw cycle to avoid repeated freezing and thawing.
[0065] 3. Mass spectrometry detection of proteins:
[0066] Definition: This section describes further processing of serum samples before mass spectrometry analysis, based on upstream processing, to facilitate subsequent mass spectrometry analysis. The initial proteomics data file of serum proteins obtained using mass spectrometry is used for subsequent qualitative and quantitative protein analysis.
[0067] The detailed steps for mass spectrometry detection of proteins are as follows:
[0068] 1) Protein treatment: Serum samples were treated with U9 lysis buffer to disrupt cell structure and release proteins. The samples were then diluted about 40 times with WCX2 buffer (pH 4.0) to reduce non-specific binding interference.
[0069] The composition of the U9 pyrolysis solution is as follows:
[0070] Buffer base: 100mM sodium phosphate, pH adjusted to 6.0;
[0071] Salt ions: 500mM NaCl;
[0072] Descaling agent / surfactant: 0.1% Triton X-100;
[0073] Organic solvent: 10 acetonitrile;
[0074] pH range: pH 4.5–6.0.
[0075] The WCX2 buffer solution has the following specific composition: 50mM sodium acetate; the WCX2 buffer solution is prepared using high-purity water (HPLC grade) and the pH is adjusted to ~4.0 using acetic acid.
[0076] 2) Chip Processing: Using a weak cation exchange (WCX2) protein chip, rinse the chip surface with 0.1% Triton X-100 solution for 2 minutes to remove impurities. Slowly add diluted serum sample (1:40 diluted in WCX2 buffer) to the chip wells and incubate at room temperature for 1 hour to allow for complete protein binding. Rinse the chip three times with 0.1% Triton X-100 solution, then rinse once with ultrapure water to remove unbound proteins. Add α-cyano-4-hydroxycinnamic acid matrix solution (10 mg / mL, acetonitrile:water = 1:1) and dry at room temperature for 10 minutes to form a co-crystallization.
[0077] 3) Mass spectrometry analysis: The instrument used was a Bruker Autoflex Speed MALDI-TOFMS (linear mode). The ion source was in positive ion reflection mode, with an accelerating voltage of 20 kV and a reflecting voltage of 23 kV. The mass range was 1000-20000 m / z. The laser parameters were a nitrogen laser (337 nm), 1000 Hz, and 70% laser energy. Each sample point was bombarded 2000 times, and the average mass spectrum was taken. Each sample was analyzed three times. The raw data was exported in .mzML format, including mass-to-charge ratio (m / z) and peak intensity information.
[0078] 4. Data Analysis
[0079] Definition: This section of the serum proteome data analysis is based on mzML mass spectrometry files, using library search software for qualitative and quantitative analysis. Subsequent data results were processed using methods such as missing value handling, standardization, and differential protein screening to obtain differentially expressed proteins from different groups, allowing us to select the proteins we need.
[0080] The detailed steps of data analysis are as follows:
[0081] Mass spectrometry data preprocessing: msconvert was used to convert .mzML files into peak lists. The acquired mass spectrometry data underwent format conversion, noise filtering, baseline correction, and peak extraction, extracting m / z values, peak intensities, and retention times. MALDI-TOFMS (m / z 1000–20000) was exported as *.mzML. A peak list was generated with thresholds of SNR > 5 and peak width 2–20 Da, removing low-quality peaks.
[0082] Raw data format conversion: The mass spectrometry output files are in the raw .raw file format. Using ProteoWizard's msconvert tool, peak centering is employed for format conversion and peak extraction. The output is a .mzML file and the corresponding mass spectrometry peak list (peaklist, text format, including m / z, intensity, and retention time).
[0083] SNR, or Signal-to-Noise Ratio, = Peak Intensity / Local Background Noise. Local background noise is defined as the non-peak region within ±50 m / z. The SNR threshold is set to 5, filtering out peaks with SNR < 5. Peaks with a width between 2-20 Da are retained, while excessively wide or narrow noise peaks are excluded.
[0084] Baseline correction is performed using the rolling ball algorithm. Specifically, a "virtual ball" of radius r is rolled at the bottom of the spectral line, and the trajectory of the ball's center is recorded as the baseline function. Then use the original signal minus Obtain the corrected signal
[0085] For the rolling ball algorithm, the window size (equivalent to the ball radius) is generally set in the range of 100–500 m / z, taking into account both the baseline drift amplitude and the preservation of peak details. The preferred value is about 300 m / z, which can effectively remove most of the slowly drifting baselines while maintaining the true peak shape without distortion.
[0086] For polynomial fitting, polynomials of order 2–5 are usually selected. When the order is too low, it is difficult to adapt to complex backgrounds, while when the order is too high, it may lead to overfitting and false peaks. Therefore, this invention prefers to use third-order fitting.
[0087] In practical processing, parameters can be optimized based on residual distribution and signal-to-noise ratio to ensure that the corrected baseline is close to zero and the spectral peaks are fully preserved.
[0088] After baseline correction, further noise smoothing is required for the remaining signal. To reduce random noise in the mass spectrometry data while preserving peak shape characteristics, this invention employs a Savitzky-Golay filter for smoothing. This method maintains peak symmetry and half-maximum width through local polynomial fitting, avoiding peak broadening that may be caused by conventional moving averages.
[0089] The specific parameters are as follows:
[0090] A third-order Savitzky–Golay filter was used, with a window size of 15 data points.
[0091] The parameter selection is based on the following: a window that is too small will result in an insignificant smoothing effect and still high noise; a window that is too large will cause peak blunting and loss of detail. Verification showed that a 3rd-order polynomial + 15-point window achieves the best smoothing effect for the mass spectrometry data (m / z 1000–20000, peak width approximately 2–20 Da) of this invention: peak height and half-maximum width remain stable; background noise is significantly reduced, and the signal-to-noise ratio is improved by approximately 20%–30%.
[0092] Specifically, in this embodiment, after mass spectrometry data preprocessing, the analysis steps are as follows:
[0093] 1) Mass spectrometry identification (protein identification): Using MSfragger software in FragPipe, the mass spectrometry data files were searched based on the FASTA file of the human proteome (UniProt Human Proteome FASTA database), which contains 20,408 records. Search criteria were: trypsin digestion specificity, with two deletions; fixed cysteine residues with carbamoyl methylation modification; variable methionine oxidation modification and protein N-terminal acetylation modification; precursor ion mass tolerance range of -20 to +20 ppm for each sample (FragPipe default settings); daughter ion mass tolerance of ±20 ppm; minimum peptide length of 8. The obtained peptides were screened, and the FDR (Free Detection Ratio) was 1% at both the peptide and protein levels.
[0094] 2) Post-identification data processing: If a peptide / protein is missing in >50% of the samples, the peak is directly deleted to avoid introducing noise. For the remaining missing values, fill them with the minimum intensity value of the peak in all samples, minimizing the impact of data filling while ensuring data integrity. Calculate the coefficient of variation (CV) between replicates of different groups to ensure data consistency.
[0095] (A) Definition and imputation strategy for missing values
[0096] To balance the left-censoring characteristics of proteomics data with the robustness of inter-group comparisons, this invention employs a hierarchical, interpretable deletion processing procedure:
[0097] Definition: If a protein is missing in >50% of the samples, it is judged as having low detectability and is directly removed to avoid introducing systematic bias through filling; if the detection rate is between 50% and 80%, it is retained and enters the stratified filling process.
[0098] If the loss of this protein is mainly concentrated in low abundance samples or specific groups, it is inferred to be MNAR / left censoring;
[0099] If the deletions are approximately evenly distributed across groups, they are considered MCAR / MAR.
[0100] Filling strategy and basis:
[0101] MNAR (Left Censorization): This method employs semi-minimum / LOD imputation (per-protein, per-group; using 50% of the lowest non-zero intensity of the non-deleted signal within the group as the imputation value), or QRILC (Quantile Regression Imputation of Left-Censored data) to perform quantile regression sampling imputation on the log2 intensity. This approach aligns with the assumption that low-abundance signals are truncated by the detection limit, preserving the directionality of inter-group differences.
[0102] MCAR / MAR: Uses KNN interpolation (k=5, based on sample similarity) to limit to nearest neighbors within the same group to reduce oversmoothing.
[0103] All filling was performed after log2 transformation; then Z-score normalization (cross-sample centering and unit variance scaling) was performed on each protein to eliminate the effects of batch and loading drift.
[0104] Standardization beforehand: Z-score standardization (cross-sample centering and unit variance scaling) is then performed on the peak intensity of each protein to make the data distribution of different groups of samples consistent and eliminate systematic bias.
[0105] ;
[0106] in For original strength, The sample mean. The standard deviation is denoted as .
[0107] The protein expression matrix was processed using Z-score normalization to ensure that the Principle Component Analysis (PCA) was unaffected by dimensions.
[0108] (B) Robustness assessment: The differential analysis was rerun using the three strategies of "semi-minimum / QRILC / KNN" respectively. The overlap rate and discriminant performance (PCA / AUC) of the differential protein sets were compared. The strategy with the highest stability was selected as the default master scheme and the consistency results were given in the figure.
[0109] 3) Outlier removal:
[0110] Principal component analysis (PCA) was performed on all proteins with a detection rate >80%, and outliers were identified using Mahalanobis distance (threshold >3×SD). Figure 1 (Marked in red in the image), and then removed before proceeding to the differential analysis. The threshold is T, where T = μ + 3σ, and μ is the mean of the detection signals based on the training sample set, and σ is the standard deviation of the detection signals based on the training sample set.
[0111] The criteria for setting the outlier removal threshold are as follows:
[0112] PCA is performed on the retained proteins (detection rate >80%), and the top d principal components that can explain ≥70%–80% of the cumulative variance are selected (d is usually 2–3).
[0113] The squared Mahalanobis distance is calculated using the Minimum Covariance Determinant (MCD) robust covariance estimation center and covariance matrix. .
[0114] The Mahalanobis distance of each data point is compared with a threshold to identify outliers and remove them.
[0115] After removing outliers, with the threshold and d unchanged, repeat PCA and differential analysis, and annotate the sample IDs and grouping ellipses (95% confidence region) before and after removal in the attached figure. Figure 1 / Figure 3 ).
[0116] Specifically, in this embodiment, all proteins with reliable quantitative analysis (detection rate > 80%) in the protein identification results are reduced to the top three principal components (PC1, PC2, PC3) using principal component analysis. Samples deviating from the population distribution are identified using Mahalanobis distance statistics (threshold T, T = μ + 3σ, where μ is the mean of the detection signals based on the training sample set, and σ is the standard deviation of the detection signals based on the training sample set). These are then identified as outliers, and the corresponding proteins are selected as candidate target proteins.
[0117] Specifically, before using PCA, the raw data were standardized (Z-score standardization mentioned above) to ensure that the PCA results were not affected by the dimensions of the variables. Principal component analysis was performed using ScreenPlot. PCA was performed on proteins with a detection rate >80% (cumulative variance ≥70%). The center and covariance were estimated using MCD robust covariance. Mahalanobis distance was calculated, and outliers were identified and removed if the threshold was >3×SD. After fixing the threshold and dimensions, PCA / difference analysis was repeated, and the IDs before / after removal and the 95% confidence ellipse were marked in the figure captions.
[0118] The Mahalanobis distance is calculated as follows, and the covariance is estimated using the MCD robust estimation.
[0119]
[0120] PCA and Mahalanobis distance are suitable for high-dimensional data and can help discover global structural anomalies. In contrast, DBSCAN is suitable for clustering anomalies but is sensitive to high-dimensional sparse data; IsolationForest is suitable for local anomalies but has poor interpretability.
[0121] 4) Differential expression analysis:
[0122] In this embodiment, PCA analysis (PC1 and PC2) was performed on all proteins with a detection rate >80%, and outliers were identified using Mahalanobis distance (threshold >3×SD). Figure 1(The items marked with red circles in the text) were removed and then entered into the differential analysis.
[0123] The Kruskal-Wallis test was performed using the `kruskal.test()` function in R, with default parameters selected. The protein peak intensities of the normal group, the unruptured abdominal aortic aneurysm group (unruptured AAA group), and the ruptured abdominal aortic aneurysm group (ruptured AAA group) were compared. P-values were calculated to preliminarily screen for differentially expressed proteins. The final set of differentially expressed proteins satisfying |Log2FC|>1 and FDR<0.05 was obtained. Figure 2 Based on this set, perform PCA dimensionality reduction again ( Figure 3 The normal group (NC), the unbroken AAA group (AAA), and the broken AAA group (rAAA) formed a clear separation in the first two principal component spaces, with 95% confidence that there was no overlap between the ellipses, verifying the significant distinguishing ability of differentially expressed proteins in the grouping.
[0124] Specifically, in this embodiment, the p-value correction is selected using Benjamini-Hochberg (BH). The Benjamini-Hochberg method is used to control the false discovery rate (FDR), and difference peaks with FDR < 0.05 are screened to reduce false positive results.
[0125] Sort the p-values from smallest to largest and calculate the BH critical value for each p-value. (m = total test number), find the maximum rank number where P < critical value, and all proteins before that are considered significant. Here, FDR < 0.05 is set as significant. If there is a zero value, fill it with 0. The criteria for significant differences are defined as |Log2FC|>1 and P.adj<0.05, resulting in a set of significantly differentially expressed proteins. If Log2FC>0, the proteins are upregulated, and if Log2FC<0, the proteins are downregulated.
[0126] In this embodiment, after Kruskal-Wallis test (KW test) and Benjamini-Hochberg correction, the final set of differentially expressed proteins satisfying |Log2FC|>1 and FDR<0.05 was obtained. Figure 2 Based on this set, perform PCA dimensionality reduction again ( Figure 3 The normal group (NC group), the unbroken AAA group (AAA group), and the broken AAA group (rAAA group) formed a clear separation in the first two principal component spaces, with no overlap between the ellipses with a 95% confidence level, verifying the significant distinguishing ability of differentially expressed proteins in the grouping.
[0127] Outputs include volcano plots, PCA clustering validation, GO / KEGG enrichment (including FDR and rich factor), and STRING-PPI (threshold, version, and centrality). The resulting set of significantly differentially expressed proteins is used for further downstream biological significance analysis.
[0128] This invention uses principal component analysis (PCA) to visualize samples in the PC1–PC2 plane based on the preprocessed protein expression matrix (deletion-level imputation + Z-score normalization). Figure 1 The three groups (normal, unruptured AAA, and ruptured AAA) are clearly distributed in two-dimensional space: samples in the same group form a relatively tight cluster distribution in the figure, and there is little overlap in the distribution between groups, suggesting that the protein expression profile can effectively distinguish different pathological states.
[0129] Table 1. Names of differentially expressed proteins, fold changes, and statistical power.
[0130] Protein Nonparametric Test (K-W) P-Value KW False Discovery Rate Fold Change (log2) ace2 5.68E-05 0.001103 2.61749 adam8 0.004327 0.030708 6.21079 adamts16 0.000911 0.009203 5.756253 adgrg1 6.48E-07 3.93E-05 2.049302 aldh3a1 0.004503 0.031429 4.618792 alpp 1.45E-05 0.00048 1.340802 ambp 0.000176 0.002481 4.086789 angpt2 2.31E-05 0.00061 4.246026 anxa10 0.004434 0.031164 3.900751 asgr1 2.83E-05 0.000711 3.920276 bcl2l11 0.004581 0.031743 3.542862 c1qtnf1 0.00482 0.032618 5.070918 calca 1.47E-06 8.25E-05 1.351071 capg 1.51E-05 0.000481 2.076731 ccdc80 0.000124 0.001954 6.142952 ccl16 0.002348 0.018876 2.224985 ccl22 0.005026 0.033857 2.295475 ccl23 0.001364 0.012726 2.93235 ccl27 0.002213 0.018088 2.545395 ccl3 0.004275 0.030494 3.116523 ccn5 1.80E-05 0.000534 2.684077 cd27 0.00024 0.003118 3.685923 cd274 0.000204 0.002694 2.974818 cd302 0.000574 0.00614 4.226989 cd38 9.97E-05 0.001662 3.266719 cd4 2.21E-05 0.000595 3.377804 cd59 0.001209 0.011421 4.72162 cd83 0.000291 0.003679 2.685932 cdh1 0.004409 0.031141 3.773965 cdh15 3.10E-05 0.000746 2.637549 cdhr2 1.76E-05 0.000534 1.866431 chrdl1 6.94E-06 0.000266 3.180482 clc 8.68E-05 0.001541 2.124685 clec4d 0.002305 0.01863 3.886812 clec5a 0.003915 0.028293 3.296176 clec6a 0.002272 0.018471 3.299013 col18a1 1.85E-06 9.98E-05 3.36437 colec12 0.001819 0.01585 7.404825 crip2 9.56E-07 5.56E-05 3.477125 crisp2 0.002097 0.017282 4.385307 cst3 7.17E-08 8.69E-06 3.996519 ctsd 2.79E-05 0.000711 4.865606 ctsl 7.20E-05 0.001326 3.102825 ctsz 0.002733 0.021493 8.532291 cxcl16 0.000846 0.008604 4.930116 cxcl17 1.86E-12 9.01E-10 2.333509 cxcl9 3.43E-05 0.00078 1.952875 dsc2 0.001196 0.011421 3.45956 eda2r 9.80E-11 2.85E-08 3.337871 efemp1 5.40E-05 0.001062 5.695181 epha1 0.001719 0.01534 8.532291 fgfr2 0.006855 0.045134 3.956773 fut3_fut5 0.00062 0.006581 2.948965 gcnt1 0.00105 0.010258 3.959015 gfra1 0.0035 0.025591 5.853001 havcr1 0.000127 0.001954 1.783137 havcr2 0.001705 0.015313 4.333487 hgf 2.10E-05 0.000588 3.681289 hspb6 2.23E-07 1.91E-05 4.063687 hspg2 0.001302 0.012226 3.787847 icam5 0.001957 0.016648 2.463541 ifnl1 6.97E-05 0.0013 3.166166 igfbp4 3.21E-06 0.000146 2.3212 igfbp6 5.18E-05 0.001049 3.405054 il15 0.000299 0.003714 3.751446 il16 0.005419 0.036337 5.636589 il18bp 0.000448 0.004902 3.184411 il19 7.37E-05 0.001341 3.45956 il2ra 0.000726 0.00762 2.98392 il4r 0.001002 0.009948 8.860777 il5ra 0.000347 0.004109 2.899684 il6 9.67E-08 1.08E-05 1.936215 klk10 0.003384 0.025247 4.218907 klk13 0.003475 0.025537 3.091713 klk4 4.07E-05 0.000884 2.251532 krt19 0.000105 0.001714 5.268066 lair1 0.000443 0.004902 4.636625 lamp3 2.17E-10 5.11E-08 1.311861 layn 0.00281 0.021903 3.179175 lcn2 8.20E-05 0.001474 4.691485 lgals4 0.0002 0.002694 4.925725 ltbp2 0.000327 0.003937 3.397433 ly96 0.000296 0.003706 5.237809 mad1l1 0.001796 0.015742 3.239212 matn3 0.000339 0.004046 3.822126 mb 2.82E-05 0.000711 2.780899 mfap5 0.004709 0.032259 6.519396 msln 1.24E-07 1.29E-05 2.437518 mzb1 0.000101 0.001662 2.954546 nbl1 0.004608 0.031778 4.446903 nefl 5.19E-05 0.001049 2.370333 nfasc 0.000162 0.002331 4.652865 nos3 0.000205 0.002694 2.897533 nppb 0.002102 0.017282 1.189019 ntprobnp 0.000728 0.00762 1.665885 pgf 2.81E-10 5.11E-08 2.955105 pi3 3.13E-05 0.000746 4.208201 pigr 0.000206 0.002694 3.417332 pik3ip1 0.003435 0.025371 8.702149 plat 3.14E-06 0.000146 2.139229 prok1 0.000402 0.004497 1.955111 prss8 3.12E-09 4.54E-07 2.51539 rbp2 0.002037 0.01703 2.85099 reg1a 0.000565 0.006093 5.974739 reg1b 0.005853 0.038885 2.295475 reg3a 0.001873 0.016107 2.579482 rspo1 1.89E-05 0.000541 2.792847 rspo3 0.001005 0.009948 4.22429 rtn4r 0.000304 0.003714 4.216223 s100a12 6.68E-05 0.001263 1.869855 scarb2 4.50E-05 0.000948 5.643784 pgf.1 2.81E-10 5.11E-08 2.955105 pi3.1 3.13E-05 0.000746 4.208201 pigr.1 0.000206 0.002694 3.417332 pik3ip1.1 0.003435 0.025371 8.702149 plat.1 3.14E-06 0.000146 2.139229 prok1.1 0.000402 0.004497 1.955111 prss8.1 3.12E-09 4.54E-07 2.51539 rbp2.1 0.002037 0.01703 2.85099 reg1a.1 0.000565 0.006093 5.974739 reg1b.1 0.005853 0.038885 2.295475 reg3a.1 0.001873 0.016107 2.579482 rspo1.1 1.89E-05 0.000541 2.792847 rspo3.1 0.001005 0.009948 4.22429 rtn4r.1 0.000304 0.003714 4.216223 s100a12.1 6.68E-05 0.001263 1.869855 scarb2.1 4.50E-05 0.000948 5.643784 trem2 0.003928 0.028293 3.925747 vsig4 0.002058 0.017108 3.488415 wfdc12 2.10E-06 0.000109 8.117387 wfdc2 6.50E-12 2.37E-09 2.125944
[0131] As shown in Table 1, this embodiment used a non-parametric Kruskal-Wallis (KW) test based on serum samples related to abdominal aortic aneurysms to analyze the differential expression levels of each detected protein among different groups. Table 1 lists the names of proteins that still showed statistical differences after multiple test correction, the corresponding KW test p-values, the KW false discovery rate (FDR), and the fold change expressed as log2. The fold change (log2) is the logarithm of the expression ratio of the target group to the control group, with a larger value indicating a higher degree of relative upregulation in the target group. As shown in Table 1, several proteins related to inflammation, vascular remodeling, and cardiovascular function, such as ACE2, CALCA, CXCL17, ANGPT2, IL-6, and LAMP3, showed significant differences among the groups, providing a basis for candidate biomarkers to subsequently construct a serum protein fingerprint of abdominal aortic aneurysms.
[0132] It should be noted that the K-WP value in Table 1 is the original P value obtained from the Kruskal-Wallis nonparametric test; the KW false discovery rate is the FDR corrected by multiple hypothesis testing; the fold difference (log2) is the logarithm of the ratio of expression levels of the target group to the control group, which is used to characterize the magnitude of protein upregulation or downregulation.
[0133] Table 2 - Top 25 differentially expressed proteins in ruptured abdominal aortic aneurysms (rAAA)
[0134] logFC (log2 fold change of protein expression levels) AveExpr (average expression level of proteins) T (t-statistic of the test) P.Value (original P-value) adj.P.Val (adjusted P-value for multiple hypothesis correction) B (Bayes' statistic) Significant (indicates whether a protein is significantly differentially expressed, up for up-regulated, down for down-regulated) Protein 1.225599 0.66514 6.566735 1.44E-09 1.81E-06 11.45137 up mmp12 1.406932 0.16968 5.557374 1.72E-07 0.000108 7.0051 up pten 0.472252 0.269219 5.118141 1.21E-06 0.000506 5.198683 up eda2r 0.719378 0.256603 4.860919 3.63E-06 0.001139 4.183539 up calca 0.741899 0.374501 4.665822 8.16E-06 0.00205 3.436362 up cxcl17 0.30079 0.148749 4.549381 1.31E-05 0.002097 3.000251 up phospho1 0.30079 0.148749 4.549381 1.31E-05 0.002097 3.000251 up vwc2 0.412949 0.227768 4.544269 1.34E-05 0.002097 2.981277 up prtg 0.644079 0.352919 4.515741 1.50E-05 0.002097 2.875662 up lamp3 0.271771 0.121284 4.277441 3.86E-05 0.004843 2.011738 up adm 0.595618 0.428894 4.176441 5.69E-05 0.006499 1.655711 up gdf15 -0.48457 -0.22032 -4.08015 8.20E-05 0.008588 1.322098 down il7r -0.4144 -0.20498 -4.02203 0.000102 0.009842 1.123525 down faslg 0.407697 0.133158 4.002602 0.00011 0.009842 1.057608 up mzb1 0.327667 0.128553 3.795845 0.000234 0.019562 0.371272 up ptk7 -0.23057 -0.08073 -3.73199 0.000293 0.021509 0.164962 down b4gat1 0.615194 0.083398 3.721766 0.000304 0.021509 0.132188 up Prss27 -0.46562 -0.1644 -3.7181 0.000308 0.021509 0.120452 down lpl -0.22877 -0.03493 -3.66093 0.000377 0.024922 -0.0614 down Prss2 0.38133 0.044988 3.634295 0.000414 0.025986 -0.14537 up nppc 0.255783 0.136272 3.616102 0.000441 0.026366 -0.20244 up cst3 0.455819 0.208492 3.584171 0.000492 0.02811 -0.30207 up krt19 -0.23485 -0.14489 -3.55278 0.000549 0.029956 -0.39932 down apom 0.316622 0.157107 3.523934 0.000605 0.030681 -0.48809 up angpt2
[0135] Based on the overall list of differentially expressed proteins, this embodiment further performed a linear model analysis on the differences between the ruptured abdominal aortic aneurysm (rAAA) group and the control group, and ranked the proteins according to the fold change and statistical significance. As shown in Table 2, the top 25 differentially expressed proteins were selected as the top 25 proteins of rAAA characteristics.
[0136] It should be noted that in Table 2, logFC represents the log2 fold change of the rAAA group relative to the control group; adj.P.Val is the p-value after multiple comparison correction; "up / down" indicates upregulation or downregulation in the rAAA group. The top 25 differentially expressed proteins listed in Table 2 are selected after being sorted by fold change and significance, and are the key differentially expressed proteins used to characterize rAAA.
[0137] The explanation regarding the top 25 differentially expressed proteins is as follows: Among all differentially expressed proteins that meet the criteria of "|Log2FC|>1 and FDR<0.05", they are first sorted according to the effect size |Log2FC| (and / or FDR), and then the top 25 proteins are selected as the "top 25 differentially expressed proteins" for subsequent functional enrichment analysis, mechanism exploration, or visualization. The selection of 25 proteins is based on a balance consideration of this dataset - it can cover the main signaling pathways and key biological information, while avoiding excessive protein count that could lead to model overfitting and interpretation difficulties.
[0138] In Table 2, logFC is the log2 fold change, representing the expression change of the rAAA group relative to the control group. logFC > 0 indicates upregulation in the rAAA group, and logFC < 0 indicates downregulation. AveExpr is the normalized mean expression level. t is the t-statistic for the difference test. P.Value is the original P-value. adj.P.Val is the P-value after multiple correction. B is the log-likelihood ratio of the difference. The "significant" column indicates the direction of the difference (up / down) and whether the preset significance standard is met. "protein" is the corresponding protein name.
[0139] As shown in Table 2, MMP12, PTEN, EDA2R, CALCA, CXCL17, LAMP3, and GDF15 were significantly upregulated in the rAAA group, while IL7R, FASLG, APOM, and LPL showed a downregulated trend. These proteins together constitute a characteristic protein combination of rAAA that is helpful in distinguishing between ruptured and unruptured states, and can be used to characterize the serum protein fingerprint of ruptured abdominal aortic aneurysms.
[0140] Table 3 - Top 25 differentially expressed proteins in abdominal aortic aneurysms (AAA)
[0141] logFC (log2 fold change in protein expression levels) T (t-value of statistical test) P.Value (Original P-value) adj.P.Val (adjusted p-value, used for correction in multiple hypothesis testing) B (Bayesian test statistic) Significant (marks whether a protein is differentially expressed; up indicates upregulation, down indicates downregulation) protein 1.19441957282848 13.5001899750957 1.82838017920702e-41 2.29644550508401e-38 82.5780702278968 up mmp12 0.683029746747323 8.30611276783789 1.01210007440243e-16 4.23732564483152e-14 27.2298252930171 up cxcl17 0.629412926803235 8.34543508551382 7.2649718614883e-17 4.23732564483152e-14 27.5501419091219 up gdf15 0.603488512076721 7.04774838250881 1.84061048998647e-12 5.77951693855752e-10 17.7763030334566 up lamp3 0.271837020010551 6.34214757407474 2.2841491108808e-10 5.73778256653258e-08 13.1522084700563 up pgf 0.714725835459387 5.95807066800537 2.56848070036862e-09 5.37668626610497e-07 10.8397125586262 up il6 0.41130732602333 5.86260918578916 4.58356572402621e-09 8.22422649910988e-07 10.2873233628033 up prss8 0.51734334439847 5.58610585011473 2.3336047876909e-08 3.66375951667472e-06 8.73761596313351 up eda2r 0.588458224715838 5.51179140961236 3.56858961568376e-08 4.98016506366533e-06 8.33385599054085 up adgrg1 -0.346414145 -5.466554961 4.60942113201509e-08 5.78943294181096e-06 8.09072499112663 down faslg 0.5534813056623 5.38424867034891 7.3058156673667e-08 8.34191316201143e-06 7.65348923376174 up wfdc2 0.987658515818959 5.34631435819182 9.01357505770954e-08 8.39829685549465e-06 7.4542013362116 up alpp 0.481463073922412 5.33945453728149 9.36115891536028e-08 8.39829685549465e-06 7.41831354123407 up msln -0.155378217 -5.348395203 8.91063284518141e-08 8.39829685549465e-06 7.46509657045737 down egfr 0.345968768018917 5.27931627666677 1.30179224350866e-07 1.09003403856458e-05 7.10566506142022 up hspb6 0.301962890521521 5.1952275483535 2.0522886072005e-07 1.61104655665239e-05 6.67443354118781 up ptgds 0.456849142643553 5.1765322954316 2.26875708331518e-07 1.67621111567286e-05 6.57949863904216 up lta4h 0.266691645207475 4.99760431280162 5.82324760692555e-07 4.06333277461027e-05 5.68819351494255 up cst3 0.809274118813176 4.9231282963432 8.54284831891213e-07 5.64727236239665e-05 5.32643271110556 up ntprobnp 0.241907581936417 4.8940683619063 9.90637149182485e-07 6.22120129686601e-05 5.1867483613261 up plaur 0.281340497219838 4.87173830757678 1.10941073911312e-06 6.63533280155273e-05 5.0799743283912 up angpt2 0.704649563315022 4.84507326200816 1.26923928942372e-06 7.24620248870994e-05 4.95311107036259 up mln 0.493488219986014 4.80021167577859 1.58931725921395e-06 8.67905425031617e-05 4.74124454526961 up wfdc12 0.286377725055568 4.73124439559578 2.23725222236326e-06 0.000117082866303678 4.41937522236932 up crip2
[0142] For the unruptured abdominal aortic aneurysm (AAA) group, this embodiment also compared the expression differences between the AAA group and the healthy control group based on the overall differentially expressed proteins. The proteins were then ranked comprehensively according to the fold change and statistical indicators, and the top 25 differentially expressed proteins were selected to describe the characteristic fingerprint of AAA. Table 3 lists the 25 proteins with the highest degree of difference in the AAA group.
[0143] The columns in Table 3 have the same meanings as those in Table 2: logFC is the log2 fold difference between the AAA group and the healthy control group, AveExpr is the average expression level, t is the test statistic, P.Value is the raw P-value, adj.P.Val is the corrected P-value, B is the log-likelihood ratio of the difference, "significant" indicates the direction of the difference and the significance marker, and "protein" is the protein name. The top 25 proteins were selected from the AAA-related differentially expressed proteins by a comprehensive ranking based on fold difference and statistical significance, and are used to represent the main characteristic protein signals of the AAA group.
[0144] As shown in Table 3, multiple proteins, including MMP12, CALCA, EDA2R, NOS3, S100A12, EPHA1, LAMP3, IL-6, CXCL17, IL5RA, PGF, and PLAT, were upregulated in the AAA group, while proteins such as UMOD, APOM, PLTP, and CNDP1 were downregulated. These top 25 proteins reflect the comprehensive changes in multiple pathways during the development of AAA, including inflammatory responses, lipid metabolism disorders, and alterations in vascular endothelial function. They provide key characteristic variables for constructing serum protein fingerprints of the AAA group and for subsequent model training.
[0145] Figure 13 This is a heatmap of differentially expressed proteins between abdominal aortic aneurysms (AAA, including ruptured and unruptured aneurysms) and healthy controls (NC). Each column represents a single subject sample (grouped by color bands: AAA—red, NC—blue), and each row represents a single protein. The color indicates the row-normalized expression level (Z-score) of the protein in the sample: red for high expression and blue for low expression, with the color range indicated by the right-hand bar (−3 to +3). The figure shows the top 25 differentially expressed proteins based on a combination of statistical significance and effect size. Hierarchical clustering was performed simultaneously in both rows and columns, and the cluster dendrogram shows that the protein set can distinguish AAA from NC as a whole (see the clustering branches at the top and left of the figure).
[0146] Figure 14This is a volcano plot showing the differentially expressed proteins in abdominal aortic aneurysms (AAA, including ruptured and unruptured aneurysms) relative to healthy controls (NC). The horizontal axis represents log2 (fold change) (log2FC, right side for upregulation, left side for downregulation), and the vertical axis represents **−log10(P-value / corrected q-value)** to reflect statistical significance. Red dots represent significantly upregulated proteins, blue dots represent significantly downregulated proteins, and gray dots represent proteins that did not reach the significance threshold. The labeled proteins are representative proteins with the highest degree of difference or significance (see the Statistical Methods section of the examples for specific thresholds and multiple correction methods).
[0147] Figure 15 A heatmap showing differentially expressed proteins between ruptured abdominal aortic aneurysms (rAAA) and "unruptured AAA + healthy controls (AAA∪NC)". The plotting method is similar to... Figure 13 Consistency: Columns represent samples (grouped by color bands: rAAA—red, AAA∪NC—blue), behavioral proteins, with color indicating row-wise Z-score (higher red, lower blue, −3 to +3). The top 25 most discriminative differentially expressed proteins are also shown. Bidirectional clustering reveals a clear separation between the overall expression pattern of rAAA and AAA∪NC in this protein set, suggesting molecular features related to fragmentation.
[0148] Figure 16 This is a volcano plot showing the differentially expressed proteins in ruptured abdominal aortic aneurysms (rAAA) compared to unruptured AAA + healthy controls (AAA∪NC). The coordinates and coloring have the same meaning. Figure 14 The x-axis represents log2FC and the y-axis represents −log10 (P or q). Red indicates proteins that are significantly upregulated in rAAA, blue indicates proteins that are significantly downregulated, and gray indicates proteins that are not significant. The labeled proteins are those with higher degree of difference / significance, which are candidate biomarkers that are more strongly associated with the rupture state.
[0149] 5) Exploration of differential protein mechanisms:
[0150] Based on the differentially expressed proteins, PCA analysis was also performed, and PCA score plots were generated, labeling the normal group, the unruptured abdominal aortic aneurysm group, and the ruptured abdominal aortic aneurysm group. The separation trend between groups was observed based on the differences between principal components (PCs) 1, 2, and 3 and the mean. The number of outliers in each group was found to be: normal group (24 people), unruptured abdominal aortic aneurysm group (1 person), and ruptured abdominal aortic aneurysm group (2 people), validating the discriminative effect of the differentially expressed proteins. Specifically, the differential protein discrimination validation steps are as follows:
[0151] An expression matrix was constructed based on the differentially expressed proteins obtained through screening. Proteins with a detection rate >80% were subjected to PCA, and the first three principal components were extracted to form a feature space. The Mahalanobis distance of each sample within this space was calculated, and an outlier threshold T = μ + 3σ was set based on the mean μ and standard deviation σ of the Mahalanobis distances in the training set. A sample was considered an outlier when its Mahalanobis distance was greater than T. Statistical results showed that the number of outliers in the normal group, the unruptured abdominal aortic aneurysm group, and the ruptured abdominal aortic aneurysm group were 24, 1, and 2, respectively. These results indicate that after removing extreme outliers, the differentially expressed proteins can effectively distinguish the three groups of subjects, providing a stable feature base for subsequent model construction.
[0152] Table 4. Protein names, functions, and expression levels (m / z and intensity) in the normal and pathological groups (unruptured and ruptured).
[0153] Protein name Function m / z Expression level (normal group) Expression level (abdominal aortic aneurysm group) Expression level (ruptured group) cxcl17 Mucosal chemokines recruit monocytes / macrophages and participate in the regulation of immune inflammation and angiogenesis. 600" 0.81 1.13 1.63 pgf Placental growth factor, VEGF family, promotes angiogenesis / permeability and vascular wall remodeling ~26,000* 0.93 1.21 1.56 lamp3 Dendritic cell-associated lysosomal membrane proteins, antigen processing and migration 000*" 0.71 1.19 1.47 il6 Pro-inflammatory cytokines, induction of acute phase response and immune regulation 000" 0.64 1.12 1.86 adgrg1 Adhesive GPCRs (GPR56) regulate ECM and smooth muscle cell behavior and immune cell adhesion. 000*" 0.95 1.13 1.31 ntprobnp N-terminal fragment of B-type natriuretic peptide precursor, hemodynamic stress and cardiovascular load markers ~8,500 0.91 1.13 1.43 eda2r The TNF receptor family (XEDAR) can activate NF-κB and participate in cytokine signaling and immune regulation. ~45,000 0.85 1.05 1.31 prss8 Prostasin, a GPI-anchored serine protease, plays a role in barrier / ion channel regulation. ~40,000 1.24 1.01 0.83 hspb6 The low heat shock protein Hsp20 is involved in smooth muscle relaxation and cell protection. 000" 1.14 0.95 0.72
[0154] Differentially expressed proteins were input into the clusterProfiler and STRING databases for GO-BP, KEGG pathway, and PPI network analysis. GO-BP enrichment results ( Figure 4 KEGG analysis showed significant enrichment in immune-inflammatory related biological processes such as "positive regulation of cytokine production," "myeloid leukocyte migration," and "humoral immune response." Figure 5 KEGG revealed significant enrichment of core pathways such as the PI3K-Akt signaling pathway, chemokine signaling pathway, JAK-STAT signaling pathway, and cytokine-receptor interactions. The PPI network ( Figure 6 The rAAA group (PPI) exhibits multiple highly connected modules, with some hub proteins playing important roles in immune regulation and vascular smooth muscle function. Correlation analysis combining aortic diameter and rupture status revealed that most immune chemotaxis-related proteins showed a significant upregulation trend in the rAAA group. For the top-ranked candidate proteins, ELISA and Western blot were used to verify the consistency of expression trends in independent cohorts. Preliminary observations in an ApoE⁻ / ⁻+AngII mouse AAA model indicated that antibody blockade could delay tumor expansion.
[0155] 6) Downstream Analysis
[0156] ① LASSO target screening based on differentially expressed proteins: Clinical grouping (normal group, unruptured abdominal aortic aneurysm group, and ruptured abdominal aortic aneurysm group) and differentially expressed proteins were used as inputs. With normal group, unruptured group, and ruptured group as dependent variables, and the differentially expressed protein matrix as independent variables, 10-fold cross-validation was used to select the optimal regularization coefficient λ. The relationship between the cross-validation error and log(λ) is as follows: Figure 9As shown, the dashed line represents the optimal λ; under this optimal λ, the regression coefficients of each protein entering the model are as follows. Figure 7 As shown (the optimal λ value was determined using 10-fold cross-validation, with the goal of minimizing the mean squared error (MSE) of the prediction risk). Finally, key proteins with non-zero regression coefficients (n=9) were retained, and β values were extracted for subsequent model construction.
[0157] Lasso algorithm is used for feature selection because it can screen a large number of proteins with a relatively small number of events, retaining key proteins to construct an aortic aneurysm risk prediction model. The formula is as follows:
[0158] ;
[0159] Here, x represents the expression level of the protein, and b is the intercept of the LASSO fitted curve.
[0160] ② Drug target analysis based on differentially expressed proteins: The differentially expressed proteins are mapped to databases such as DrugBank, DGIdb, and ChEMBL to identify known or druggable interacting molecules; the feasibility of antibody blocking or ligand competition is assessed by stratification according to pharmacological category (enzyme / receptor / cytokine / adhesion molecule, etc.) and subcellular localization (secretory / extracellular region); the tissue-specific information of HumanProteinAtlas and KEGG pathway are combined to calculate "network proximity / pathway coverage", and a reversal signature score is generated based on LINCS / CMap to screen small molecules with reuse potential; the "Target Accessibility Score (TAS)" and priority list are output for evaluation in in vitro functional experiments and clinical translation.
[0161] Therefore, this embodiment also provides a method for drug target analysis based on proteomics data, including the following steps:
[0162] Differential protein database mapping: Map significantly screened differential proteins to DrugBank, ChEMBL, or DGIdb databases to identify known protein-drug interactions;
[0163] Calculate the target accessibility score (TAS) and stratify the screening based on the potential of antibody blocking, antagonist or small molecule intervention;
[0164] Drug signature prediction: Generate a list of compounds with reverse expression patterns based on LINCS / CMap and validate them in order of priority.
[0165] Preferably, the drug screening includes drugs that target adhesion molecules and cytokine-related proteins.
[0166] ③ Risk scoring and clinical decision assessment analysis based on differentially expressed proteins: A weighted protein risk score (PRS = Σβ_i·E_i) was constructed based on the LASSO-reserved variables and corresponding β values in ①. Discriminative power (ROC / AUC, PR-AUC), calibration (Hosmer–Lemeshow, calibration curve), and net clinical benefit (Decision Curve Analysis) were assessed in the training and validation sets. Net reclassification index (NRI) and integrated discriminant improvement (IDI) were compared with a baseline model based solely on aortic diameter / traditional risk factors. Stratified validation was conducted in key subgroups (diameter <55mm, female, smokers, etc.), and a nomogram and stratified follow-up strategy were constructed to guide early intervention and monitoring. Specifically, its NRI and IDI were compared with the traditional diameter model. Individualized risk prediction tools were output using the nomogram for early screening and stratified follow-up reference. In the independent validation set (UKBiobank, n=52034), the AUC value was 0.91 (ruptured) / 0.85 (not ruptured).
[0167] ;
[0168] The corresponding code for calculating the risk score is as follows:
[0169] #Assume beta is a named vector; E is a normalized expression matrix (column names = protein).
[0170] prs<-as.numeric(E[,names(beta),drop=FALSE]%*%beta)
[0171] #ROC and AUC (95% CI)
[0172] library(pROC)
[0173] roc_obj<-roc(label,prs);auc(roc_obj);ci.auc(roc_obj,method="bootstrap",boot.n=2000)
[0174] #Calibration (logistic example)
[0175] library(rms)
[0176] cal<-calibrate(lrm(label~prs),method="boot",B=1000)# Plot the calibration curve
[0177] #DCA (Net Income)
[0178] #Available rmda::decision_curve(label~prs,thresholds=seq(0,0.5,0.01),bootstraps=1000).
[0179] The protein risk score (PRS) model constructed in this embodiment differs from traditional assessment methods that rely on aortic diameter or a few clinical variables, as explained below:
[0180] Structurally: The model is based on the multi-protein combination screened by LASSO, with the formula PRS=Σβ_i·E_i, making full use of high-dimensional omics features.
[0181] Regarding parameter settings: the regularization parameter λ is optimized through 10-fold cross-validation to ensure that the model is simple and robust.
[0182] In terms of training and validation: ROC, calibration curve and decision curve analysis were performed using independent cohorts (including large samples from UK Biobank), and compared with traditional models, the AUC was improved to 0.91 (rupture prediction), showing better discrimination and clinical applicability.
[0183] These designs ensure the model's predictive performance and generalization ability, and enable its application in high-risk screening for early small aneurysms.
[0184] The detection and analysis results in this embodiment are summarized as follows:
[0185] (1) PCA analysis (results are as follows) Figure 1 , Figure 3 (as stated)
[0186] The separability of PCA is closely related to the co-variation of immune inflammation-related proteins (such as chemokine and cytokine pathway-related proteins); when these features are upregulated in the ruptured group, the sample shifts along the same direction on the PC axis, resulting in intergroup separation. This invention uses principal component analysis (PCA) in the PC1–PC2 plane to visualize samples based on the preprocessed protein expression matrix (deletion-level imputation + Z-score normalization). Figure 1 The three groups (normal, unruptured AAA, and ruptured AAA) are clearly distributed in two-dimensional space: samples in the same group form a relatively tight cluster distribution in the figure, and there is little overlap in the distribution between groups, suggesting that the protein expression profile can effectively distinguish different pathological states.
[0187] (2) Differential protein volcano plot ( Figure 2 )
[0188] Figure 2This is a differential protein volcano plot (horizontal axis: maximum |log2FC|; vertical axis: -log10(FDR)). It visually shows the overall distribution and number of significantly up / downregulated proteins (131 / 0).
[0189] (3) BP / KEGG pathway enrichment ( Figure 4 , Figure 5 )
[0190] Using differentially expressed proteins as input, the Benjamini–Hochberg method was employed for multiple validation (FDR < 0.05) to perform enrichment analysis on biological processes (GO-BP) and KEGG pathways. Results showed that GO-BP was significantly enriched in immune-inflammatory processes such as positive regulation of cytokine production, myeloid leukocyte migration / chemotaxis, and humoral immune response. KEGG was significantly enriched in core pathways such as the PI3K–Akt signaling pathway, chemokine signaling pathway, JAK–STAT signaling pathway, and cytokine–cytokinesereceptor interaction. These results collectively point to the pathological axis of "inflammatory chemotaxis—cytokine cascade—vascular wall remodeling," which is biologically consistent with an increased risk of aortic rupture. The figure captions simultaneously provide the enrichment factor (Gene Ratio / Rich Factor), the number of genes hit, and the FDR, along with representative proteins for each pathway for verification.
[0191] (4) PPI network and Hub protein ( Figure 6 )
[0192] Differentially expressed proteins were mapped to the STRING database (human species, default interaction score threshold ≥ 400) to construct a protein interaction network. Hub proteins were defined as the union of the top 10% of nodes based on degree centrality and betweenness centrality. The network exhibited multiple highly connected modules, primarily enriched in functions related to immune regulation, extracellular matrix (ECM) tissue, and vascular smooth muscle. Hub proteins, as key cross-pathway connectors, may be located in upstream regulatory positions; combined with drug target search results, they were prioritized for inclusion in subsequent in vitro validation and intervention feasibility assessment. Figure captions indicate the STRING version, species, and scoring threshold, and the centrality calculation method is indicated to ensure reproducibility.
[0193] (5) LASSO target screening Figure 7 )
[0194] Using the three categories (normal / unruptured / ruptured) as the dependent variable and the differential protein expression matrix as the independent variable, 10-fold cross-validation was used to select the optimal regularization coefficient λ, preserving the feature that the regression coefficient (|β|) is non-zero (n=9 in this example). Figure 7 The bar charts of the retained features and their selection strengths (|β|) are given as candidate molecule sets for subsequent construction of weighted risk scoring and decision analysis.
[0195] Table 5 lists the specific names and β values of the nine key proteins that ultimately entered the model.
[0196] Protein name Select intensity (β value) Difference factor (log2) KW Inspection Efficiency cxcl17 0.396535 2.333509 9.01E-10 pgf 0.132865 2.955105 5.11E-08 lamp3 0.11289 1.311861 5.11E-08 il6 0.078543 1.936215 1.08E-05 adgrg1 0.041175 2.049302 3.93E-05 ntprobnp 0.010493 1.665885 0.00762 eda2r 0.008071 3.337871 2.85E-08 prss8 0.002327 2.51539 4.54E-07 hspb6 6.29E-05 4.063687 1.91E-05
[0197] This embodiment also provides a method for drug target analysis based on proteomics data, including the following steps:
[0198] Differential protein database mapping: Map significantly screened differential proteins to DrugBank, ChEMBL, or DGIdb databases to identify known protein-drug interactions;
[0199] Calculate the target accessibility score (TAS) and stratify the screening based on the potential of antibody blocking, antagonist or small molecule intervention;
[0200] Drug signature prediction: Generate a list of compounds with reverse expression patterns based on LINCS / CMap and validate them in order of priority.
[0201] Preferably, the drug screening includes drugs that target adhesion molecules and cytokine-related proteins.
[0202] This embodiment also provides a method for assessing the risk of abdominal aortic aneurysm, including the following steps:
[0203] Differential protein extraction: The weight coefficients of retained variables are calculated based on protein expression levels and LASSO regression. In this application, LASSO regression can automatically perform variable selection in high-dimensional proteomics data, remove redundant or noisy features, and retain only the key protein combinations most relevant to the disease, thereby avoiding overfitting and improving model stability and interpretability. Compared with traditional methods, the LASSO regression used in this application has more advantages in processing high-throughput mass spectrometry data and solving the problem of "high-dimensional small sample size", and can significantly improve sensitivity and specificity.
[0204] Risk score construction: Calculate individual patient PRS values using variable weighting;
[0205] Risk model assessment: Calculate metrics such as ROC, AUC, and calibration curves from both training and validation data.
[0206] Clinical decision generation: Based on decision curve analysis, output stratified follow-up strategies for different aortic diameters and genders.
[0207] Preferably, the calibration curve is adjusted for its fit quality according to the Hosmer–Lemeshow test.
[0208] In this embodiment, the risk assessment model is mainly based on LASSO regression, with its core hyperparameter being the regularization parameter λ. 10-fold cross-validation is used to assess the risk within a preset range (e.g., 10). -4 ~10 1 The model was progressively optimized, with the mean squared error (MSE) of predicted risk and model stability serving as the comprehensive evaluation metrics. Finally, the value of λ that minimizes the validation set error was selected to ensure a good balance between the training and validation sets.
[0209] Figure 9 A graph showing the λ selection for the LASSO model; Figure 10 The plot shows the ROC curve, where AAA (AUC 0.743) and AAA (AUC 0.898). Figure 11 For calibration curves; Figure 12 This is a DCA curve, showing the low threshold range (0%–2%).
[0210] Figure 9 This is a graph showing the selection of the optimal regularization parameter λ using 10-fold cross-validation in the LASSO regression model. The horizontal axis represents log(λ), and the vertical axis represents the cross-validation error. Each black dot represents the average error under different λ values, and the dashed line indicates the optimal λ corresponding to the minimum error.
[0211] Figure 11 The figure shows the calibration curves for the abdominal aortic aneurysm rupture risk prediction model. The horizontal axis represents the model-predicted risk, and the vertical axis represents the actual observed event risk. The black dots represent the actual incidence rates after risk grouping. The solid green line is the smoothed calibration curve, and the dashed gray line is the ideal calibration line. The figure also shows the Brier coefficient and the slope of the calibration curve to evaluate the consistency between the model's predicted probability and the actual risk.
[0212] In the low threshold range (0%–2%), the net benefit curve of the model of this invention is consistently higher than that of ultrasound examination, with a visible difference between the two. This invention demonstrates a higher overall net benefit in early screening scenarios where there is a preference for "more screening." In the medium threshold range (2%–6%), the net benefit of the model of this invention remains at 0.004–0.008, while ultrasound gradually decreases to approximately 0.002–0.003, with a stable difference. This range represents the clinically commonly used "medium risk" assessment standard, and this invention still has a significant advantage. In the high threshold range (6%–10%), the net benefits of both methods gradually approach 0, but the model of this invention maintains a positive net benefit of approximately 0.001 even near the 10% threshold, while the ultrasound method approaches 0.
[0213] The risk model provided in this embodiment is used to evaluate the model's detection capability and false positive risk at different thresholds, ensuring high sensitivity in early clinical screening; the Hosmer-Lemeshow test and calibration curve are used to test the consistency between the predicted probability and the actual incidence.
[0214] In this embodiment, decision curve analysis (DCA) is used to measure the net benefit of the model at different risk thresholds, demonstrating its value in clinical applications.
[0215] In this embodiment, PR-AUC (area under the Precision-Recall curve) is a metric that better reflects the stability of the model when the sample distribution is uneven (e.g., there are relatively few broken AAA samples).
[0216] A single AUC cannot fully reflect the predictive efficacy and clinical applicability of a model. This application combines multi-dimensional evaluation to ensure the reliability and generalizability of the risk model in real clinical scenarios.
[0217] In the abdominal aortic aneurysm risk assessment method provided in this embodiment, L1 regularization was used to automatically constrain the regression coefficients during LASSO regression modeling, causing most irrelevant variables to be zeroed out and retaining only a few key proteins, naturally acting as a "pruning" mechanism. After constructing the risk scoring model, the model complexity and overfitting risk were further evaluated to ensure that the number of required input variables was limited and that clinical feasibility was high.
[0218] In this embodiment, the risk score is output as a continuous value, and the appropriate stratification threshold is determined based on ROC analysis, the Youden index, and the sensitivity / specificity balance point under different thresholds.
[0219] In this embodiment, the population can be divided into low-risk, medium-risk, and high-risk groups:
[0220] Low risk (score below the lower limit threshold): Regular follow-up is recommended.
[0221] Medium risk (score in the middle range): It is recommended to shorten the follow-up interval and add imaging examinations if necessary.
[0222] High risk (score above the upper limit threshold): This indicates a significantly increased risk of rupture, and early intervention or surgical evaluation should be considered in conjunction with clinical factors.
[0223] The basis for the tiered recommendations mainly includes:
[0224] The model's ability to distinguish between training and validation queues;
[0225] Gain analysis after comparison with clinically commonly used aortic diameter standards (such as NRI, IDI).
[0226] Current clinical guidelines outline management strategies for different diameters and risk stratifications.
[0227] In this embodiment, the serum protein fingerprint of abdominal aortic aneurysm is detected by MALDI-TOF mass spectrometry and combined with the expression profile of key proteins screened by machine learning algorithm. It is used to: distinguish AAA patients from healthy people, predict the risk of AAA rupture (especially for early small aneurysms), and guide clinical risk stratification and individualized treatment decisions.
[0228] In this embodiment, the key protein expression profile refers to a set of specific proteins and their quantitative expression levels selected through the following process, which serve as the core biomarker combination for constructing the abdominal aortic aneurysm (AAA) risk assessment model:
[0229] (1) Basic screening:
[0230] 131 significantly differentially expressed proteins were screened from the serum proteome using differential expression analysis (Kruskal-Wallis test + FDR correction) (Table 1), satisfying |Log2FC|>1 and FDR<0.05.
[0231] (2) Machine Learning Refinement:
[0232] The LASSO regression algorithm (with 10-fold cross-validation to optimize the λ value) was used to further screen 131 differentially expressed proteins, and finally 9 key proteins with non-zero regression coefficients were retained (Table 3) to form the core expression profile.
[0233] The core functions of this core expression profile are as follows:
[0234] (1) Constructing a risk assessment model (PRS):
[0235] Based on the regression coefficients (β) and expression levels (E_i) of nine key proteins, a weighted protein risk score (PRS) was calculated:
[0236] PRS=Σ(β_i×E_i)
[0237] The model demonstrated high discriminative power in the validation cohort (breakage risk AUC = 0.91, non-breakage risk AUC = 0.85).
[0238] (2) Clinical decision support:
[0239] Auxiliary diagnosis: Differentiating AAA patients from healthy individuals (especially early-stage small aneurysms).
[0240] Predicting rupture risk: superior to traditional aortic diameter assessment (verified by NRI and IDI indices).
[0241] Guide stratified management: Develop individualized follow-up and intervention strategies based on PRS scores (e.g., for high-risk patients with diameter <55mm).
[0242] (3) Pathological mechanism analysis:
[0243] Key proteins are enriched in pathways such as inflammatory chemotaxis (CXCL17 / IL6), vascular remodeling (PGF / ADGRG1), and smooth muscle dysfunction (HSPB6). Figure 4 , Figure 5 , Figure 6 This provides a molecular basis for understanding the pathology of AAA.
[0244] (4) Drug target screening:
[0245] After being mapped to databases such as DrugBank, some proteins (such as CXCL17 and IL6) are identified as potential drug targets and can be used to develop antibodies or small molecule inhibitors.
[0246] In this application, the key protein expression profile specifically refers to nine core serum proteins and their expression level combinations screened by machine learning. Its core value lies in: serving as input variables for the risk assessment model (PRS) to achieve accurate prediction of AAA rupture risk; and providing target resources for pathological mechanism analysis and drug development.
[0247] In summary, the detection and analysis method provided in this embodiment has the following characteristics:
[0248] (1) Integration of multi-omics and machine learning: Combining proteomics (mass spectrometry identification + differential analysis) with machine learning (LASSO regression) to extract proteins at risk of abdominal aortic aneurysm rupture, overcoming the limitations of traditional biomarkers. Screening key protein combinations using LASSO yields high accuracy and strong interpretability.
[0249] (2) High sensitivity and specificity detection process: The proteomics detection process (U9 lysis buffer + WCX2 chip + MALDI-TOFMS) is optimized to achieve efficient capture and quantification of low abundance proteins, and the detection sensitivity is improved to the fmol level.
[0250] (3) Data stability: By using Z-score standardization and missing value imputation strategies, technical bias is reduced and data stability is ensured (CV<10%).
[0251] (4) Clinical significance: It can identify high-risk subgroups of proteins in patients with early abdominal aortic aneurysms (diameter <55mm). The differentially screened proteins provide a theoretical basis for the development of targeted drugs, help clinical drug development, achieve early intervention, and reduce rupture mortality.
[0252] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments. Even if various changes are made to the present invention, if these changes fall within the scope of the claims of the present invention and their equivalents, they shall still fall within the protection scope of the present invention.
Claims
1. A method for detecting and analyzing serum protein fingerprints of abdominal aortic aneurysms based on the LASSO regression algorithm, characterized in that, Includes the following steps: a) Collect blood samples from the subjects and prepare serum; b) The serum samples were pretreated using lysis buffer and buffer solution; the lysis buffer was U9 lysis buffer and the buffer solution was WCX2 buffer with a pH of approximately 4.0, diluted at a ratio of 1:
40. c) Proteins in the pretreated sample were captured using a weak cation exchange chip, followed by rinsing and matrix crystallization. The weak cation exchange chip was a WCX2 chip, the rinsing solution was 0.1% Triton X-100 solution, and the matrix solution was an aqueous solution of acetonitrile in α-cyano-4-hydroxycinnamic acid. d) The processed chip was analyzed by mass spectrometry using a MALDI-TOF mass spectrometer to obtain mass spectrometry data of serum proteins; e) Analyze the mass spectrometry data: perform preprocessing, protein identification, outlier removal and differential expression analysis on the mass spectrometry data to screen out differentially expressed proteins associated with abdominal aortic aneurysm; f) The LASSO regression algorithm was used to screen the differentially expressed proteins to obtain key protein combinations and their regression coefficients. g) Detection and analysis were performed using an abdominal aortic aneurysm risk assessment model constructed based on the aforementioned key protein combination and its regression coefficients; Step e) specifically includes: Data preprocessing: Convert the raw data into a peak list, perform peak filtering, handle missing values, and Z-score standardization; Protein identification: Mass spectrometry data were identified using library search software. Parameters included: maximum allowable number of missed cleavage sites for trypsin digestion (2), peptide length range of 8-50 amino acids, fixed modification urea methylation, variable modification oxidation and acetylation, precursor ion mass tolerance ±20 ppm, daughter ion mass tolerance ±20 ppm, and false discovery rate (FDR) ≤1%. Outlier detection: First, Z-score is used to standardize the original protein expression matrix data. Principal component analysis is performed on proteins with a detection rate >80%. Mahalanobis distance is used to identify and remove outlier samples. Differential expression analysis: The Kruskal-Wallis test was used to compare the protein peak intensities between groups, and the Benjamini-Hochberg method was used to control the false discovery rate. The screening criteria were |Log2FC|>1 and FDR<0.
05. Specifically, data preprocessing includes: For the initial mass spectrometry data in .raw file format, the msconvert tool of ProteoWizard was used to perform format conversion and peak extraction by peak centering, and the output .mzML file and corresponding mass spectrometry peak list results were obtained; the peak list was generated with a threshold of SNR>5 and peak width of 2–20 Da, and low-quality peaks were removed. The rolling ball algorithm is used to roll a radius of [missing information] at the bottom of the spectral line. The "virtual sphere" records the trajectory of the sphere's center as a baseline function. Then use the original signal minus Obtain the corrected signal The rolling ball algorithm uses a window size of 100–500 m / z to effectively remove most of the slowly drifting baselines while maintaining the true peak shape without distortion. Z-score standardization is used to normalize the data and eliminate systematic errors between samples, ensuring that the data conforms to a standard normal distribution. The formula is as follows: in The original peak intensity, For sample error, The standard deviation is denoted as .
2. The detection and analysis method according to claim 1, characterized in that, The mass spectrometry detection in step d) has a mass range of m / z 1000-20000, uses linear mode and positive ion reflection mode, and the number of laser bombardments is not less than 2000 times / point.
3. The detection and analysis method according to claim 1, characterized in that, Step f) Using the LASSO regression algorithm to screen the differentially expressed proteins for features, and obtaining the key protein combinations and their regression coefficients, includes the following steps: K-fold cross-validation is used to optimize the regularization parameter λ to determine the optimal λ value in the Lasso algorithm: all samples are randomly divided into 10 similarly sized subsets; 10 rounds are performed. In each round: one fold is used as the validation set, and 10-1 folds are used as the training set. On the training set, a series of different λ values are used to fit the LASSO model. Each fitted model is then applied to the validation set, and the mean squared error between the predicted and actual values is calculated. For each λ value, there are 10 MSEs from the 10 different validation sets. The average MSE corresponding to this λ value is calculated. A curve is obtained showing the average MSE as a function of λ value. A disease risk prediction model is constructed by using the Lasso algorithm to select features and retain key proteins. The formula is as follows: ,in Here, x represents the expression level of the protein, and b is the intercept of the LASSO fitting curve.
4. The detection and analysis method according to claim 1, characterized in that, The risk assessment model described in step g) is the weighted protein risk score PRS, which is calculated as follows: PRS=Σ(β_i*E_i), where β_i is the LASSO regression coefficient of the i-th key protein and E_i is the expression level of the protein.
5. An abdominal aortic aneurysm risk assessment system, characterized in that, It is configured to perform the method of any one of claims 1-4.
6. The use of the method according to any one of claims 1-4 in the preparation of reagents or systems for assisting in the diagnosis of abdominal aortic aneurysm, assessing the risk of rupture of abdominal aortic aneurysm, or guiding the stratified management of patients with abdominal aortic aneurysm.
Citation Information
Patent Citations
Molecular model established by detecting serum specimen based on proteomics and used for assisting in evaluating renal injury progress of diabetic nephropathy as well as construction method and application of molecular model
CN119269801A