Wilson disease low-abundance urine protein biomarker detection method
By combining DIA mass spectrometry and PRM mass spectrometry, low-abundance protein biomarkers in the urine of Wilson's disease were screened and validated, which solved the problem of insufficient specificity of diagnostic biomarkers in existing technologies and enabled efficient urine proteomics analysis and early diagnosis.
Patent Information
- Application Number
- CN202511078186.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-28
AI Technical Summary
Current technologies for diagnosing Wilson's disease rely on indicators such as serum ceruloplasmin levels, 24-hour urinary copper excretion, and liver tissue copper content. These technologies are either highly invasive or lack specificity, and cannot comprehensively reflect changes in the low-abundance proteome in urine, resulting in relatively limited diagnostic marker functions.
We used DIA mass spectrometry combined with liquid chromatography-ion mobility separation to perform unbiased in-depth analysis of urine samples, screened for potential low-abundance protein biomarkers specific to Wilson's disease, and validated them by PRM mass spectrometry. We then used statistical and bioinformatics methods to screen candidate biomarkers and constructed a support vector machine algorithm model for early screening and monitoring.
This study enabled efficient identification and validation of low-abundance urinary proteins in Wilson's disease, improving the quantity and reliability of protein identification, providing multiple potential protein targets to support early screening and monitoring, and constructing a highly sensitive and specific diagnostic model.
Smart Images

Figure CN121027347A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Wilson's disease technology, and in particular to a method for detecting low-abundance urinary protein biomarkers in Wilson's disease. Background Technology
[0002] Wilson's disease (WD) is a hereditary copper metabolism disorder caused by mutations in the ATP7B gene. It is characterized by abnormal accumulation of copper ions in the body, potentially leading to serious complications such as cirrhosis and neurological damage. Currently, diagnosis relies primarily on indicators such as serum ceruloplasmin levels, 24-hour urinary copper excretion, and liver tissue copper content. However, these diagnostic methods are either highly invasive or lack specificity, thus necessitating the search for more sensitive and reliable early biomarkers.
[0003] In recent years, the development of ultra-sensitive and high-throughput proteomics detection technologies has made it possible to conduct large-scale proteomics analysis using body fluid samples. Proteomics research has revealed that although low-abundance proteins constitute a small proportion of total protein, they are often closely related to specific pathological processes of diseases and may play a crucial role in the early or subclinical stages of disease, serving as important sources of potential disease biomarkers and drug targets. However, due to technological limitations, many potential low-abundance biomarkers may have been overlooked in previous studies, limiting a comprehensive understanding of disease mechanisms and diagnostic biomarkers.
[0004] Compared to blood, urine samples exhibit unique advantages in the study of low-abundance proteins. Plasma contains extremely high levels of high-abundance proteins (such as albumin and immunoglobulins), with a dynamic range of 12 to 13 orders of magnitude. This can mask the signals of low-abundance proteins, affecting their accurate detection. The dynamic range of proteins in urine is smaller than that in blood, making it more suitable for detecting low-abundance proteins. Furthermore, blood components are strictly regulated by the body's homeostasis mechanisms, typically showing significant changes only in the later stages of disease; urine, however, is not subject to this limitation. As a carrier of metabolic waste, it can reflect abnormal changes in the body earlier, which is beneficial for the early and sensitive detection of biomarkers. In addition, urine collection is non-invasive, simple to perform, and has high subject compliance, making it ideal for disease screening and dynamic monitoring.
[0005] Traditional methods for reducing sample complexity to detect low-abundance proteins include immunoaffinity depletion of high-abundance proteins, sample fractionation, and enrichment of specific proteins. However, these methods are often complex and time-consuming, and may suffer from non-specific adsorption leading to the accidental loss of proteins outside the target, thus affecting the accurate detection of low-abundance proteins. With the advancement of mass spectrometry technology, high-throughput, unbiased, and in-depth proteomics analysis of biological fluid samples has become possible. Traditional data-dependent acquisition (DDA) proteomics methods randomly select peptides with high-intensity signals for fragmentation in a single analysis, thus tending to detect high-abundance components and easily missing low-abundance signals, resulting in a large number of missing values; this defect is more pronounced in large-scale sample analysis. Data-independent acquisition (DIA) technology divides the mass spectrometry scan range into several windows, cyclically fragmenting and detecting all ions within each window, which can acquire fragment information of all ions in the sample without omission, significantly reducing data loss. In particular, combining ion mobility separation technology with DIA and using Trapped Ion Mobility Spectroscopy (TIMS) to separate peptide ions in two dimensions can further improve the accuracy and reliability of quantification, achieve higher detection depth and wider coverage, and is more suitable for proteomics analysis of large-scale samples.
[0006] Mass spectrometry-based proteomics offers the potential for discovering novel biomarkers for WD. Previous studies have preliminarily confirmed the feasibility of using mass spectrometry to screen for WD biomarkers in urine; however, existing methods focus on pre-assumed pathways, resulting in biomarkers with limited functionality that cannot comprehensively reflect changes in the urinary proteome in WD. Therefore, it is necessary to develop a more comprehensive and efficient technical approach that can deeply explore low-abundance protein biomarkers in the urine of WD patients without requiring pre-existing assumptions. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of existing technologies that focus on pre-assumed pathways, and the screened biomarkers have relatively simple functions and cannot fully reflect the changes in the urinary proteome of Wilson's disease. This invention provides a method for detecting low-abundance urinary protein biomarkers in Wilson's disease.
[0008] To address the aforementioned technical problems, the present invention provides the following technical solution: A method for detecting low-abundance urinary protein biomarkers in Wilson's disease includes the following steps: S1: Collect urine samples from the Wilson disease test subjects and preprocess the urine samples; S2: DIA mass spectrometry was used to analyze the pretreated urine samples to identify and quantify global protein components in the urine samples in an unbiased manner. S3: Screen candidate biomarker proteins that are significantly associated with Wilson's disease in the global protein composition using statistical and bioinformatics methods; S4: The selected candidate biomarker proteins are validated by targeted mass spectrometry to obtain biomarker protein data.
[0009] As a preferred embodiment of the present invention, the pretreatment of the urine sample in step S1 includes: centrifuging the urine sample to remove impurities and precipitates, and then using a filter membrane-assisted protease digestion method and a C18 solid-phase extraction column to extract, enzymatically digest and purify the proteins in the supernatant.
[0010] As a preferred embodiment of the present invention, step S2 further includes: optimizing the liquid chromatography and mass spectrometry parameters of the pretreated urine sample using nano-level reversed-phase liquid chromatography coupled with quadrupole-time-of-flight tandem mass spectrometry.
[0011] As a preferred embodiment of the present invention, the analysis of the pretreated urine sample using DIA mass spectrometry in step S2 includes: S21: Determine the optimal DIA isolation window parameters and collect data based on the DIA isolation window parameters; S22: The collected data was analyzed using DIA-NN software, and a spectral library-free search mode was used to identify and quantify peptides and proteins.
[0012] As a preferred embodiment of the present invention, the determination of the optimal DIA isolation window parameters in step S21 includes: firstly, using CompassHyStar software to preliminarily determine the optimal window range of the isolation window; secondly, introducing the py_diAID algorithm to automatically generate a dynamic isolation window that matches the ion density based on the two-dimensional distribution of precursor ions in the urine sample; and finally, manually integrating the optimal window range and the width of the dynamic isolation window to form the optimal DIA isolation window parameters.
[0013] As a preferred embodiment of the present invention, step S3 includes: S31: The Limma differential analysis method was used to compare the relative protein abundance between the Wilson disease test group and the healthy control group. The screening threshold was set as adjusted p value < 0.05 and fold change |FC| > 1.5 to identify differentially regulated or downregulated proteins. S32: For the obtained list of differentially expressed proteins, functional enrichment analysis was performed using bioinformatics analysis tools; S33: Calculate the Pearson correlation coefficient of each protein expression and perform agglomerative hierarchical clustering to examine the correlation and molecular grouping characteristics among candidate biomarker proteins, and finally identify several candidate biomarker proteins that are closely related to Wilson's disease and have potential diagnostic value.
[0014] As a preferred embodiment of the present invention, step S4 includes: selecting a suitable characteristic peptide and precursor ion as a target based on the information of the candidate biomarker protein; The screening criteria for the characteristic peptides include: the peptide is a specific and unique peptide of the target protein; its length is between 7 and 12 amino acids; it carries 1 or 2 charges; it has sufficient peak intensity signal in DIA data; the peptide sequence does not contain variable modification sites; and it has no undigested enzyme cleavage sites. The selection criteria for the precursor ion are: the charged ion with the highest response intensity in the peptide segment; and the absence of other co-efferentiation interfering ions.
[0015] As a preferred embodiment of the present invention, step S4 further includes: using nano-level reversed-phase liquid chromatography coupled with quadrupole-time-of-flight tandem mass spectrometry, while switching the mass spectrometry acquisition mode to prm-PASEF mode for mass spectrometry monitoring. The target of mass spectrometry monitoring is the selected characteristic peptide precursor and its fragment ions suitable as targets. The retention time window, ion mobility range and fragment ion information to be detected for each characteristic peptide precursor are pre-input into the CompassHyStar mass spectrometry software. A PRM scan schedule is established according to the data provided by the Skyline software for accurate and targeted detection of candidate peptides.
[0016] As a preferred embodiment of the present invention, step S4 further includes: manually checking the chromatographic peak retention time of each target peptide in each urine sample, aligning the retention time drift between different samples, and removing peptides that did not detect the expected peak shape or signals with significant interference noise. Based on the preset peak morphology and signal-to-noise ratio criteria, high-quality peptide signals and corresponding fragment ions for quantitative analysis were screened out. The screening criteria were: fragment ions and precursor peptides were co-eluted with consistent relative abundance, dot product value > 0.8, no obvious interference peaks, and peak intensity signal-to-noise ratio > 3. For each protein, the most abundant fragment ions in its corresponding peptide that meet the above screening criteria are selected, and the sum of the chromatographic peak areas of these fragment ions is used as the relative quantitative index of the protein. The selected peak area data were normalized using Tukey median smoothing to obtain high-confidence quantitative results for each candidate biomarker protein in each sample.
[0017] Compared with the prior art, the advantages of the present invention are as follows: This invention employs DIA mass spectrometry combined with liquid chromatography-ion mobility separation to perform unbiased, in-depth analysis of the urinary proteome in patients with Wilson's disease and healthy controls. Optimized detection parameters enhance the quantity and reliability of protein identification, screening for potential low-abundance protein biomarkers specific to Wilson's disease. Further targeted validation of candidate biomarkers using parallel reaction monitoring mass spectrometry (PRM) ensures the reliability of the screening results. Finally, support vector machine algorithms and recursive feature elimination (RFE) feature selection are utilized. This invention lays the foundation for developing high-performance early screening and monitoring tools for Wilson's disease based on urine samples, while also providing multiple potential protein targets to support research into new therapies for Wilson's disease. Attached Figure Description
[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. In the drawings: Figure 1 This is a flowchart of a method for detecting low-abundance urinary protein biomarkers for Wilson's disease, as described in Embodiment 1 of the present invention. Figure 2 These are different precursor ion isolation windows for the detection method of low-abundance urinary protein biomarkers for Wilson's disease described in Embodiments 1 and 3 of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0020] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, the terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance, or suggesting any such actual relationship or order between these entities or operations. Additionally, the terms "connected," "linked," etc., can refer to a direct connection between components or an indirect connection via other components.
[0021] Example 1 A method for detecting low-abundance urinary protein biomarkers in Wilson's disease, such as Figure 1 As shown, it includes the following steps: S1: Collect urine samples from the Wilson disease test subjects and preprocess the urine samples; Specifically, the pretreatment of the urine sample in step S1 includes: collecting urine samples from WD (Wilson's disease) subjects, centrifuging to remove impurities and precipitates, and then using filter-assisted protease digestion (FASP) and a C18 solid-phase extraction column (StageTip) to extract, enzymatically digest, and purify the proteins in the supernatant. The treated peptide samples are reconstituted with 0.1% formic acid aqueous solution, and the peptide concentration is adjusted to a final concentration of approximately 50 ng / μL for mass spectrometry analysis.
[0022] S2: DIA mass spectrometry was used to analyze the pretreated urine samples to identify and quantify global protein components in the urine samples in an unbiased manner. Specifically, step S2 also includes coupling nano-level reversed-phase liquid chromatography (Bruker NanoElute) with quadrupole-time-of-flight tandem mass spectrometry (Brukertims TOFPro2). The chromatographic column used was a 25cm × 75μm PepSepUltraC18 column with a 1.5μm particle size, and the column temperature was 50℃. Mobile phase A was 0.1% formic acid in water, and mobile phase B was 0.1% formic acid and 80% acetonitrile in water. The flow rate was 0.3μL / min. The injection volume for each sample was 4μL. The optimized gradient elution program was 60 min, with a total run time of approximately 90 min. Mass spectrometry data were acquired in positive ion mode and dia-PASEF scanning mode; MS1 mass resolution was 60,000 (200 m / z), and MS2 resolution was 30,000 (200 m / z); mass-to-charge ratio (m / z) scan range was 467.4–1067.4; trapping ion mobility range was 0.75–1.40 Vs / cm²; ion mobility ramp separation time and ion accumulation time were both set to 150 ms. These HPLC and mass spectrometry parameters were optimized to achieve high identification depth and stable quantitative repeatability while ensuring a cycle scan time of <2 s.
[0023] Step S2, which involves analyzing the pretreated urine sample using DIA mass spectrometry, includes: S21: Determine the optimal DIA isolation window parameters and collect data based on the DIA isolation window parameters; The determination of the optimal DIA isolation window parameters in step S21 includes: first, using CompassHyStar software to initially determine the optimal window range of the isolation window; second, introducing the py_diAID algorithm to automatically generate a dynamic isolation window that matches the ion density based on the two-dimensional distribution of precursor ions in the urine sample; and finally, manually integrating the optimal window range and the width of the dynamic isolation window to form the optimal DIA isolation window parameters.
[0024] Specifically, compared to the DDA method, the DIA method requires pre-setting a precursor ion isolation window and cyclically fragmenting and acquiring secondary spectra of all peptide precursor ions within the window. A narrower precursor ion isolation window can reduce spectral complexity but increases the total number of windows and prolongs the cycle time required to cover the entire window range; a wider precursor ion isolation window can reduce the total number of windows but significantly increases the difficulty of spectral interpretation. Optimizing the size and distribution of the isolation window is crucial for improving the protein identification capability of the DIA method. The isolation window can be manually set by software or automatically generated by computer algorithms.
[0025] First, the optimal region of the isolation window was initially determined using CompassHyStar software. Based on the distribution characteristics of multi-charged ions in the peptide, the isolation window range was manually set. Under the premise that the average estimated cycle time is <2s, narrow and compact isolation windows were set (e.g., ...). Figure 2 (as shown in the diagram at the top left) and wide and extensive isolation window 2 (as shown in the diagram at the top left) Figure 2 (As shown in the upper right figure). To simultaneously meet the requirements of an average estimated cycle time of <2s and a reduction in the number of windows, the width of each window in isolation window 2 was appropriately widened, and a segmented strategy was adopted. The protein identification rates under the two isolation window settings are shown in Table 1. It can be found that a narrower and more compact isolation window can significantly improve the protein identification rate.
[0026] Although CompassHyStar software supports custom isolation window ranges, it can only generate isolation windows of uniform width. Based on the sample's mass-to-charge ratio-ion mobility two-dimensional distribution, it can be observed that the precursor ions exhibit a non-uniform distribution, with significant differences in mass spectrometry resolution difficulty across different ion density regions. Therefore, the py_diAID software was used to automatically generate a dynamic isolation window 3 that matches the precursor ion density distribution, such as... Figure 2 As shown in the diagram at the bottom left.
[0027] Table 1 Protein identification values of mixed samples with different isolation windows Table 1 shows that the protein identification capacity of isolation window 3 and isolation window 1 is comparable, and the identification ability is not significantly improved. Figure 2 As shown in the diagram at the bottom left, the dynamically generated isolation window by the software includes some mobility regions lacking multi-charged precursor ions, resulting in an overly wide window range. Therefore, by combining the window range of isolation window 1 and the dynamic window width of isolation window 3, isolation window 4 is constructed (as shown in the diagram). Figure 2 (See the diagram in the lower right corner). This window achieved the highest protein identification rate. By integrating the advantages of manual settings and automatic algorithm optimization, the method successfully balanced selectivity (narrower isolation window) and sensitivity (fewer scans), obtaining the optimal isolation window parameters, which were then used to establish the subsequent DIA method.
[0028] S22: The collected data was analyzed using DIA-NN software, and peptides and proteins were identified and quantified using a spectral library search mode; Specifically, the samples were subjected to mass spectrometry analysis using the optimized DIA method described above to obtain DIA data for the urinary proteome. The acquired data was analyzed using DIA-NN software, employing a library-free search mode to identify and quantify peptides and proteins. DIA-NN uses a deep neural network to predict the electrospray mass spectrometry (ESM) patterns of sequences from the UniProt human protein database, constructing a virtual library, which was then analyzed and matched against the actual DIA data. DIA-NN analysis parameters included: protease set to trypsin (allowing a maximum of one missed cleavage site); fixed modifications including N-terminal formamide and cysteine alkylation (IAM); variable modifications including methionine oxidation; an identification threshold (FDR) of 1%; and a mass accuracy tolerance of 10 ppm for both primary and secondary mass spectrometry. The analysis output included the identified proteins and their relative abundance in each sample. Mass spectrometry quality control: To ensure data reliability, strict quality control (QC) measures were implemented during the DIA mass spectrometry analysis process. Specifically, one QC standard sample (HeLa cell protease digestion product, 100 μg / mL) was analyzed after every 10 urine samples, with a QC injection volume of 2 μL each time. Two QC analyses were run before and after the experiment, with a 15-minute wash gradient added before and after each QC. The Bruker PaSER 2023 real-time database search platform was used for real-time identification and monitoring of the QC samples. An instrument performance pass threshold was defined as the number of identified HeLa proteins > 5000 and a TIMS mass deviation of ±5 ppm. If the identification results of consecutive QC samples fell below the threshold, the system automatically stopped the injection and performed instrument calibration or tuning to eliminate potential faults before continuing sample analysis. Throughout the experiment, the protein identification count of all QC samples remained above 5000 without significant fluctuations, indicating that the mass spectrometry platform was stable and the data acquisition was accurate and reliable.
[0029] S3: Screen candidate biomarker proteins that are significantly associated with Wilson's disease in the global protein composition using statistical and bioinformatics methods; Step S3 includes: S31: The Limma differential analysis method was used to compare the relative protein abundance between the Wilson disease test group and the healthy control group. The screening threshold was set as adjusted p value < 0.05 and fold change |FC| > 1.5 to identify differentially regulated or downregulated proteins. S32: For the obtained list of differentially expressed proteins, functional enrichment analysis was performed using bioinformatics analysis tools; S33: Calculate the Pearson correlation coefficient of each protein expression and perform agglomerative hierarchical clustering to examine the correlation and molecular grouping characteristics among candidate biomarker proteins, and finally identify several candidate biomarker proteins that are closely related to Wilson's disease and have potential diagnostic value.
[0030] Specifically, from the vast proteome identified by DIA, candidate biomarkers significantly associated with WD are screened using statistical and bioinformatics methods. Preferably, Limma differential analysis is used to compare the relative abundance of proteins between the WD patient group and the healthy control group. The screening threshold can be set as an adjusted p-value (Benjamini-Hochberg correction) < 0.05 and a fold change |FC| > 1.5 to identify significantly upregulated or downregulated differentially expressed proteins (DEPs). For the obtained list of differentially expressed proteins, functional enrichment analysis can be performed using bioinformatics tools, such as GeneOntology (GO) and KEGG pathway analysis using the DAVID database, to understand the biological processes and pathways involved in the candidate biomarkers. Furthermore, Pearson correlation coefficients for each protein expression can be calculated and hierarchical clustering can be performed to examine the association and molecular grouping characteristics among the candidate biomarkers. Through the above analysis, several candidate urinary protein biomarkers closely related to WD and with potential diagnostic value are identified.
[0031] S4: LC-PRM-MS Targeted Validation: The selected candidate biomarker proteins are validated by targeted mass spectrometry to obtain biomarker protein data.
[0032] Step S4 includes: establishing a PRM detection method, and selecting suitable characteristic peptides and precursor ions as targets based on the information of the candidate biomarker proteins; The screening criteria for the characteristic peptides include: the peptide is a specific and unique peptide of the target protein; its length is between 7 and 12 amino acids; it carries 1 or 2 charges; it has sufficient peak intensity signal in DIA data; the peptide sequence does not contain variable modification sites; and it has no undigested enzyme cleavage sites. The selection criteria for the precursor ion are: the charged ion with the highest response intensity in the peptide segment; and the absence of other co-efferentiation interfering ions.
[0033] Specifically, based on information about candidate proteins, suitable characteristic peptides and precursor ions are selected as targets. Preferred peptide screening criteria include: a) the peptide is a specific and unique peptide of the target protein; b) its length is between 7 and 12 amino acids; c) it carries one or two charges (+1 / +2); d) it has sufficient peak intensity signal in DIA data (signal-to-noise ratio > 3); e) the peptide sequence does not contain variable modification sites; f) it has no undigested restriction enzyme sites (i.e., the peptide sequence does not contain dibasic sequences such as "KK", "RR", "RK", or "KR"). The selection criteria for precursor ions are: a) the charged ion with the highest response intensity in the peptide; b) no other co-eluting interfering ions. According to the above criteria, this invention screened 51 urinary protein biomarkers from candidate proteins, and the corresponding 97 precursor ions were used for subsequent PRM method detection.
[0034] Step S4 further includes: using nano-level reversed-phase liquid chromatography coupled with quadrupole-time-of-flight tandem mass spectrometry, while switching the mass spectrometry acquisition mode to prm-PASEF mode for mass spectrometry monitoring. The target of mass spectrometry monitoring is the selected characteristic peptide precursor and its fragment ions suitable as targets. The retention time window, ion mobility range, and fragment ion information to be detected for each characteristic peptide precursor are pre-input into the CompassHyStar mass spectrometry software. A PRM scan schedule is established according to the data provided by the Skyline software for accurate targeted detection of candidate peptides. The injection volume of each sample is 4 μL, and the other mass spectrometry source and ion mobility parameters remain unchanged to ensure comparability with the detection conditions in the DIA stage.
[0035] Step S4 also includes: analyzing the acquired PRM data using Skyline-daily software. For each sample, the retention time of each target peptide peak is manually checked, retention time drift between different samples is aligned, and peptides that do not show the expected peak shape or signals with significant interference noise are removed. Based on pre-set peak morphology and signal-to-noise ratio standards, high-quality peptide signals and corresponding fragment ions for quantitative analysis are selected: all fragment ions must co-elute with the precursor peptide and have a consistent relative abundance ratio, a dot product value > 0.8, no obvious interference peaks, and a peak intensity signal-to-noise ratio > 3. For each protein, preferably, a maximum of six fragment ions with the highest abundance in its corresponding peptide that meet the above standards are selected, and the sum of the chromatographic peak areas of these fragment ions is used as the relative quantitative indicator of the protein. To correct for differences in signal intensity of different fragment ions, the selected peak area data is normalized using Tukey's median smoothing method. Through the above steps, high-confidence quantitative results for each candidate biomarker protein in each sample can be obtained.
[0036] Construction of a Disease Diagnosis Model Based on SVM: A machine learning model for WD diagnosis was constructed using biomarker protein data obtained from PRM quantification. First, the dataset was preprocessed: missing values were imputed using the K-nearest neighbor algorithm, and protein concentration values from different samples were standardized using the Z-score method to eliminate dimensional differences. Then, the processed data was randomly divided into a training set and a test set in a 7:3 ratio, ensuring that the two sets matched and did not overlap in clinical variables such as gender and age, thus guaranteeing the independence of model training and evaluation. During model training, a support vector machine classification algorithm was introduced, and an appropriate kernel function (such as the radial basis function kernel) was selected based on the data characteristics. The recursive feature elimination (RFE) algorithm was used to iteratively select the subset of biomarkers that contributed most to the classification results, determining the optimal feature combination for the model. Next, 10-fold cross-validation (10-fold CV) was used to evaluate the model's generalization ability on the training set, and the penalty coefficient and other hyperparameters of the SVM model were tuned using a grid search method, ultimately establishing an optimized WD urinary protein biomarker combination diagnostic model. After the model was established, independent test set data were input into the model for blinded prediction. The receiver operating characteristic (ROC) curve and its area under the curve, sensitivity, specificity, and other indicators were used to evaluate the model's performance in distinguishing WD patients from healthy controls. A confusion matrix was generated to calculate evaluation parameters such as accuracy and positive predictive value. The model evaluation results of this invention show that the constructed SVM model can effectively distinguish WD from healthy individuals on both the training and test sets, exhibiting high diagnostic sensitivity and specificity (see examples for details), demonstrating the clinical application potential of the selected combination of urinary protein biomarkers.
[0037] External Validation: To further validate the clinical applicability of the aforementioned model biomarkers, this invention conducted ELISA testing on some biomarkers in an independent population. Specifically, in an independent validation cohort (e.g., 10 newly recruited WD patients and 10 healthy controls), four protein biomarkers included in the SVM model were randomly selected, and the concentrations of the target proteins in urine samples were quantitatively detected using a commercial enzyme-linked immunosorbent assay (ELISA) kit. The experiment was conducted according to the kit instructions, with each sample tested three times repeatedly, and corresponding standards and positive and negative controls were established. The results showed that the expression levels of these biomarkers in the urine of WD patients were consistent with the mass spectrometry quantification results and significantly different from those in the healthy control group, further supporting the correlation between the biomarker combination and the WD disease state, and validating the reliability of the model constructed in this invention.
[0038] The method provided by this invention, through multi-level technological innovation, achieves efficient identification and verification of low-abundance urinary protein biomarkers in Wilson's disease, and has the following significant effects: Improving Protein Identification Depth: By optimizing the liquid chromatography gradient and ion mobility separation conditions, the detection coverage and number of proteins identified in urine were significantly improved. For example, extending the gradient elution time from 20 min to 60 min increased the number of identified proteins by approximately 400, while further extending it to 120 min only identified about 100 more proteins, indicating a diminishing recognition benefit. Considering both identification capability and throughput, 60 min was determined to be the optimal gradient. Furthermore, under the 60 min gradient conditions, a wider ion mobility range (0.75–1.40 Vs / cm²) identified more proteins than a narrower range (0.85–1.30 Vs / cm²). Extending the TIMS ion mobility ramp separation time from 100 ms to 150 ms (ensuring a cycle time <2 s) fully utilizes low-abundance ion signals, significantly increasing the number of proteins detected. The optimized parameter combination was experimentally verified to maximize the identification of low-abundance proteins in urine, providing a rich information basis for subsequent analysis.
[0039] Optimized Isolation Window Improves Identification Rate: This invention employs an isolation window optimization strategy combining manual experience and algorithms to balance the spectral complexity and scanning throughput of the DIA method. By introducing the py_diAID algorithm to generate a dynamic window and combining it with a manual approach, the final isolation window scheme significantly increases the number of proteins identified compared to a conventional fixed-width window. This scheme enables the window width to adaptively adjust according to the precursor ion density, improving identification sensitivity and selectivity while maintaining scanning speed, and fully exploring more useful signals in the sample.
[0040] Large-scale proteomic differential analysis: Using an optimized DIA method, thousands of urinary proteins can be identified in a single experiment. For example, in this embodiment, approximately 2263 urinary proteins were identified, and after data filtering and statistical analysis, 1771 proteins were reliably quantified. Among these, 447 proteins showed significant differential expression between WD patients and healthy individuals (adjusted p < 0.05, FC > 1.5), including 68 upregulated and 379 downregulated proteins. These results demonstrate that the method of this invention can sensitively capture extensive changes in the urinary proteome under disease states, providing a strong basis for discovering novel biomarkers related to WD.
[0041] Validating data reliability and consistency: Re-detection of candidate biomarkers using PRM-targeted mass spectrometry effectively validated the reliability of the DIA screening results. The PRM quantification results of most selected candidate proteins in independent samples were consistent with the DIA quantification trend, and the coefficient of variation was within an acceptable range, indicating that the differences detected in the DIA stage were reproducible. Rigorous peptide screening criteria and data quality control procedures ensured the accuracy of PRM quantification and eliminated interference from false positive biomarkers.
[0042] The model demonstrates outstanding diagnostic efficacy: Based on screened and validated biomarkers, the SVM machine learning model constructed in this invention can efficiently distinguish between WD patients and healthy individuals. Through feature optimization and parameter optimization, the model achieves excellent fitting results on the training set and also exhibits high sensitivity and specificity on the independent test set, proving that the model does not suffer from overfitting. ROC curve analysis shows that the model's AUC value for WD classification is close to 1, indicating that the model has excellent diagnostic discrimination ability. Compared with single biomarkers, the combined biomarker model of this invention can more comprehensively reflect disease characteristics, thereby significantly improving diagnostic accuracy.
[0043] Platform stability and method repeatability: Quality control measures throughout the experiment ensured the stable operation of the mass spectrometry platform and the consistency of data. Analysis results of QC standard samples inserted for every 10 samples showed that the number of proteins identified by the QC samples remained above 5000 without significant fluctuations throughout the experiment, and the instrument performance consistently met the preset requirements. This result demonstrates that the method of this invention has good repeatability and robustness, and the obtained data is reliable and suitable for the detection and analysis of large-scale samples.
[0044] Example 2 This embodiment describes the use of nano-level reversed-phase liquid chromatography coupled with quadrupole-time-of-flight tandem mass spectrometry to optimize the liquid chromatography and mass spectrometry parameters of pretreated urine samples in the detection method for low-abundance urinary protein biomarkers of Wilson's disease described in Example 1. Specifically, this embodiment focuses on the establishment phase of non-targeted proteomics detection for DIA, and optimizes the liquid chromatography gradient program and ion mobility mass spectrometry parameters to determine the conditions that maximize the number of proteins identified.
[0045] Experimental Methods: A certain number of urine samples from WD patients were collected, mixed, and aliquoted to prepare test samples for method optimization. The effects of the following conditions on protein identification were investigated in DDA (Data-Dependent Acquisition) mode: liquid chromatography gradient times of 20 min, 60 min, and 120 min; ion mobility separation ranges of 0.75–1.40 Vs / cm² (wide range) and 0.85–1.30 Vs / cm² (narrow range); and TIMS ion mobility enhancement slope separation time and cumulative time of 100 ms and 150 ms, respectively. The number of proteins identified in real time under different conditions was statistically compared using the PaSER real-time database search platform.
[0046] Results Analysis: With increasing HPLC gradient time, protein detection coverage and identification count gradually increased. Extending the gradient to 60 min significantly improved the number of identified proteins, increasing by approximately 400 proteins compared to the 20 min gradient. Further extending the gradient to 120 min, while increasing the number of identified proteins by approximately 100, resulted in a significantly smaller increase. Considering both protein identification capability and analytical throughput, the 60 min gradient effectively covered the main proteomic information, thus 60 min was selected as the elution time for subsequent analyses. Regarding mobility range optimization, except for the 20 min elution gradient, a wider mobility range (0.75–1.40 Vs / cm²) yielded higher protein identification counts under all other conditions compared to a narrower range (0.85–1.30 Vs / cm²). This indicates that expanding the ion mobility separation range is beneficial for detecting more peptide ion signals with different properties. For the TIMS ramp separation time, extending it from 100 ms to 150 ms significantly improved ion separation and utilization, increasing protein identification count while ensuring the cycle time did not exceed 2 s. Taking into account various factors, the optimal parameter combination for DDA mode was finally determined to be 60 min gradient elution, 0.75–1.40 Vs / cm2 mobility range, and 150 ms slope separation time, and a DIA detection method was established accordingly.
[0047] Example 3 This embodiment is an optimization experiment of the precursor ion isolation window for a method for detecting low-abundance urinary protein biomarkers in Wilson's disease as described in Example 1. This embodiment optimizes and compares the precursor ion isolation window settings unique to the DIA method, aiming to balance spectral complexity and scanning efficiency, thereby further improving protein identification performance.
[0048] Experimental Methods: Based on the optimized chromatographic and mass spectrometric parameters of Example 1, different isolation window partitioning schemes were designed for DIA data acquisition within the mass spectrometry m / z scan range of 467.4–1067.4. First, the isolation window was initially determined using CompassHyStar software based on the distribution characteristics of multi-charged ions in the peptide: Under the premise of an average cycle time <2 s, a narrow and compact isolation window 1 (e.g., ...) was established. Figure 2 (as shown in the diagram at the top left) and wide and extensive isolation window 2 (as shown in the diagram at the top left) Figure 2 (As shown in the diagram at the top right). Isolation window 1 has a larger number of windows, but each window has a narrower range, while isolation window 2 is the opposite. Subsequently, the py_diAID software was used to automatically generate dynamic isolation windows that matched the precursor ion density distribution (isolation window 3, as shown in the diagram at the top right). Figure 2(As shown in the diagram at the bottom left). The characteristic of isolation window 3 is that its width varies with the precursor ion density: a narrow window is used in high ion density regions, while a wider window is used in low density regions. Finally, combining the window range of isolation window 1 and the dynamic window width of isolation window 3, isolation window 4 is obtained (as shown in the diagram at the bottom left). Figure 2 (As shown in the illustration at the bottom right).
[0049] Results and Analysis: The impact of different isolation window schemes on the number of proteins identified was evaluated by comparing the results of DIA-NN analysis. Under the same mass spectrometry analysis time constraints, isolation window 1 (narrow and compact) identified significantly more protein types than isolation window 2 (wide and extensive). This indicates that an excessively wide window leads to too many ions being selected simultaneously, increasing spectral complexity and reducing the software's ability to effectively fragment and identify low-abundance ions. Isolation window 3 (dynamic isolation window) generally yielded similar identification results to isolation window 1 and did not bring the expected gain. In-depth analysis revealed that while isolation window 3 reduced the window size in high-density regions to improve selectivity, it generated some excessively wide windows in ion-sparse m / z regions. These windows contained very few effective precursor ions, thus wasting valuable cycle time. To compensate for the shortcomings of each scheme, combining the window boundary setting of isolation window 1 with the dynamic width adjustment strategy of isolation window 3 (isolation window 4) can simultaneously ensure fine ion separation in high-density regions and reduce the number of windows in low-density regions, thereby achieving a significant increase in the number of proteins identified. Isolation window 4 combines the advantages of manual setting and automatic algorithm optimization, successfully balancing the selectivity (narrower isolation window) and sensitivity (fewer scans) of the method, obtaining the optimal isolation window parameters, and using them to establish the subsequent DIA method.
[0050] Example 4 This embodiment is an SVM diagnostic model construction and ELISA independent validation experiment for WD based on the detection method of low abundance urinary protein biomarkers for Wilson's disease described in Example 1: Based on the above-validated biomarker proteins, this embodiment constructs a support vector machine model for the auxiliary diagnosis of WD, and validates the model through ELISA detection of independent samples.
[0051] SVM Model Construction: The biomarker proteins that passed the PRM quantitative validation in Example 1 and their expression data in each sample were used for machine learning modeling. The samples included two categories: WD patients and healthy controls, with each data matrix having a dimension of (number of samples × number of biomarkers). First, the data was preprocessed: missing data points were filled using the KNN algorithm, and then the expression values of each biomarker were Z-score standardized to eliminate the adverse effects of differences in the absolute content of different proteins on model training. Then, 70% of the samples were randomly selected as the training set, and the remaining 30% as an independent test set for final model evaluation. The ratio of WD to healthy samples in the training set was the same as the overall population to ensure the balance of the model learning process. A support vector machine classification model was constructed using the e1071 library in R. After trial and error, the radial basis function kernel (RBF) was selected because it exhibited good non-linear classification ability on this dataset. Next, the recursive feature elimination (RFE) method was used to select the most discriminative feature subset: the RFE algorithm repeatedly trained the SVM model and eliminated biomarkers that contributed less to classification, ultimately selecting approximately 5-10 optimal feature combinations. After defining the features, the hyperparameters of the SVM model, such as the penalty coefficient C and the kernel function parameter γ, were adjusted using grid search cross-validation to avoid overfitting and achieve optimal classification performance. During optimization, 10-fold cross-validation was used to evaluate the model's average accuracy on the training set until the metrics stabilized and stopped improving. The final optimal SVM model contained multiple WD urinary protein biomarker features and could output a 0 / 1 (binary classification) result to determine whether an input sample was WD. Test set data was input into the model, and the model's prediction results for unknown samples were calculated and compared with the actual labels. Model performance was evaluated by plotting ROC curves and confusion matrices: In this embodiment, the SVM model's ROC curve on the test set was close to the upper left peak, the area under the curve (AUC) reached over 0.95, the sensitivity and specificity were both around 90%, and the accuracy also exceeded 90%. These metrics indicate that the model has good generalization performance and clinical application potential.
[0052] Independent ELISA Validation: To further validate the applicability of the biomarkers relied upon by the model in a large population, four proteins with commercially available detection kits and significant variations in WD were selected from the biomarkers used in the SVM model. ELISA was used for validation in an independent cohort. The independent validation cohort included 10 newly diagnosed WD patients and 10 healthy volunteers. Morning midstream urine samples were collected from each subject. Commercially available enzyme-linked immunosorbent assay (ELISA) kits were used to quantify the four target proteins, strictly following the instructions, including sample dilution, incubation, enzyme-catalyzed colorimetric development, and quantification. Standards and controls were also set on each ELISA plate to plot a standard curve and monitor experimental errors. Results showed that the mean concentrations of these four proteins in the urine of WD patients were significantly different from those in the healthy control group (p<0.01), and the direction of change was consistent with the mass spectrometry analysis results. The dispersion of the values of each indicator within the WD group was greater than that in the control group, consistent with the increased variability of biomarkers in disease states. A simple logistic regression model was constructed using the concentrations of these four biomarkers measured by ELISA, achieving a similar ability to distinguish between WD and other diseases as the mass spectrometry SVM model, thus validating the stability and clinical feasibility of mass spectrometry-based biomarker screening. Therefore, the combination of urinary biomarkers identified in this invention demonstrates good discriminative power across different detection platforms and populations, providing a basis for the future development of rapid on-site detection tools for WD.
[0053] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0054] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. A method for detecting low-abundance urinary protein biomarkers in Wilson's disease, characterized in that, Includes the following steps: S1: Collect urine samples from Wilson's disease subjects and preprocess the urine samples; S2: DIA mass spectrometry was used to analyze the pretreated urine samples to identify and quantify global protein components in the urine samples in an unbiased manner. S3: Screen candidate biomarker proteins that are significantly associated with Wilson's disease in the global protein composition using statistical and bioinformatics methods; S4: The selected candidate biomarker proteins are validated by targeted mass spectrometry to obtain biomarker protein data.
2. The method for detecting low-abundance urinary protein biomarkers in Wilson's disease according to claim 1, characterized in that, The pretreatment of the urine sample in step S1 includes: centrifuging the urine sample to remove impurities and precipitates, and then using a filter membrane-assisted protease digestion method and a C18 solid-phase extraction column to extract, enzymatically digest, and purify the proteins in the supernatant.
3. The method for detecting low-abundance urinary protein biomarkers for Wilson's disease according to claim 1, characterized in that, Step S2 further includes optimizing the liquid chromatography and mass spectrometry parameters of the pretreated urine sample using nano-level reversed-phase liquid chromatography coupled with quadrupole-time-of-flight tandem mass spectrometry.
4. A method for detecting low-abundance urinary protein biomarkers for Wilson's disease according to claim 1 or 3, characterized in that, Step S2, which involves analyzing the pretreated urine sample using DIA mass spectrometry, includes: S21: Determine the optimal DIA isolation window parameters and collect data based on the DIA isolation window parameters; S22: The collected data was analyzed using DIA-NN software, and a spectral library-free search mode was used to identify and quantify peptides and proteins.
5. The method for detecting low-abundance urinary protein biomarkers for Wilson's disease according to claim 4, characterized in that, The determination of the optimal DIA isolation window parameters in step S21 includes: first, using Compass HyStar software to initially determine the optimal window range of the isolation window; second, introducing the py_diAID algorithm to automatically generate a dynamic isolation window that matches the ion density based on the two-dimensional distribution of precursor ions in the urine sample; and finally, manually integrating the optimal window range and the width of the dynamic isolation window to form the optimal DIA isolation window parameters.
6. The method for detecting low-abundance urinary protein biomarkers for Wilson's disease according to claim 1, characterized in that, Step S3 includes: S31: The Limma differential analysis method was used to compare the relative protein abundance between the Wilson disease test group and the healthy control group. The screening threshold was set as adjusted p value < 0.05 and fold change |FC| > 1.5 to identify differentially regulated or downregulated proteins. S32: For the obtained list of differentially expressed proteins, functional enrichment analysis was performed using bioinformatics analysis tools; S33: Calculate the Pearson correlation coefficient of each protein expression and perform agglomerative hierarchical clustering to examine the correlation and molecular grouping characteristics among candidate biomarker proteins, and finally identify several candidate biomarker proteins that are closely related to Wilson's disease and have potential diagnostic value.
7. The method for detecting low-abundance urinary protein biomarkers for Wilson's disease according to claim 1, characterized in that, Step S4 includes: selecting a suitable characteristic peptide and precursor ion as a target based on the information of the candidate biomarker protein; The screening criteria for the characteristic peptides include: the peptide is a specific and unique peptide of the target protein; its length is between 7 and 12 amino acids; it carries 1 or 2 charges; it has sufficient peak intensity signal in DIA data; the peptide sequence does not contain variable modification sites; and it has no undigested enzyme cleavage sites. The selection criteria for the precursor ion are: the charged ion with the highest response intensity in the peptide segment; and the absence of other co-efferentiating interfering ions.
8. The method for detecting low-abundance urinary protein biomarkers for Wilson's disease according to claim 7, characterized in that, Step S4 also includes: using nano-level reversed-phase liquid chromatography coupled with quadrupole-time-of-flight tandem mass spectrometry, while switching the mass spectrometry acquisition mode to prm-PASEF mode for mass spectrometry monitoring. The target of mass spectrometry monitoring is the selected characteristic peptide precursor and its fragment ions suitable as targets. The retention time window, ion mobility range and fragment ion information to be detected for each characteristic peptide precursor are pre-input into the Compass HyStar mass spectrometry software. A PRM scan schedule is established according to the data provided by the Skyline software for accurate and targeted detection of candidate peptides.
9. The method for detecting low-abundance urinary protein biomarkers for Wilson's disease according to claim 8, characterized in that, Step S4 also includes: manually checking the chromatographic peak retention time of each target peptide in each urine sample, aligning the retention time drift between different samples, and removing peptides that did not detect the expected peak shape or signals with significant interference noise. Based on the preset peak morphology and signal-to-noise ratio criteria, high-quality peptide signals and corresponding fragment ions for quantitative analysis were screened out. The screening criteria were: fragment ions and precursor peptides were co-eluted with consistent relative abundance, dot product value > 0.8, no obvious interference peaks, and peak intensity signal-to-noise ratio > 3. For each protein, the most abundant fragment ions in its corresponding peptide segment that meet the screening criteria described above are selected, and the sum of the chromatographic peak areas of these fragment ions is used as the relative quantitative index of the protein. The selected peak area data were normalized using Tukey median smoothing to obtain high-confidence quantitative results for each candidate biomarker protein in each sample.