Identification method of sea cucumber producing area characteristic protein based on label-free data independent acquisition mass spectrometry technology
By employing label-free data-independent acquisition mass spectrometry technology and DDA/DIA mass spectrometry library construction, the problems of poor reproducibility and high cost of low-abundance proteins in sea cucumber origin tracing have been solved, achieving high coverage depth and high accuracy of origin identification, which is suitable for large-scale sample cohort analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN POLYTECHNIC UNIVERSITY
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies for tracing the origin of sea cucumbers suffer from random ion selection bias, resulting in poor reproducibility and limited coverage of low-abundance proteins. Furthermore, the labeling methods are costly and difficult to apply to large-scale sample cohort analysis.
Label-free data-independent acquisition mass spectrometry was employed, and exploratory proteomics methods were combined with DDA and DIA mass spectral libraries to identify characteristic proteins of sea cucumber origin, thereby constructing a spectral library with broader coverage and higher specificity, and achieving efficient identification and quantification of low-abundance proteins.
It significantly improves the accuracy and reproducibility of sea cucumber origin identification, simplifies the sample pretreatment process, reduces costs, and is suitable for the analysis of large-scale sample cohorts.
Smart Images

Figure CN122017105A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of food science and analytical chemistry, specifically to a method for identifying sea cucumber place-specific proteins based on label-free data-independent acquisition mass spectrometry. Background Technology
[0002] Sea cucumber, as a high-value marine food, has its quality and market value closely related to its place of origin. Sea cucumbers from Liaoning, China, enjoy a higher market premium due to their superior collagen integrity. However, the molecular basis of this geographical difference, especially at the proteomics level, remains unclear. Currently, proteomics studies on food geographical traceability mostly employ data-dependent acquisition mass spectrometry (DDA) or label-based quantitative methods. These methods generally suffer from random ion selection bias, leading to poor reproducibility and limited coverage of low-abundance proteins. While labeling methods (such as TMT and iTRAQ) improve the accuracy of quantification, their sample processing is complex and costly, and the number of channels is limited, making them difficult to apply to the analysis of large-scale sample cohorts.
[0003] Therefore, there is an urgent need in this field for a new mass spectrometry data acquisition mode suitable for large-scale sample analysis, in order to overcome the randomness problem of traditional data acquisition mass spectrometry technology, thereby achieving higher reproducibility, deeper coverage and more accurate quantitative results, thus providing sufficient theoretical and technical support for the origin traceability and identification of marine foods, especially high-value marine foods such as sea cucumbers. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the present invention aims to provide a method for identifying sea cucumber origin-specific proteins based on label-free data-independent acquisition mass spectrometry (DDA). This method uses DDA to acquire and process data from sea cucumber samples, identifying over 6,000 proteins in a single analysis. Its coverage depth is significantly superior to existing label-free methods such as SWATH and MaxQuant, completely overcoming the defect of the general DDA method in which the spectral library of sea cucumber samples is severely biased towards high-abundance proteins, thus significantly improving the accuracy of sea cucumber origin identification.
[0005] To achieve the above objectives, the following technical solution is provided: This invention provides a method for identifying sea cucumber origin-specific proteins based on label-free data-independent acquisition mass spectrometry, the method comprising the following steps: (1) Sample collection and protein extraction Sea cucumbers from known and unknown origins were collected, and proteins were extracted from the body wall tissue of each sea cucumber sample and from the body wall tissue of the mixture of all sea cucumber samples. (2) Peptide digestion and fractionation The proteins extracted in step (1) were digested with trypsin using the membrane-assisted sample preparation method, centrifuged, and the supernatant was collected. The enzymatic digest supernatant of proteins extracted from the body wall tissue of all sea cucumber samples was fractionated, desalted, and quantified using a high pH reverse phase fractionation kit. (3) Construction of DDA spectral library All the graded and quantified peptides from all sea cucumber samples in step (2) were mixed in equal amounts and subjected to DDA mass spectrometry analysis to construct a sea cucumber DDA mass spectrometry library. (4) Construction of DIA mass spectrometry library The supernatant of each sea cucumber sample obtained in step (2) was used with the same mass spectrometry platform as in step (3) and mass spectrometry detection was performed using DIA mode to construct a DIA mass spectrometry library. (5) Data Analysis Using specialized software, the DDA spectral library constructed in step (3) is used to search the DDA data, thereby enabling protein identification and quantification for each sample. (6) Identification of characteristic proteins of sea cucumber origin Standard methods for exploratory proteomics research were used to analyze known and unknown proteins. Based on t-tests, characteristic proteins of sea cucumber origin were identified with a p-value <0.05 and an absolute change fold >1.5 or <0.67.
[0006] In one embodiment, the protein extraction in step (1) can be performed using existing extraction processes.
[0007] In one implementation, the known origin in step (1) is specifically the Liaoning sea area; the unknown origin is a non-Liaoning sea area, including the Shandong sea area and the Fujian sea area.
[0008] In one embodiment, the specific process of protein extraction in step (1) is as follows: After grinding the sea cucumber body wall tissue with liquid nitrogen, UA buffer was added to lyse the homogenate product, followed by ultrasonic treatment and boiling water bath heating, centrifugation, and collection of the supernatant.
[0009] In one embodiment, the UA buffer contains 8 mol / L urea, 0.15 mol / L Tris-HCl, and pH 8.0.
[0010] In one embodiment, the filter membrane-assisted sample preparation method in step (2) specifically involves: reducing the protein with dithiothreitol, followed by alkylation with iodoacetamide at room temperature in the dark, transferring the reaction mixture to a 10 kDa molecular weight cutoff centrifuge filter, and washing it sequentially with UA buffer and ammonium bicarbonate; after washing, adding trypsin for enzymatic hydrolysis.
[0011] In one embodiment, the conditions for the enzymatic hydrolysis reaction in step (2) are: temperature of 35~37℃ and time of 15~18h.
[0012] In one embodiment, the centrifugation parameters in step (2) are: 10000×g~15000×g, and the time is 5~10min.
[0013] In one embodiment, the desalting in step (2) is performed using a C18 solid-phase extraction column.
[0014] In one embodiment, the ultraviolet quantification in step (2) is performed at a wavelength of 280 nm.
[0015] In one embodiment, the chromatographic conditions for the DDA mass spectrometry analysis in step (3) are as follows: the chromatographic column is a reversed-phase C18; the mobile phase A is 0.1% formic acid and the mobile phase B is 84% acetonitrile containing 0.1% formic acid.
[0016] In one embodiment, the mass spectrometry detection conditions for the DDA mass spectrometry analysis in step (3) are as follows: positive ion mode is used, and the acquisition parameters are as follows: the mass range of the full scan mass spectrum is set to 350-1800 m / z, the resolution is 60000 (at 200 m / z), and the automatic gain control target value is 1×10⁻⁶. 6 The maximum injection time was 50 milliseconds, and the dynamic exclusion time was set to 10.0 seconds to reduce repeated fragmentation. After each full scan, 20 data-dependent secondary mass spectrometry scans were triggered based on a "preset inclusion list" generated from pre-experiments and theoretical knowledge of sea cucumber samples. During secondary spectrum acquisition, precursor ions were fragmented via a 1.5 m / z isolation window and subjected to high-energy collision dissociation at a normalized collision energy of 30 eV. The secondary spectrum recording resolution was 30,000 (at m / z 200), and the automatic gain control target value was 1 × 10⁻⁶. 5 The maximum injection time is 50 milliseconds.
[0017] In one implementation, the parameters for mass spectrometry detection in the DIA mode of step (4) are: The mass spectrometry acquisition parameters are as follows: each DIA cycle includes one full scan and 44 consecutive isolation windows covering 350–1800 m / z; the full scan is acquired in profile mode with a resolution of 120,000 (at m / z 200) and an automatic gain control target value of 3 × 10⁻⁶. 6 Maximum injection time: 30 milliseconds; DIA window fragmentation at 30 eV collision energy; second-order spectrum resolution: 30000 (at m / z 200); automatic gain control target value: 3 × 10⁻⁶. 6 The maximum injection time is automatically adjusted.
[0018] Beneficial effects: The present invention provides a method for identifying sea cucumber origin-specific proteins based on label-free data-independent acquisition mass spectrometry. Compared with existing technologies, this method has the following advantages: (1) The customized acquisition method of this invention maximizes the acquisition of real and diverse fragmented spectra, especially those of low-abundance targets, from the "front end"; while the customized retrieval method interprets and filters these spectra with stricter rules from the "back end". This collaborative design between the front end and the back end ensures that the final constructed spectral library not only has a wider coverage, but also higher specificity and reliability, completely overcoming the defect of the general DDA method when used for sea cucumber samples, which is seriously biased towards high-abundance proteins; (2) Compared with the existing technology that uses standard and universal DDA acquisition parameters, the spectral library constructed by the method of the present invention can significantly improve the identification ratio of low-abundance functional proteins (such as signaling pathway-related proteins and transcription regulatory factors) while maintaining a similar or slightly improved total number of identifications. This provides a more realistic and in-depth reference spectral library for subsequent DIA quantitative analysis, solving a key technical obstacle in sea cucumber proteomics research. (3) Extremely high protein coverage depth: The method of the present invention can identify more than 6,000 proteins in a single analysis. Its coverage depth is significantly better than existing label-free methods such as SWATH and MaxQuant, and is comparable to or even better than high-cost labeling methods such as TMT and iTRAQ. (4) Excellent reproducibility and quantitative accuracy: DIA technology eliminates the randomness of ion selection, which greatly improves the reproducibility of detection of low-abundance proteins, and the coefficient of variation of technical reproducibility remains stable at a low level. (5) High cost-effectiveness and high throughput: The label-free design eliminates the need for expensive isotope labeling and simplifies the sample pretreatment process, making the present invention particularly suitable for geographic origination studies of large-scale sample cohorts. Attached Figure Description
[0019] Figure 1 This is a three-dimensional principal component analysis diagram of the protein expression profiles of quality control samples and experimental samples in an embodiment of the present invention; Figure 2 This refers to the average data point distribution of chromatographic peaks in the embodiments of the present invention; Figure 3 This is a statistical graph of the peak capacity of the chromatographic column in the DIA experiment of this invention. Detailed Implementation
[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. The specific embodiments described below further illustrate the present invention.
[0021] The source of raw materials involved in this invention: Trypsin was purchased from Promega (Cat. No. 317107).
[0022] Example 1 A method for identifying sea cucumber origin-specific proteins based on label-free data-independent acquisition mass spectrometry includes the following steps: (1) Sample pretreatment Two hundred sea cucumbers were collected from aquaculture farms along the Chinese coast, including 80 from Liaoning waters (4 sampling sites) and 120 from non-Liaoning waters (3 sampling sites in Fujian Province and 3 sampling sites in Shandong Province). All samples were transported on dry ice and stored in an ultra-low temperature freezer at -80°C, and underwent only one freeze-thaw cycle before protein extraction. Protein extraction was completed within one month after sample collection. (2) Protein extraction and peptide preparation Protein extraction: Equal amounts of body wall tissue (approximately 2 g) were cut from 200 sea cucumber samples and mixed. After grinding with liquid nitrogen, 500 μL of UA buffer (8 mol / L urea, 0.15 mol / L Tris-HCl, pH 8.0) was added to lyse the homogenate. The homogenate was then sonicated and heated in a boiling water bath for 15 minutes. After centrifugation at 14000×g for 40 minutes (4°C), the supernatant was collected and protein was quantified using a BCA protein quantification kit. Peptide preparation: The supernatant (containing 20 μg of protein) was reduced with 10 mmol dithiothreitol at 37°C for 1.5 h (600 rpm orbital oscillation), followed by alkylation with 20 mmol iodoacetamide at room temperature in the dark for 30 min. The reaction mixture was transferred to a 10 kDa molecular weight cutoff centrifuge filter and washed sequentially with 100 μL UA buffer (8 mol urea, 150 mmol Tris-HCl pH 8.0; 3 times) and 25 mmol ammonium bicarbonate (2 times). After washing, trypsin was added at a 1:50 (mass ratio) enzyme-protein ratio and enzymatically digested at 37°C for 15-18 h (overnight). The filtrate was collected by centrifugation at 14000×g, which yielded the released peptides. The collected peptides were fractionated into 10 fractions using the Thermo Scientific Pierce High pH Reverse Phase Peptide Fractionation Kit. Both individual and fractionated peptides were desalted using a C18 solid-phase extraction column. The desalted peptides were lyophilized, resuspended in 40 μL of 0.1% formic acid solution, and quantified using a NANODROP 2000C spectrophotometer at 280 nm. Before mass spectrometry analysis, each sample was incorporated with an iRT calibration peptide at a 2:1 (sample:iRT) ratio for retention time correction. (3) Construction of a customized data dependency acquisition (DDA) spectral library for sea cucumbers All peptides fractionated and quantified in step (2) were mixed in equal volumes and analyzed using a Thermo Scientific Q-Exactive HF-X mass spectrometer coupled with an Easy-nLC 1200 nanofluid chromatography system. Peptide separation was achieved using a reverse-phase C18 analytical column, with elution using a constant flow rate of 300 nanoliters / minute and a linear gradient of mobile phase A (0.1% formic acid) and mobile phase B (84% acetonitrile containing 0.1% formic acid). Mass spectrometry detection was performed in positive ion mode, with the following acquisition parameters: the mass range of the full scan mass spectrum was set to 350-1800 m / z, the resolution was 60000 (at m / z 200), and the automatic gain control target value was 1×10⁻⁶. 6 The maximum injection time was 50 milliseconds, and the dynamic exclusion time was set to 10.0 seconds to reduce repeated fragmentation. After each full scan, 20 data-dependent secondary mass spectrometry scans were triggered based on a "preset inclusion list" generated from pre-experiments and theoretical knowledge of sea cucumber samples. During secondary spectrum acquisition, precursor ions were fragmented via a 1.5 m / z isolation window and subjected to high-energy collision dissociation at a normalized collision energy of 30 eV. The secondary spectrum recording resolution was 30,000 (at m / z 200), and the automatic gain control target value was 1 × 10⁻⁶. 5 Maximum injection time: 50 milliseconds; The preset list includes: peptides detected in sea cucumber pre-experiments, their precise mass-to-charge ratio (m / z) and retention time (RT), as well as a spectral library of known functionally important peptides obtained from theoretical knowledge or public database information that were not captured during sea cucumber pre-experiments due to low expression levels. (4) Protein extraction and peptide preparation Protein extraction: Approximately 2 grams of body wall tissue was cut from each sea cucumber sample, ground in liquid nitrogen, and then lysed in 500 μL of UA buffer (8 mol / L urea, 0.15 mol / L Tris-HCl, pH 8.0). The homogenate was then sonicated and heated in a boiling water bath for 15 minutes. After centrifugation at 14000×g for 40 minutes (4°C), the supernatant was collected and protein was quantified using a BCA protein quantification kit. Peptide preparation: The supernatant (containing 20 μg protein) was reduced with 10 mmol dithiothreitol at 37°C for 1.5 h (600 rpm orbital oscillation), followed by alkylation with 20 mmol iodoacetamide at room temperature in the dark for 30 min. The reaction mixture was transferred to a 10 kDa molecular weight cutoff centrifuge filter and washed sequentially with 100 μL UA buffer (8 mol urea, 150 mmol Tris-HCl pH 8.0; 3 times) and 25 mmol ammonium bicarbonate (2 times). After washing, trypsin was added at a 1:50 (w / w) enzyme-protein ratio and enzymatically hydrolyzed at 37°C for 15-18 h (overnight). The filtrate was collected by centrifugation at 14000×g. (5) Data Independence Acquisition Analysis Guided by DDA (DIA) DIA data were acquired from the filtrate of each sample collected in step (4). The same mass spectrometry platform was used for DIA data analysis, and the mass spectrometry acquisition parameters were specially designed and optimized for the complex composition and wide dynamic range of sea cucumber proteome samples. The mass spectrometry acquisition parameters are as follows: each DIA cycle includes one full scan and 44 consecutive isolation windows covering 350-1800 m / z; the full scan is acquired in profile mode with a resolution of 120,000 (at 200 m / z) and an automatic gain control target value of 3 × 10⁻⁶. 6 Maximum injection time: 30 milliseconds; DIA window fragmentation at 30 eV collision energy; second-order spectrum resolution: 30000 (at m / z 200); automatic gain control target value: 3 × 10⁻⁶. 6 (6) Mass spectrometry data analysis DDA spectral library construction: The raw data collected in step (3) was processed using data software, and the FASTA sequence database from UniProt was searched; to enhance retention time calibration, iRT peptide sequences were appended to the database; the search parameters were set as follows: enzyme specificity was set to trypsin, and the maximum number of missed cleavage sites was 1; the fixed modification was cysteine aminomethylation; the variable modifications included methionine oxidation and N-terminal protein acetylation; all reported data were based on a protein identification confidence level of 99%, and the confidence level was determined by the false discovery rate (FDR) ≤ 1%, forming a peptide spectral matching library; DIA Data Processing: The DIA dataset was analyzed using the same software platform, with spectral matching against a pre-established DDA library. Key software parameters included: retention time prediction based on dynamic iRT; secondary spectral level interference correction; cross-run intensity normalization; post-analysis filtering was performed using a Q-value threshold of 0.01, with the false detection rate for corresponding peptides and proteins controlled to <1%; thus, a protein quantification matrix was formed, including the identified proteins (more than 6000) for each of the 200 sea cucumbers, as well as their expression levels in the sample. (7) Differential identification of proteins The standard methods for exploratory proteomics research were used to identify the differences between proteins from Liaoning and non-Liaoning regions. Preliminary screening was based on t-tests with P-values <0.05 and absolute fold changes >1.5 or <0.67. A total of 64 differentially regulated proteins were identified, of which 37 were upregulated (absolute fold change >1.5) and 27 were downregulated (absolute fold change <0.67). Detailed results are shown in Table 1. Table 1. Significantly differentially expressed proteins (DEPs) detected between sea cucumbers from Liaoning and those from non-Liaoning (P<0.05).
[0023]
[0024] (8) Quality control sample evaluation: To monitor system stability and experimental reproducibility, a quality control sample was inserted for every 7 test samples in the data analysis (this study used a mixed sample); the consistency of the 24 quality control replicates was evaluated using two orthogonal methods: coefficient of variation analysis and principal component analysis; such as Figure 1 Principal component analysis showed that the median coefficient of variation of the quality control samples was approximately 30%, and the quality control replicates showed tight clustering, which together indicate that the technical variation was minimized and the system performance was robust.
[0025] DIA Workflow Validation: Evaluating DIA Performance from Three Key Parameters (1) Chromatographic peak data points: The results are as follows Figure 2 As shown, an average of 6 data points were collected for each chromatographic peak (minimum threshold > 5 points) to ensure sufficient sampling for accurate peak integration; (2) Peak capacity: The results are as follows Figure 3 As shown, the peak capacity of all samples exceeded 550, which is better than the typical performance benchmark of nanoliter liquid chromatography systems (about 200), confirming that it has sufficient separation complexity to achieve deep proteomic coverage. (3) iRT retention time: All 11 internal standard peptides in the iRT kit were detected and showed stable elution curves across all runs (median retention time variation <30%), enabling reliable retention time correction between samples. This consistency allows for direct comparative analysis of HPLC-MS data for individual samples.
[0026] The embodiments provided above are not intended to limit the scope of the invention, nor are the described steps intended to limit the order of execution. Any obvious modifications made to the invention by those skilled in the art based on existing common knowledge also fall within the scope of protection defined by the claims.
Claims
1. A method for identifying sea cucumber place-specific proteins based on label-free data-independent acquisition mass spectrometry, characterized in that, The method includes the following steps: (1) Sample collection and protein extraction Sea cucumbers from known and unknown origins were collected, and proteins were extracted from the body wall tissue of each sea cucumber sample and from the body wall tissue of all sea cucumber samples. (2) Peptide digestion and fractionation The proteins extracted in step (1) were digested with trypsin using the membrane-assisted sample preparation method, centrifuged, and the supernatant was collected. The enzymatic digest supernatant of proteins extracted from the body wall tissue of all sea cucumber samples was fractionated, desalted, and quantified using a high pH reverse phase fractionation kit. (3) Construction of DDA spectral library All the graded and quantified peptides from all sea cucumber samples in step (2) were mixed in equal amounts and subjected to DDA mass spectrometry analysis to construct a sea cucumber DDA mass spectrometry library. (4) Construction of DIA mass spectrometry library The supernatant of each sea cucumber sample obtained in step (2) was used with the same mass spectrometry platform as in step (3) and mass spectrometry detection was performed using DIA mode to construct a DIA mass spectrometry library. (5) Data Analysis Using specialized software, the DDA spectral library constructed in step (3) is used to search the DDA data, thereby enabling protein identification and quantification for each sample. (6) Identification of characteristic proteins of sea cucumber origin Standard methods for exploratory proteomics research were used to analyze known and unknown proteins. Based on t-tests, characteristic proteins of sea cucumber origin were identified with a p-value <0.05 and an absolute change fold >1.5 or <0.
67.
2. The method according to claim 1, characterized in that, The protein extraction in step (1) can be performed using existing extraction processes.
3. The method according to claim 1, characterized in that, The specific steps of the filter membrane-assisted sample preparation method in step (2) are as follows: the protein is reduced with dithiothreitol, and then alkylated with iodoacetamide at room temperature in the dark. The reaction mixture is transferred to a 10 kDa molecular weight cutoff centrifuge filter and washed sequentially with UA buffer and ammonium bicarbonate. After washing, trypsin is added for enzymatic hydrolysis.
4. The method according to claim 1, characterized in that, The conditions for the enzymatic hydrolysis reaction in step (2) are: temperature 35~37℃ and time 15~18h.
5. The method according to claim 1, characterized in that, The centrifugation parameters in step (2) are: 10000×g~15000×g, and the time is 5~10min.
6. The method according to claim 1, characterized in that, The desalting in step (2) is performed using a C18 solid-phase extraction column.
7. The method according to claim 1, characterized in that, The ultraviolet quantification in step (2) is performed at a wavelength of 280 nm.
8. The method according to claim 1, characterized in that, The chromatographic conditions for the DDA mass spectrometry analysis in step (3) are as follows: the chromatographic column is a reversed-phase C18; the mobile phase A is 0.1% formic acid and the mobile phase B is 84% acetonitrile containing 0.1% formic acid.
9. The method according to claim 1, characterized in that, The mass spectrometry detection conditions for the DDA mass spectrometry analysis in step (3) are as follows: positive ion mode is used, and the acquisition parameters are as follows: the mass range of the full scan mass spectrum is set to 350-1800 m / z, the resolution is 60000, and the automatic gain control target value is 1×10⁻⁶. 6 The maximum injection time is 50 milliseconds, and the dynamic exclusion time is set to 10.0 seconds to reduce repeated fragmentation. During secondary spectrum acquisition, precursor ions are isolated using a 1.5 m / z window and fragmented via high-energy collision dissociation at a normalized collision energy of 30 eV. The secondary spectrum recording resolution is 30,000, and the automatic gain control target value is 1 × 10⁻⁶. 5 The maximum injection time is 50 milliseconds.
10. The method according to claim 1, characterized in that, The parameters for mass spectrometry detection in the DIA mode described in step (4) are as follows: each DIA cycle includes one full scan and 44 consecutive isolation windows covering 350-1800 m / z; the full scan is acquired in profile mode with a resolution of 120,000 and an automatic gain control target value of 3 × 10⁻⁶. 6 Maximum injection time: 30 milliseconds; DIA window fragmentation at 30 eV collision energy; secondary spectrum resolution: 30000; automatic gain control target value: 3×10⁻⁶. 6 The maximum injection time is automatically adjusted.