System for identifying wild and culture sources of black carps and application of system
By using label-free quantitative proteomics and the OPLS-DA model, specific biomarkers for wild and farmed grass carp were screened, solving the problems of low accuracy and high cost in existing grass carp identification techniques, and achieving efficient and low-cost identification results.
Patent Information
- Application Number
- CN202511919999.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-06
AI Technical Summary
Existing technologies struggle to accurately distinguish between wild and farmed grass carp. Morphological methods have low accuracy, isotope analysis is costly, and metabolomics struggles to differentiate between short-term stress and long-term adaptation.
Label-free quantitative proteomics was used to analyze grass carp samples by liquid chromatography-mass spectrometry. Combined with the OPLS-DA classification and discriminant analysis model, protein biomarkers with origin-specific characteristics were screened to construct an identification system.
It achieves efficient and low-cost identification of wild and farmed grass carp, with good stability and specificity, does not rely on large instruments, and is suitable for traceability of aquatic food products.
Smart Images

Figure CN121613035A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aquatic food traceability, specifically to a system for identifying the wild and farmed origins of grass carp and its application. Background Technology
[0002] As an important freshwater economic fish in my country, the wild population of grass carp is on the verge of decline due to overfishing and habitat destruction. Meanwhile, the large-scale development of the aquaculture industry has led to confusion between wild and farmed products in the market. Against this backdrop, establishing accurate and efficient technologies for identifying the wild and farmed origins of grass carp has become crucial for supporting resource conservation law enforcement and market supervision.
[0003] Compared to the limitations of metabolic biomarkers, which are susceptible to environmental fluctuations, protein biomarkers, with their high stability and specificity, demonstrate unique value in the field of biomarker identification. As direct products of gene expression, proteins' expression patterns better reflect the genetic and physiological characteristics of an organism's long-term adaptation to its environment, and are less prone to significant fluctuations due to short-term changes in feeding conditions. Among existing identification techniques, morphological methods rely on empirical judgment, resulting in low accuracy; isotope analysis, while reflecting food chain differences, is costly and dependent on large instruments; and while metabolomics can capture immediate physiological states, it struggles to distinguish between short-term stress and long-term adaptive differences. The maturity of proteomics technology provides a new approach to overcoming these bottlenecks. By systematically analyzing the differences in protein expression profiles between wild and farmed black carp populations, it is hoped that biomarker combinations with origin-specific characteristics can be screened, laying the foundation for establishing standardized identification methods. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a system for identifying the wild and farmed origins of grass carp and its application. The specific details of this invention are as follows: In a first aspect, this invention analyzes proteins from wild and farmed grass carp using label-free quantitative proteomics. To ensure the high reliability and statistical significance of the screened differentially expressed proteins and their corresponding peptides, strict screening parameters are set. Specifically, at the protein quantification level, the Protein Group FDR (False Discovery Rate) is set to >1%; at the peptide quantification level, the Precursor FDR is set to >1%. Only peptides and their corresponding proteins that simultaneously meet these two core parameters are included in the subsequent differential analysis candidate library. Based on this, a system for identifying the wild and farmed origins of grass carp is provided, the system comprising at least: Sample processing module: The sample processing module is at least used to process the grass carp sample to obtain the raw data of the polypeptide as shown in any of SEQ ID NO. 1-42, and to obtain the proteomic component response of the sample to be tested; Data processing module: The data processing module is at least used to import the proteomic component responses of the sample to be tested into the OPLS-DA classification and discriminant analysis model to obtain a score plot for subsequent discriminant analysis of the sample data; Result output module: The result output module is used at least to output results.
[0005] Further, the specific operation of the sample processing module is as follows: pre-processing the sample to be tested for proteomics determination to obtain a polypeptide solution including at least any of the polypeptides described in SEQ ID NO.1-42, and determining the obtained polypeptide solution by liquid chromatography-mass spectrometry to obtain the proteomics component response of the sample to be tested.
[0006] Furthermore, the liquid chromatography conditions for determining the polypeptide solution by liquid chromatography-mass spectrometry are as follows: Column: bioZen TM Peptide PS-C18 (100 mm × 2.1 mm, 1.6 μm) (Phenomenex, Torrance, CA, USA); Column temperature: 40℃; Mobile phase A: 0.1% formic acid-acetonitrile, Mobile phase B: 0.1% formic acid-water; Mobile phase gradient: 0–2 min, 5% A; 2–27 min, 5–20% A; 27–37 min, 20–55% A; 37–39 min, 55–80% A; 39–42 min, 80% A; 42–46 min, 5% A; Flow rate: 0.25 mL / min; Injection volume: 40 μL; Elution time: 46 min; Mass spectrometry conditions: TOF scan range: 350-1500 Da; Positive ion reaction mode, GS1: 60, GS2: 50, Curtain Gas: 40, ISVF: 5500, TEM: 525, DP: 100, CE: 10.
[0007] Furthermore, the method for preparing the OPLS-DA classification and discriminant analysis model in the judgment module is as follows: S1: Protein extraction step: The muscle powder from wild and farmed grass carp samples is pretreated for proteomic analysis to obtain a polypeptide solution containing at least the polypeptide markers described in SEQ ID NO. 1-42. S2: Protein detection step: The polypeptide solution obtained in step S1 is determined by liquid chromatography-mass spectrometry to obtain the liquid chromatography-mass spectrometry analysis results; S3: Import the liquid chromatography-mass spectrometry analysis results from step S2 into PEAKS Online software for database retrieval and relative quantitative analysis. The parameters are set as follows: the database is the Cyprinidae database (sourced from NCBI); the enzyme digestion method is set to Trypsin; the maximum number of missed cleavage sites is set to 2; the precursor ion mass error is 10 ppm; the fragment ion mass error is 0.02 Da; and the false positive rate (FDR) thresholds for both peptide and protein levels are set to 1%. S4: The protein identification list and normalized expression levels of each protein in all samples obtained from the PEAKS software in step S3 are used as the initial data matrix. This matrix and the corresponding wild and farmed sample grouping information are imported into the SIMCA software (MKS Umetrics AB) to obtain the OPLS-DA classification discriminant analysis model. Differential proteins are screened by comparing the OPLS-DA models of wild and farmed animals. The screening parameters are set as follows: VIP>1; in S-plot analysis, |p[corr]|>0.75; the confidence interval of the load plot knife-cut method does not include the zero point; in pairwise comparisons of OPLS-DA, the parameter change factor is >1.4 (or <0.71) and p<0.05.
[0008] Furthermore, the specific steps for discriminant analysis are as follows: use the score plot output by the model to perform discriminant analysis on subsequent sample data. The sample region represented by the position where the subsequent data is projected in the score plot is predicted as the growth mode of the test sample.
[0009] Furthermore, the polypeptide is derived from at least one of the following: thiamine pyrophosphokinase 1 (TPK1), thiamine pyrophosphokinase 1-like (TPK1-L), adenylosuccinate synthetase isozyme 1 isoform X1, myozenin-1b isoform X2, myosin light polypeptide chain 2 (partial), myosin light chain 1, skeletal muscle isoform, and fructose-bisphosphate aldolase A.
[0010] Optionally, the polypeptide derived from thiamine pyrophosphokinase 1 (TPK1) includes at least one of the sequences shown in SEQ ID NO. 1-4.
[0011] Optionally, the polypeptide derived from thiamine pyrophosphokinase 1-like (TPK1-L) includes at least IMETPDQDLTDFTK (SEQ ID NO.5).
[0012] Optionally, the polypeptide derived from adenylosuccinate synthase isozyme 1 isoform X1 includes at least one of the sequences shown in SEQ ID NO. 6-9.
[0013] Optionally, the peptide derived from myozenin-1b isoform X2 includes at least MALPFGGFDK (SEQ ID NO.10).
[0014] Optionally, the polypeptide derived from myosin light polypeptide chain 2 (partial) includes at least one of the sequences shown in SEQ ID NO. 11-35.
[0015] Optionally, the polypeptide derived from myosin light chain 1, skeletal muscle isoform, includes at least one of the sequences shown in SEQ ID NO.36-40.
[0016] Optionally, the polypeptide derived from fructose-bisphosphate aldolase A includes at least one of the sequences shown in SEQ ID NO.41-42.
[0017] Furthermore, the polypeptides described in SEQ ID NO.1-42 are detected by peak elution time, as follows: The elution times of the peptides derived from thiamine pyrophosphokinase 1 (TPK1) were: SEQ ID NO.1 RT: 34.69 min; SEQ ID NO.2 RT: 24.61 min; SEQ ID NO.3 RT: 23.84 min; SEQ ID NO.4 RT: 32.91 min.
[0018] The peak elution time of the peptide derived from thiamine pyrophosphokinase 1-like (TPK1-L) was SEQ ID NO.5 RT: 23.84 min.
[0019] The peak elution times of the peptide derived from adenylosuccinate synthetase isozyme 1 isoform X1 were: SEQ ID NO.6 RT: 35.89 min; SEQ ID NO.7 RT: 23.9 min; SEQ ID NO.8 RT: 25.81 min; SEQ ID NO.9 RT: 34.39 min.
[0020] The peak elution time of the peptide derived from myozenin-1b isoform X2 was: SEQ ID NO.10 RT: 26.76 min.
[0021] The elution times of the peptides derived from myosin light polypeptide chain 2 (partial) were as follows: SEQ ID NO.11 RT: 31 min; SEQ ID NO.12 RT: 31.93 min; SEQ ID NO.13 RT: 33.66 min; SEQ ID NO.14 RT: 11.7 min; SEQ ID NO.15 RT: 32.63 min; SEQ ID NO.16 RT: 24.01 min; SEQ ID NO.17 RT: 31.22 min; SEQ ID NO.18 RT: 17.86 min; SEQ ID NO.19 RT: 21.34 min; SEQ ID NO.20 RT: 22.91 min; SEQ ID NO.21 RT: 37 min; SEQ ID NO.22 RT: 36.97 min; SEQ ID NO.23 RT: 6.85 min; SEQ ID NO.24 RT: 32.74 min; SEQ ID NO. NO.25RT: 29.37 min; SEQ ID NO.26 RT: 33.64 min; SEQ ID NO.27 RT: 23.99 min; SEQ ID NO.28RT: 26.76 min; SEQ ID NO.29 RT: 13.18 min; SEQ ID NO.30 RT: 22.92 min; SEQ ID NO.31RT: 17.06 min; SEQ ID NO.32 RT: 31.01 min; SEQ ID NO.33 RT: 31.06 min; SEQ ID NO.34 RT: 27.99 min; SEQ ID NO.35 RT: 27.86 min.
[0022] The elution times of the peptides derived from myosin light chain 1, skeletal muscle isoform were: SEQ ID NO.36 RT: 36.33 min; SEQ ID NO.37 RT: 39.28 min; SEQ ID NO.38 RT: 10.83 min; SEQ ID NO.39 RT: 34.25 min; SEQ ID NO.40 RT: 36.37 min.
[0023] The elution times of the peptides derived from fructose-bisphosphate aldolase A were: SEQ ID NO.41 RT: 1.2 min; SEQ ID NO.42 RT: 20.53 min.
[0024] In a second aspect, the present invention provides a method for identifying the wild and farmed origins of grass carp, the method comprising the step of using the system described above to analyze and determine the origin of the grass carp.
[0025] Furthermore, the method includes at least the following steps: S1: Protein extraction step: Perform pretreatment of the sample to be tested for proteomics analysis to obtain a protein solution comprising at least one of the polypeptides described in SEQ ID NO. 1-42: S2: Protein detection step: The protein solution obtained in step S1 is determined by liquid chromatography-mass spectrometry to obtain the liquid chromatography-mass spectrometry analysis results of any of the polypeptides described in SEQ ID NO. 1-42; S3: Data analysis step: Import the liquid chromatography-mass spectrometry analysis results obtained in step S2 into the system to obtain the analysis results.
[0026] Furthermore, the pretreatment for proteomics assay in step S1 includes at least the step of adding trypsin to the sample protein solution for enzymatic digestion.
[0027] Furthermore, the specific operation of step S1 is as follows: S1-1: Weigh a certain amount of fish sample, homogenize it into powder, add protein extraction solution, shake to extract protein, centrifuge at high speed and low temperature, and take the supernatant. Optionally, the protein extraction solution includes 8M urea and 50mM NH4HCO3. S1-2: Add dithiothreitol to the supernatant obtained in step S1-1 and react for a certain time to obtain reaction solution 1; S1-3: Add the freshly prepared iodoacetamide solution to the reaction solution 1 obtained in step (2) which has been cooled to room temperature, and react the reaction solution 2 at room temperature in the dark; S1-4: Use a 10K filter membrane for ultrafiltration for 25-35 min to obtain reaction solution 2 in step (3), and repeatedly rinse the filter membrane with ammonium bicarbonate solution to obtain proteome solution; S1-5: Digest the proteome solution obtained in step (5) with protease solution for 16-18 hours; S1-6: Ultrafiltration is performed using a 10K membrane, and the filtrate collected from the lower layer is the polypeptide solution.
[0028] Optionally, steps S1-4 are repeated at least 3 times and the proteome solutions obtained each time are combined.
[0029] Optionally, the proteases described in S1-5 include at least one of serine protease, cysteine protease, aspartic protease, metalloproteinase, threonine protease, and glutamate protease.
[0030] Furthermore, the serine protease includes at least one of trypsin, thrombin, elastase, and subtilisin; the cysteine protease includes at least one of papain, bromelain, and cathepsin; the aspartic protease includes at least one of pepsin and renin; and the metalloproteinase includes at least one of collagenase, thermophilic protease, and angiotensin-converting enzyme.
[0031] In a specific embodiment of the present invention, the protease is a serine protease, preferably a trypsin.
[0032] Furthermore, the specific operation of step S2 is as follows: Detection was performed using liquid chromatography-quadrupole / time-of-flight mass spectrometry. Mobile phase A: 0.1% formic acid-acetonitrile, Mobile phase B: 0.1% formic acid-water. Mobile phase gradient: 0–2 min, 5% A; 2–27 min, 5–20% A; 27–37 min, 20–55% A; 37–39 min, 55–80% A; 39–42 min, 80% A; 42–46 min, 5% A.
[0033] Optionally, the method includes at least the following steps: (1) Weigh 1g of sample, homogenize it into powder, add 10mL of protein extraction solution (8M urea, 50mM NH4HCO3), shake to extract protein, centrifuge at 4℃ at high speed (10000r / min), and transfer 400µl of supernatant to EP tube; (2) Add 8µL of 1mol / L DTT to the above protein solution, shake in a water bath at 60℃, and react for 30 min; (3) Cool to room temperature, then add 5µL of freshly prepared 1mol / L IAA and react at room temperature in the dark for 1 hour; (4) Use a 10K filter membrane with 8000g for ultrafiltration for 20 minutes, rinse the upper layer of the filter membrane repeatedly with 200µL of 50mmol / L ammonium bicarbonate solution, and transfer it to a new EP tube; (5) Add 200µL of ammonium bicarbonate solution and repeat this step. Combine the solutions to complete the extraction of proteins under the membrane. (6) Use a nucleic acid protein concentration analyzer to determine the protein solution concentration, add trypsin to the protein solution at an enzyme / substrate ratio of 1:50 (w / w), mix well, and then enzymatically digest on a membrane at 37 °C for 16-18 h; (7) Use a 10K filter membrane with 8000g for ultrafiltration for 20 minutes, collect the lower layer of peptide filtrate, and wait for instrument testing.
[0034] (ii) On-machine testing, Detection was performed using liquid chromatography-quadrupole / time-of-flight mass spectrometry. Liquid chromatography conditions: Column: bioZen TM Peptide PS-C18 (100 mm × 2.1 mm, 1.6 μm) (Phenomenex, Torrance, CA, USA); Column temperature: 40℃; Mobile phase A: 0.1% formic acid-acetonitrile, Mobile phase B: 0.1% formic acid-water; Mobile phase gradient: 0–2 min, 5% A; 2–27 min, 5–20% A; 27–37 min, 20–55% A; 37–39 min, 55–80% A; 39–42 min, 80% A; 42–46 min, 5% A; Flow rate: 0.25 mL / min; Injection volume: 40 μL; Elution time: 46 min; Mass spectrometry conditions: TOF scan range: 350-1500 Da; Positive ion reaction mode, GS1: 60, GS2: 50, Curtain Gas: 40, ISVF: 5500, TEM: 525, DP: 100, CE: 10.
[0035] Using the pretreatment and liquid chromatography-mass spectrometry analysis results described above, the obtained raw data were imported into PEAKSOnline software (version 12, Bioinformatics Solutions Inc.) for database retrieval and relative quantitative analysis. The following parameters were used: the database was the Cyprinidae database (sourced from NCBI); the enzyme digestion method was set to Trypsin; the maximum number of missed cleavage sites was set to 2; the precursor ion mass error was 10 ppm; the fragment ion mass error was 0.02 Da; and the false positive rate (FDR) thresholds for both peptide and protein levels were set to 1%.
[0036] Subsequently, the protein identification list exported from PEAKS software and the normalized expression levels of each protein across all samples were used as the initial data matrix. This matrix, along with the corresponding wild- and farmed sample grouping information, was imported into SIMCA software (MKSUmetrics AB) to construct a multivariate analysis model. Before OPLS-DA modeling, the data underwent Pareto scaling preprocessing to enhance the model's stability and interpretability.
[0037] Diagram of the constructed OPLS-DA classification and discriminant analysis model.
[0038] The score plot output by the model is used for discriminant analysis of subsequent sample data. The sample region represented by the position where subsequent data is projected onto the score plot is predicted as the growth mode of the test sample.
[0039] pass R 2 X (cum) R 2 Y (cum) and Q 2 The quality of the OPLS-DA model was evaluated using parameters such as (cum). The robustness of the model was assessed using 200 permutation tests. Differential proteins were screened by comparing OPLS-DA models from wild and domesticated farming methods. The screening parameters were set as follows: VIP>1; in S-plot analysis, |p[corr]|>0.75; the confidence interval of the load plot knife-cut method did not include the zero point; in pairwise comparisons of OPLS-DA, the parameter change factor was >1.4 (or <0.71), and p<0.05.
[0040] The peptides from grass carp include SEQ ID NO.1-42, and the constructed OPLS-DA model... R 2 X (cum) R 2 Y (cum) Q 2 The cum values were 0.562, 0.993, and 0.903, respectively, indicating that the model has good fitting and predictive abilities. The evaluation results of 200 permutation tests show that... R 2 Y and Q 2 The intercepts of the regression lines are 0.926 and -0.695, respectively, and the slope of the regression lines is close to 0, indicating that the model has good robustness.
[0041] In a third aspect, the present invention provides a kit for identifying wild and farmed grass carp, the kit comprising at least reagents for quantitative and / or qualitative detection of any of the peptides described in SEQ ID NO. 1-42.
[0042] Furthermore, the kit includes at least one of a protein extraction reagent and a protein detection reagent.
[0043] Further, the protein extraction reagent includes at least one of the following: protein extraction solution, DTT (dithiothreitol), IAA (iodoacetamide), ammonium bicarbonate solution, and protease solution; the protein detection reagent includes at least one mobile phase solution. Optionally, the protease includes at least one of serine protease, cysteine protease, aspartic protease, metalloproteinase, threonine protease, and glutamate protease; the serine protease includes at least one of trypsin, thrombin, elastase, and subtilisin; the cysteine protease includes at least one of papain, bromelain, and cathepsin; the aspartic protease includes at least one of pepsin and renin; the metalloproteinase includes at least one of collagenase, thermophilic protease, and angiotensin-converting enzyme; the mobile phase solution includes mobile phase A comprising 0.1% formic acid-acetonitrile, and mobile phase B comprising mobile phase B: 0.1% formic acid-water.
[0044] In a specific embodiment of the present invention, the kit includes at least a lysis buffer, a reducing agent, an alkylating agent, a protease, and a buffer.
[0045] Optionally, the lysis buffer includes at least one of urea extract, SDS denaturing buffer, and Tris-HCl solution; The reducing agent includes dithiothreitol (DTT). The alkylating agent includes iodoacetamide (IAA); The protease includes at least one of trypsin, Lys-C protease, and Glu-C protease; The buffer solution includes an ammonium bicarbonate solution.
[0046] In a fourth aspect, the present invention provides the application of the system, method, or kit described herein in the preparation of products for identifying wild and farmed grass carp.
[0047] Furthermore, the identification of wild and farmed sources of grass carp includes at least one of wild-farmed grass carp from the Huai River, wild-farmed grass carp from Qili Lake, and domesticated grass carp.
[0048] The beneficial effects of the present invention include, but are not limited to: The system disclosed in this invention for identifying the wild and farmed origins of grass carp can be effectively used to identify the wild and farmed origins of grass carp. It has high stability and strong specificity, does not require the use of large instruments, has low detection costs, and has broad application prospects in the traceability of aquatic food products.
[0049] The obtained proteomic component responses were combined with data from wild-caught and farmed fish samples to form a multivariate analysis matrix. After OPLS-DA classification and discrimination, an analysis system was constructed. The OPLS-DA model constructed in this system... R 2 X (cum) R 2 Y (cum) Q 2 The cum values were 0.562, 0.993, and 0.903, respectively, indicating that the system has good fitting and predictive abilities. In a specific embodiment of this aspect, the evaluation results of 200 permutation tests show that... R 2 Y and Q 2 The intercepts of the regression lines are 0.926 and -0.695, respectively, and the slope of the regression lines is close to 0, indicating that the system has good robustness. Attached Figure Description
[0050] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1This is a proteomic fingerprint of grass carp in two samples, wild and farmed, in this embodiment of the invention. P1 is thiamine pyrophosphokinase 1 (TPK1), P2 is thiamine pyrophosphokinase 1-like (TPK1-L), P3 is adenylosuccinate synthetase isozyme 1 isoform X1, P4 is inositol-1b isoform X2, P5 is myosin light polypeptide chain 2 (partial), P6 is myosin light chain 1 (skeletal muscle isoform), and P7 is fructose-bisphosphate aldolase A. Figure 2 These are proteomic chromatographic fingerprints of wild and farmed grass carp in embodiments of the present invention; Figure 3 These are the proteomic mass spectrometry fingerprints of wild and farmed grass carp in this embodiment of the invention; Figure 4 This is a diagram of the OPLS-DA classification and discriminant analysis model for wild and farmed grass carp in this embodiment of the invention; Figure 5 This is a graph showing the results of 200 permutations of the OPLS-DA model for wild and farmed grass carp in this embodiment of the invention. Figure 6 This is a schematic diagram illustrating the predictive analysis results of the wild and farmed modes of the OPLS-DA classification discriminant analysis model for the test sample in this embodiment of the invention. Detailed Implementation
[0051] The present invention is described in detail below with reference to the embodiments, but the present invention is not limited to these embodiments. Unless otherwise specified, the raw materials and catalysts in the embodiments of the present invention are all purchased through commercial channels.
[0052] Example 1: Fingerprint analysis of proteome samples from both wild and cultured sources. Wild and farmed grass carp samples were selected from the Huai River and Qili Lake basins. Samples were taken from the upper back muscles of the grass carp, and immediately frozen at -80℃ after collection for later use. First, the samples were pretreated for proteomic analysis, as follows: (1) Weigh 1g of sample and homogenize it into powder. Add 10mL of protein extraction solution (8M urea, 50mM NH4HCO3) and shake to extract the protein. Centrifuge at 4℃ at high speed and low temperature (10000r / min). Take 400µl of the supernatant and transfer it to an EP tube. (2) Add 8µL of 1mol / L DTT (dithiothreitol) to the above protein supernatant, shake in a water bath at 60℃, and react for 30 minutes; (3) Cool to room temperature, then add 5µL of freshly prepared 1mol / L IAA (iodoacetamide), and react at room temperature in the dark for 1 hour; (4) Use a 10K filter membrane with 8000g for ultrafiltration for 20 minutes, rinse the upper layer of the filter membrane repeatedly with 200µL of 50mmol / L ammonium bicarbonate solution, and transfer it to a new EP tube; (5) Add 200 μL of ammonium bicarbonate solution and repeat this step. Combine the solutions to complete the submembrane protein extraction and obtain the protein solution. (6) Use a nucleic acid protein concentration analyzer to determine the protein solution concentration, add trypsin to the protein solution at an enzyme / substrate ratio of 1:50 (w / w), mix well, and then enzymatically digest on a membrane at 37 °C for 16-18 h; (7) Use a 10K filter membrane with 8000g for ultrafiltration for 20 minutes, collect the lower layer of polypeptide solution filtrate, and wait for instrument testing.
[0053] Secondly, the protein composition of the solution was analyzed by liquid chromatography-mass spectrometry (LC-MS). The method of LC-MS is as follows: Detection was performed using liquid chromatography-quadrupole / time-of-flight mass spectrometry. Liquid chromatography conditions: Column: bioZen TM Peptide PS-C18 (100 mm × 2.1 mm, 1.6 μm) (Phenomenex, Torrance, CA, USA); Column temperature: 40℃; Mobile phase A: 0.1% formic acid-acetonitrile, Mobile phase B: 0.1% formic acid-water; Mobile phase gradient: 0–2 min, 5% A; 2–27 min, 5–20% A; 27–37 min, 20–55% A; 37–39 min, 55–80% A; 39–42 min, 80% A; 42–46 min, 5% A; Flow rate: 0.25 mL / min; Injection volume: 40 μL; Elution time: 46 min; Mass spectrometry conditions: TOF scan range: 350-1500 Da; Positive ion reaction mode, GS1: 60, GS2: 50, Curtain Gas: 40, ISVF: 5500, TEM: 525, DP: 100, CE: 10.
[0054] In proteomics analysis, quality control samples were prepared using a pooled sample method, with 100 μL of each extract taken and thoroughly mixed to serve as the quality control samples. The injection sequence and frequency of the quality control samples were arranged as follows: Before the injection sequence, the QC sample was injected three times consecutively to equilibrate the entire LC-MS system; during the sequence, a QC sample was inserted every 12 samples, for a total of 6 insertions to cover the entire injection sequence and ensure stability assessment of the entire analysis process. Simultaneously, throughout the entire injection sequence, calibration was performed using positive ion mass axis calibration solution at a frequency of once every 5 samples.
[0055] The proteome was analyzed by mass spectrometry and statistical software, and its content differed between wild and farmed grass carp samples (see...). Figure 1 ).
[0056] Example 2: Chromatographic fingerprints of the proteome in wild and farmed grass carp The protein solution prepared in Example 1 was subjected to chromatographic analysis using the same method as in Example 1. Column: bioZen TM Peptide PS-C18 (100 mm × 2.1 mm, 1.6 μm) (Phenomenex, Torrance, CA, USA); Column temperature: 40℃; Mobile phase A: 0.1% formic acid-acetonitrile, Mobile phase B: 0.1% formic acid-water; Mobile phase gradient: 0–2 min, 5% A; 2–27 min, 5–20% A; 27–37 min, 20–55% A; 37–39 min, 55–80% A; 39–42 min, 80% A; 42–46 min, 5% A; Flow rate: 0.25 mL / min; Injection volume: 40 μL; Elution time: 46 min; The results showed that the chromatograms of the proteome differed between the wild and farmed grass carp samples (see...). Figure 2 ).
[0057] The peak elution times of the peptides derived from thiamine pyrophosphokinase 1 (TPK1) were as follows: SEQ ID NO.1 RT: 34.69 min; SEQ ID NO.2 RT: 24.61 min; SEQ ID NO.3 RT: 23.84 min; SEQ ID NO.4 RT: 32.91 min.
[0058] The peak elution time of the peptide derived from thiamine pyrophosphokinase 1-like (TPK1-L) was SEQ ID NO.5 RT: 23.84 min.
[0059] The peak elution times of the peptide derived from adenylosuccinate synthetase isozyme 1 isoform X1 were: SEQ ID NO.6 RT: 35.89 min; SEQ ID NO.7 RT: 23.9 min; SEQ ID NO.8 RT: 25.81 min; SEQ ID NO.9 RT: 34.39 min.
[0060] The peak elution time of the peptide derived from myozenin-1b isoform X2 was: SEQ ID NO.10 RT: 26.76 min.
[0061] The elution times of the peptides derived from myosin light polypeptide chain 2 (partial) were as follows: SEQ ID NO.11 RT: 31 min; SEQ ID NO.12 RT: 31.93 min; SEQ ID NO.13 RT: 33.66 min; SEQ ID NO.14 RT: 11.7 min; SEQ ID NO.15 RT: 32.63 min; SEQ ID NO.16 RT: 24.01 min; SEQ ID NO.17 RT: 31.22 min; SEQ ID NO.18 RT: 17.86 min; SEQ ID NO.19 RT: 21.34 min; SEQ ID NO.20 RT: 22.91 min; SEQ ID NO.21 RT: 37 min; SEQ ID NO.22 RT: 36.97 min; SEQ ID NO.23 RT: 6.85 min; SEQ ID NO.24 RT: 32.74 min; SEQ ID NO. NO.25RT: 29.37 min; SEQ ID NO.26 RT: 33.64 min; SEQ ID NO.27 RT: 23.99 min; SEQ ID NO.28RT: 26.76 min; SEQ ID NO.29 RT: 13.18 min; SEQ ID NO.30 RT: 22.92 min; SEQ ID NO.31RT: 17.06 min; SEQ ID NO.32 RT: 31.01 min; SEQ ID NO.33 RT: 31.06 min; SEQ ID NO.34 RT: 27.99 min; SEQ ID NO.35 RT: 27.86 min.
[0062] The elution times of the peptides derived from myosin light chain 1, skeletal muscle isoform were: SEQ ID NO.36 RT: 36.33 min; SEQ ID NO.37 RT: 39.28 min; SEQ ID NO.38 RT: 10.83 min; SEQ ID NO.39 RT: 34.25 min; SEQ ID NO.40 RT: 36.37 min.
[0063] The elution times of the peptides derived from fructose-bisphosphate aldolase A were: SEQ ID NO.41 RT: 1.2 min; SEQ ID NO.42 RT: 20.53 min.
[0064] Example 3: Mass spectrometric fingerprints of the proteome in wild and farmed grass carp The polypeptide solution prepared in Example 1 was subjected to mass spectrometry analysis. The mass spectrometry analysis method and parameters were the same as in Example 1. The results showed that the mass spectra of the polypeptide differed between the wild and farmed grass carp samples (see...). Figure 3 ).
[0065] Example 4: Construction of OPLS-DA models and screening of differentially expressed peptides in wild and farmed grass carp Using the pretreatment and liquid chromatography-mass spectrometry analysis results described in Examples 1-3, the obtained raw data were imported into PEAKS Online software (version 12, Bioinformatics Solutions Inc.) for database retrieval and relative quantitative analysis. The following parameters were used: the database was the Cyprinidae database (sourced from NCBI); the enzyme digestion method was set to Trypsin, and the maximum number of missed cleavage sites was set to 2; the precursor ion mass error was 10 ppm, and the fragment ion mass error was 0.02 Da; the false positive rate (FDR) threshold for both peptide and protein levels was set to 1%. The database search results are shown in Table 1.
[0066] Table 1
[0067] Among them, XP_016140938.1 is thiamine pyrophosphokinase 1 (TPK1), XP_016431480.1 is thiamine pyrophosphokinase 1-like (TPK1-L), XP_016113702.1 is adenylosuccinate synthase isozyme 1 isoform X1, XP_043109845.1 is inositol-1b isoform X2, ACV87365.1 is myosin light polypeptide chain 2 (partial), and XP_058643928.1 is myosin light chain 1, skeletal muscle subtype (myosin light chain 1). 1, skeletal muscle isoform), RXN03471.1 is fructose-bisphosphate aldolase A.
[0068] Subsequently, the protein identification list exported from PEAKS software and the normalized expression levels of each protein across all samples were used as the initial data matrix. This matrix, along with the corresponding wild- and farmed sample grouping information, was imported into SIMCA software (MKSUmetrics AB) to construct a multivariate analysis model. Before OPLS-DA modeling, the data underwent Pareto scaling preprocessing to enhance the model's stability and interpretability.
[0069] The constructed OPLS-DA classification and discriminant analysis model is shown in the figure. Figure 4 (Each point in the figure represents a standard sample).
[0070] The score plot output by the model is used for discriminant analysis of subsequent sample data. The sample region represented by the location where subsequent data is projected onto the score plot is predicted as the growth pattern of the test sample. R 2 X (cum) R 2 Y (cum) and Q 2 The quality of the OPLS-DA model was evaluated using parameters such as (cum). The robustness of the model was assessed using 200 permutation tests. Figure 5As shown, differentially expressed proteins were screened by comparing OPLS-DA models of wild and domesticated animals. The screening parameters were set as follows: VIP>1; in S-plot analysis, |p[corr]|>0.75; the confidence interval of the load plot knife-cut method did not include the zero point; in pairwise comparisons of OPLS-DA, the parameter change factor was >1.4 (or <0.71) and p<0.05.
[0071] The polypeptide biomarkers of grass carp include SEQ ID No. 1-42, and the constructed OPLS-DA model... R 2 X (cum) R 2 Y (cum) Q 2 The cum values were 0.562, 0.993, and 0.903, respectively, indicating that the model has good fitting and predictive abilities. The evaluation results of 200 permutation tests show that... R 2 Y and Q 2 The intercepts of the regression lines are 0.926 and -0.695, respectively, and the slope of the regression lines is close to 0, indicating that the model has good robustness.
[0072] Example 5: Sample processing and detection steps for grass carp to be tested Three wild grass carp samples were taken for testing, and the processing and analysis steps were the same as those for the standard sample processing described in Example 1.
[0073] Using the OPLS-DA classification and discriminant analysis model for wild and farmed grass carp samples constructed in this invention, the proteomic component responses of the test samples are input into the model for predictive analysis, and the species and growth mode of the test samples are automatically calculated by the model.
[0074] The results of the identification are as follows Figure 6 As shown, the data points of the tested sample are all clustered in the wild sample section on the right, indicating that it belongs to the wild-caught farming method, and the model has good robustness. Therefore, the tested sample can be classified and analyzed using the model, and the analysis results are accurate.
[0075] The above description is merely an embodiment of the present invention, and the scope of protection of the present invention is not limited to these specific embodiments, but is determined by the claims of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the technical concept and principle of the present invention should be included within the scope of protection of the present invention.
Claims
1. A system for discriminating between wild and farmed origin of bluefish, comprising, The system at least comprises: a sample processing module: the sample processing module is used at least for processing grass carp samples to obtain raw data of polypeptides as shown in any one of SEQ ID NO. 1-42, and obtain protein group component response of the sample to be tested; a data processing module: the data processing module is used at least for importing the protein group component response of the sample to be tested into an OPLS-DA classification discriminant analysis model, obtaining a score plot for subsequent discriminant analysis of sample data; a result output module: the result output module is used at least for outputting results.
2. The system of claim 1, wherein, The specific operation of the sample processing module is as follows: performing protein group determination pretreatment on the sample to be tested to obtain a polypeptide solution comprising at least the polypeptide markers as shown in any one of SEQ ID NO. 1-42, and determining the obtained polypeptide solution by liquid chromatography-mass spectrometry to obtain the protein group component response of the sample to be tested.
3. The system of claim 1, wherein, The preparation method of the OPLS-DA classification discriminant analysis model in the data processing module is as follows: S1: protein extraction step: performing protein group determination pretreatment on muscle powder of wild and domestic grass carp samples to obtain a polypeptide solution comprising polypeptides as shown in SEQ ID NO. 1-42: S2: protein detection step: determining the polypeptide solution obtained in step S1 by liquid chromatography-mass spectrometry to obtain liquid chromatography-mass spectrometry analysis results; S3: importing the liquid chromatography-mass spectrometry analysis results in step S2 into PEAKS Online software for database retrieval and relative quantification analysis, and the parameters are set as follows: the database is Cyprinidae (Cyprinidae) database; the enzyme cutting mode is Trypsin (trypsin), the maximum number of missed cutting sites is 2; the precursor ion mass error is 10ppm, the fragment ion mass error is 0.02Da; the false positive rate (FDR) threshold at the peptide and protein levels is set to 1%; S4: taking the protein identification list exported by the PEAKS software obtained in step S3 and the normalized expression of each protein in all samples as an initial data matrix, importing the matrix and the corresponding grouping information of wild and domestic samples into SIMCA software to obtain an OPLS-DA classification discriminant analysis model, wherein the differential proteins are screened by comparing the OPLS-DA models of wild and domestic samples, and the screening parameters are set as follows: VIP>1; |p[corr]|>0.75 in S-plot analysis; the zero point is not included in the confidence interval of the loading chart; in the pairwise comparison of OPLS-DA, the parameter change multiple is >1.4 (or <0.71), and p<0.
05.
4. The system of claim 1, wherein, The steps of discriminant analysis are as follows: using the score plot output by the model to perform discriminant analysis of subsequent sample data, and the position where the subsequent data is projected in the score plot represents the sample area of the test sample, which is predicted as the growth mode.
5. A method for discriminating between wild and farmed origin of Mylopharyngodon piceus, characterized by, The method comprises the steps of applying the system of any one of claims 1-4 to analyze and determine the source of grass carp.
6. A kit for discriminating between wild and farmed origin of Mylopharyngodon piceus, characterized by, The kit at least comprises reagents for quantitative and / or qualitative detection of the polypeptide markers as shown in any one of SEQ ID NO. 1-42.
7. The kit of claim 6, wherein The kit at least comprises at least one of a protein extraction reagent and a protein detection reagent.
8. The kit of claim 7, wherein The protein extraction reagent at least comprises at least one of a protein extraction solution, DTT (dithiothreitol), IAA (iodoacetamide), an ammonium bicarbonate solution, and a protease solution; and the protein detection reagent at least comprises a mobile phase solution.
9. Use of the system of any one of claims 1-4 or the method of claim 5 or the kit of any one of claims 6-8 in the preparation of a product for identifying wild or farmed origin of grass carp.
10. Use according to claim 9, characterized in that, The identification of wild or farmed origin of grass carp comprises at least one of wild farming in the Huaihe River, wild farming in the Qili Lake, and domestication.