Silkworm cocoon source discrimination and quality evaluation method and system based on molecular marker analysis
By combining alkaline degumming treatment, liquid chromatography-mass spectrometry, and deep learning models, the problems of subjectivity and reproducibility in silkworm cocoon origin identification and quality assessment have been solved, enabling accurate traceability and quantitative evaluation of silkworm cocoon origin.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies suffer from strong subjectivity, poor reproducibility, and difficulty in achieving high precision and sensitivity in silkworm cocoon origin identification and quality assessment. In particular, they lack objective and quantitative criteria when dealing with subtle differences in intrinsic quality caused by variations in rearing environment and mulberry leaf variety.
The degumming rate was determined by alkaline degumming treatment, and molecular marker analysis was performed using liquid chromatography-mass spectrometry and a deep learning model. A standard curve was established by combining the external standard method, so as to achieve accurate traceability of the source of silkworm cocoons and quantitative evaluation of their intrinsic quality.
It enables precise traceability of silkworm cocoon origin and objective, quantitative evaluation of intrinsic quality, thereby improving the accuracy and precision of silkworm cocoon identification.
Smart Images

Figure CN121633358A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of cocoon detection, and particularly relates to a cocoon source discrimination and quality evaluation method and system based on molecular marker analysis. BACKGROUND
[0002] Cocoon is the basic raw material of the silk industry, and its origin and internal quality directly affect the quality of raw silk and the value of textiles. At present, the origin discrimination and quality evaluation of cocoon mainly rely on sensory inspection, physical index testing (such as unwinding rate and cleanliness) or single chemical component analysis. These traditional methods have obvious limitations: sensory inspection is highly subjective and has poor reproducibility; physical index testing is a post-evaluation and cannot achieve effective sorting and quality prediction at the front end of processing; and conventional chemical analysis is difficult to cope with the synergistic effect of multiple markers in the complex cocoon matrix, and lacks comprehensive discrimination ability with high precision and high sensitivity. In particular, for subtle internal quality differences caused by differences in feeding environment, mulberry leaf varieties, etc., existing technical means often cannot provide objective and quantitative basis for differentiation. SUMMARY
[0003] The purpose of the present application is to provide a cocoon source discrimination and quality evaluation method and system based on molecular marker analysis, in order to solve the problems in the prior art and achieve accurate tracing of cocoon origin and objective and quantitative evaluation of internal quality.
[0004] One embodiment of the present application provides a cocoon source discrimination and quality evaluation method based on molecular marker analysis, which comprises: Alkaline degumming treatment is performed on the cocoon sample, the degumming rate of the cocoon sample at the initial degumming stage is determined, and preliminary discrimination of cocoon from different feeding sources is realized based on the degumming rate difference; The cocoon sample is extracted by an organic solvent to prepare a sample solution for liquid chromatography-mass spectrometry analysis; Liquid chromatography-mass spectrometry technology is used to separate and detect the sample solution, and chromatography-mass spectrometry data of the cocoon sample are obtained; A deep learning model is used to extract features and learn similarity of the chromatography-mass spectrometry data, and when the feature similarity score is higher than a preset confidence threshold, the target compound is confirmed, and qualitative analysis of the molecular marker is completed; Based on the qualitative analysis result, a standard curve is established by an external standard method, the content of the target compound in the cocoon is calculated according to the peak area of the target compound, quantitative analysis of the molecular marker and comprehensive evaluation of the cocoon quality are realized.
[0005] Optionally, the alkaline degumming treatment of the cocoon sample, the determination of the degumming rate of the cocoon sample at the initial degumming stage, and the preliminary discrimination of cocoon from different feeding sources based on the degumming rate difference comprise: Accurately weigh the initial weight of the dried silkworm cocoon sample and record it as the cocoon weight data before degumming; The weighed silkworm cocoon sample was placed in a boiling 0.02 mol / L sodium carbonate solution and subjected to the first degumming treatment for 1 minute at a weight-to-volume ratio of 1:400. Take out the silkworm cocoon sample after the first degumming, wash it three times with pure water, and dry it in a 60℃ oven. Record its weight as the cocoon weight data after 1 minute of degumming. The same silkworm cocoon sample was degummed for another 10 minutes, and after washing and drying again, the weight was recorded as the cocoon weight data after 10 minutes of degumming. Based on the cocoon weight data before degumming, the cocoon weight data 1 minute after degumming, and the cocoon weight data 10 minutes after degumming, calculate the degumming rate value within the initial 1 minute. The calculated degumming rate is compared with a preset rate threshold, and a preliminary judgment on the origin of the silkworm cocoons is output based on the comparison results.
[0006] Optionally, the step of extracting silkworm cocoon samples with organic solvents to prepare a sample solution for liquid chromatography-mass spectrometry analysis includes: The silkworm cocoon sample was cut into uniform fragments to obtain pretreated silkworm cocoon material; Accurately weigh 0.5 grams of pretreated silkworm cocoon material and add 5 ml of 75% ethanol aqueous solution at a weight-to-volume ratio of 1:10; The mixed solution was subjected to ultrasonic-assisted extraction for 30 minutes to generate a primary extract; The primary extract was heated at 60°C for 4 hours, and then allowed to stand overnight to form a layered solution. The supernatant solution was filtered through a 0.22-micron filter membrane to obtain a clear and transparent sample solution.
[0007] Optionally, the step of using liquid chromatography-mass spectrometry (LC-MS) to separate and detect the sample solution to obtain chromatographic-mass spectrometric data of the silkworm cocoon sample includes: The sample solution to be tested is injected into an ultra-high performance liquid chromatography system, and a separation channel is established by equipping a C18 reversed-phase column. Start the gradient elution program with mobile phase A being an aqueous solution containing 0.1% formic acid and mobile phase B being acetonitrile, and maintain a constant flow rate of 0.400 mL / min. The ultraviolet absorption chromatogram was obtained by scanning at a wavelength of 254 nm using a photodiode array detector; The chromatographic effluent was introduced into a single quadrupole mass spectrometer and scanned in negative ion electrospray ionization mode in the mass-to-charge ratio range of 50-1000 Da. Simultaneously record total ion current chromatograms and mass spectrometry fragment information to generate a complete chromatogram-mass spectrometry dataset.
[0008] Optionally, the step of using a deep learning model to extract features and learn similarity from the chromatographic-mass spectrometry data, and confirming the target compound when the feature similarity score is higher than a preset reliability threshold, thereby completing the qualitative analysis of the molecular marker, includes: Load the standard database, which contains retention time, mass-to-charge ratio, and characteristic fragment ion information of the target molecule markers; The chromatography-mass spectrometry dataset is input into a pre-trained deep learning model for multi-dimensional feature extraction and dimensionality reduction. Calculate the cosine similarity between the feature vector of the sample and the feature vector of the standard, and generate a feature similarity score; The feature similarity score is compared with a preset reliability threshold to output the compound confirmation result; Integrate all confirmation results to generate a qualitative analysis report of molecular markers.
[0009] Optionally, the step of establishing a standard curve based on the qualitative analysis results using the external standard method, and calculating the content of the target compound in silkworm cocoons based on its peak area, thereby achieving quantitative analysis of molecular markers and comprehensive evaluation of silkworm cocoon quality, includes: Based on the target molecular markers identified in the qualitative analysis report, a series of standard solutions with concentration gradients were prepared. The standard solution was analyzed by liquid chromatography-mass spectrometry, and the chromatographic peak area data corresponding to each concentration was recorded. A standard curve is generated by fitting a linear regression equation with the concentration of the standard as the x-axis and the chromatographic peak area as the y-axis. Substitute the chromatographic peak area of the target compound in the silkworm cocoon sample into the standard curve equation to calculate its concentration value in the actual sample. By combining the quantitative results of various molecular markers with principal component analysis to perform multidimensional data fusion, a comprehensive evaluation report on silkworm cocoon quality is generated.
[0010] Another embodiment of this application provides a system for identifying the origin and evaluating the quality of silkworm cocoons based on molecular marker analysis, the system comprising: The degumming module is used to perform alkaline degumming treatment on silkworm cocoon samples, measure the degumming rate in the initial degumming stage, and make preliminary distinctions between silkworm cocoons from different feeding sources based on the difference in degumming rate. The preparation module is used to extract silkworm cocoon samples with organic solvents to prepare a sample solution for liquid chromatography-mass spectrometry analysis. The detection module is used to separate and detect the sample solution to be tested using liquid chromatography-mass spectrometry (LC-MS) to obtain chromatographic-mass spectrometry data of the silkworm cocoon sample. The qualitative module is used to extract features and learn similarity from the chromatographic-mass spectrometry data using a deep learning model. When the feature similarity score is higher than a preset reliability threshold, the target compound is confirmed, and the qualitative analysis of the molecular marker is completed. The quantitative module is used to establish a standard curve based on the qualitative analysis results using the external standard method, and to calculate the content of the target compound in the silkworm cocoon according to the peak area, so as to realize the quantitative analysis of molecular markers and the comprehensive evaluation of silkworm cocoon quality.
[0011] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.
[0012] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.
[0013] Compared with existing technologies, this invention provides a method for identifying the origin and evaluating the quality of silkworm cocoons based on molecular marker analysis. The method involves degumming silkworm cocoon samples using an alkaline method, and using differences in degumming rates to initially identify cocoons from different feeding sources. The cocoon samples are then extracted using organic solvents to prepare a sample solution for liquid chromatography-mass spectrometry (LC-MS). LC-MS is used to separate and detect the sample solution, obtaining LC-MS data of the silkworm cocoon samples. A deep learning model is employed to extract features and learn similarity from the LC-MS data, completing the qualitative analysis of molecular markers. Based on the qualitative analysis results, a standard curve is established using the external standard method, and the content of the target compound in the cocoon is calculated based on the peak area. This enables precise traceability of the silkworm cocoon's origin and objective, quantitative evaluation of its intrinsic quality. Attached Figure Description
[0014] Figure 1 A hardware structure block diagram of a computer terminal for a method of silkworm cocoon origin identification and quality evaluation based on molecular marker analysis provided in an embodiment of the present invention; Figure 2 A flowchart illustrating a method for identifying the origin and evaluating the quality of silkworm cocoons based on molecular marker analysis, provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a silkworm cocoon origin identification and quality evaluation system based on molecular marker analysis, provided in an embodiment of the present invention. Detailed Implementation
[0015] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0016] This invention first provides a method for identifying the origin and evaluating the quality of silkworm cocoons based on molecular marker analysis. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers.
[0017] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware block diagram of a computer terminal for a method of silkworm cocoon origin identification and quality evaluation based on molecular marker analysis, provided as an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0018] The non-volatile storage medium can store an operating system and a computer program. This computer program includes program instructions that, when executed, cause the processor to perform any method for identifying the origin and evaluating the quality of silkworm cocoons based on molecular marker analysis.
[0019] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0020] The internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any method for identifying the origin and quality evaluation of silkworm cocoons based on molecular marker analysis.
[0021] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0022] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0023] See Figure 2 The embodiments of the present invention provide a method for identifying the origin and evaluating the quality of silkworm cocoons based on molecular marker analysis, which may include the following steps: S201, alkaline degumming treatment was performed on silkworm cocoon samples, and the degumming rate in the initial degumming stage was measured. Based on the difference in degumming rate, a preliminary distinction was made between silkworm cocoons from different feeding sources. Specifically, the initial weight of the dried silkworm cocoon sample can be accurately weighed and recorded as the cocoon weight data before degumming; This step is the basic data acquisition stage for the degumming rate determination. The core is to ensure the accuracy of the silkworm cocoon sample weight measurement, providing reliable initial data support for subsequent degumming rate calculations. The key lies in controlling the sample drying state, weighing environment, and instrument accuracy to avoid errors caused by external factors. The specific implementation method is as follows: First, the silkworm cocoon samples need to be pre-treated and dried to ensure they are free of free moisture—moisture directly affects the accuracy of weight measurements and may interfere with the hydrolysis and shedding efficiency of sericin during degumming. The silkworm cocoon samples to be tested are placed in a clean glass petri dish and dried in an oven set at 60℃ for 2 hours. This temperature effectively removes moisture without damaging the sericin protein structure of the cocoons (the thermal denaturation temperature of sericin protein is above 80℃). After drying, the samples are removed and placed in a desiccator to cool at room temperature for 30 minutes to prevent the samples from absorbing moisture from the air, which could lead to weight changes.
[0024] The weighing process must be conducted in a constant temperature and humidity laboratory, with the ambient temperature controlled at 25℃±2℃ and relative humidity at 50%±5%, to minimize the impact of temperature and humidity fluctuations on the weighing results. An analytical balance with an accuracy of 0.1 mg should be used. Calibration must be performed before weighing using standard weights (100g, 50g, and 10g) at three points to ensure the balance error is ≤±0.1 mg. Carefully place the cooled silkworm cocoon sample in the center of the weighing pan, avoiding direct contact with the sample by hand (sweat and oil from hands will contaminate the sample and add extra weight). Record the data after the balance reading has stabilized for 3 seconds, accurate to four decimal places.
[0025] For example, after drying and cooling, the pre-degummed weight of a certain feed cocoon sample was 0.2881 grams; the pre-degummed weight of a certain mulberry leaf cocoon sample was 0.3859 grams. These data must be recorded immediately in the experimental record sheet, along with the sample number, weighing date, and time, to ensure data traceability. To verify weighing repeatability, each sample must be weighed three times in parallel, and the average of the three measurements is taken as the final pre-degummed cocoon weight. If the relative standard deviation (RSD) of the three measurements is greater than 0.1%, the sample must be dried again and weighed again until the accuracy requirements are met.
[0026] The weighed silkworm cocoon sample was placed in a boiling 0.02 mol / L sodium carbonate solution and subjected to the first degumming treatment for 1 minute at a weight-to-volume ratio of 1:400. This step is the core operation for determining the degumming rate. It requires precise control of the concentration, temperature, solid-liquid ratio, and processing time of the degumming solution to ensure consistent degumming conditions for different samples, thereby guaranteeing the comparability of degumming rates. The specific implementation method is as follows: First, prepare a 0.02 mol / L sodium carbonate solution and calculate the required mass of sodium carbonate: Weigh 2.12 g of anhydrous sodium carbonate (sodium carbonate has a molar mass of 10⁶ g / mol, 0.02 mol / L × 1 L × 10⁶ g / mol = 2.12 g), place it in a 1 L volumetric flask, add approximately 800 mL of purified water, stir with a glass rod until completely dissolved, then dilute to 1 L with purified water and shake well before use. This concentration of sodium carbonate solution is weakly alkaline, which effectively promotes the hydrolysis and shedding of sericin without excessively corroding the fibroin structure of silk, meeting the mild conditions required for degumming experiments.
[0027] A 500 ml three-necked flask equipped with a reflux condenser was used for degumming to prevent the solution from evaporating and increasing the concentration during boiling. Based on the pre-degumming cocoon weight data, the required volume of sodium carbonate solution was calculated at a weight-to-volume ratio (w / v) of 1:400 – that is, 1 gram of cocoon sample corresponds to 400 ml of solution. For example, for a feed cocoon sample with a pre-degumming weight of 0.2881 g, approximately 0.2881 × 400 ≈ 115.24 ml of solution was added. In practice, 115 ml was used (error ≤ 0.5 ml) to ensure accurate solid-liquid ratio. The solution was poured into the three-necked flask and heated on a heating mantle. After the solution boiled vigorously (temperature stabilized at 100℃), the weighed cocoon sample was carefully placed into the solution, and a stopwatch was started simultaneously to begin the first degumming process.
[0028] During the degumming process, the solution must be kept boiling continuously. The temperature should be stabilized by adjusting the power of the heating mantle (maintaining a power of 500 watts) to avoid localized overheating or temperature fluctuations. Simultaneously, the solution should be gently stirred with a glass rod to ensure the silkworm cocoon sample is completely submerged, preventing it from floating on the surface and causing uneven degumming. The initial degumming process should be strictly controlled to 1 minute. The stopwatch should start at the moment the sample is completely submerged and end at the moment 60 seconds have elapsed. A designated person must be responsible for timing and operation to avoid time errors caused by operational delays.
[0029] Take out the silkworm cocoon sample after the first degumming, wash it three times with pure water, and dry it in a 60℃ oven. Record its weight as the cocoon weight data after 1 minute of degumming. The core of this step is to thoroughly remove residual sericin and sodium carbonate solution from the sample surface to avoid these residues affecting the accuracy of subsequent weight measurements. At the same time, a standardized drying process ensures consistent sample dryness. The specific implementation method is as follows: After the one-minute degumming timer ended, the silkworm cocoon sample was immediately removed from the boiling sodium carbonate solution using tweezers and quickly transferred to a beaker containing 50 ml of purified water for the first wash. During the wash, the sample was gently stirred with a glass rod for 30 seconds to ensure that the hydrolyzed sericin and alkaline solution adhering to the sample surface were fully dissolved in the water. The sample was then removed, and the wastewater in the beaker was discarded. The same procedure was repeated for the second and third washes using fresh 50 ml of purified water, each wash lasting 30 seconds. A total of 150 ml of purified water was used for the three washes to ensure that the sample surface was not slippery (a typical characteristic of sericin residue) and that the pH of the water after washing was close to 7 (neutral), proving that the residual sodium carbonate had been completely removed.
[0030] After cleaning, spread the silkworm cocoon samples evenly in clean glass petri dishes and place them in a preheated oven at 60°C for 2 hours. A tray containing silica gel desiccant is placed inside the oven to absorb water vapor generated during drying, maintaining a dry environment and preventing the samples from absorbing moisture. During drying, ensure sufficient space between the petri dishes to guarantee even hot air circulation and thorough drying. After drying, remove the samples along with the petri dishes and place them in a desiccator to cool at room temperature for 30 minutes. Once the sample temperature matches the ambient temperature, perform a weight measurement—directly weighing hot samples will result in a lower reading due to air convection and rapid moisture evaporation.
[0031] A calibrated analytical balance is still used for weighing, and the procedure is the same as for measuring the cocoon weight before degumming. Each sample is weighed three times in parallel, and the average value is taken as the cocoon weight data after 1 minute of degumming, recorded to four decimal places. For example, the cocoon weight data obtained after 1 minute of degumming for the above feed cocoon sample after washing and drying is 0.2326 grams. This data must be recorded in correspondence with the cocoon weight data before degumming to ensure that the data correlation is accurate.
[0032] The same silkworm cocoon sample was degummed for another 10 minutes, and after washing and drying again, the weight was recorded as the cocoon weight data after 10 minutes of degumming. The purpose of this step is to ensure that the sericin on the surface of the silkworm cocoon is completely removed, obtaining the pure weight of the silk fibroin, thus providing a denominator parameter for calculating the degumming rate. This operation requires continuing the previous degumming conditions to ensure thorough degumming. The specific implementation method is as follows: After the initial degumming for 1 minute, followed by washing and drying, the same silkworm cocoon sample is returned to the original three-necked flask (ensuring a consistent degumming environment). The remaining sodium carbonate solution in the flask needs to be replenished to the original volume (to bring the solution to 115 ml, as boiling and evaporation may reduce the solution). The sample is then reheated to a vigorous boil and degumming continues. Timing is accumulated from the initial degumming, and the total degumming time must be precisely controlled to 10 minutes, i.e., continue degumming for 9 minutes (10 minutes - 1 minute of degummed time). During the timing process, the solution is kept continuously boiling to avoid temperature fluctuations affecting the sericin removal efficiency.
[0033] After a total degumming time of 10 minutes, immediately remove the sample with tweezers and wash it three times with purified water (50 ml each time, stirring for 30 seconds) according to the cleaning procedure to thoroughly remove residual sericin and solution. After cleaning, spread the sample flat in a new glass petri dish, place it in a 60℃ oven to dry for 2 hours, and then cool it in a desiccator for 30 minutes. Measure the weight and record it as the cocoon weight 10 minutes after degumming. Repeat this process three times and take the average value, accurate to four decimal places.
[0034] Experimental verification shows that after boiling silkworm cocoons in a 0.02 mol / L sodium carbonate solution for 10 minutes, the sericin removal rate can reach over 99%, and the remaining weight is the pure weight of the fibroin. Further extending the degumming time at this point does not significantly change the sample weight (change ≤ 0.1 mg). Therefore, the weight after 10 minutes of degumming can be used as the baseline weight of the fibroin to calculate the degumming rate within the initial minute, reflecting the initial hydrolysis and shedding characteristics of sericin. For example, the weight of the feed cocoon sample after 10 minutes of degumming was 0.2135 g, providing a key denominator parameter for subsequent degumming rate calculations.
[0035] Based on the cocoon weight data before degumming, the cocoon weight data 1 minute after degumming, and the cocoon weight data 10 minutes after degumming, calculate the degumming rate value within the initial 1 minute. This step is the core calculation process for determining the degumming rate. It requires strict adherence to the preset formula for quantitative calculation to ensure the accuracy and standardization of the calculation process. Simultaneously, the physical meaning of each parameter must be clearly defined. The specific implementation method is as follows: The formula for calculating the degumming rate is: Degumming rate = (cocoon weight before degumming - cocoon weight after 1 minute of degumming) / (cocoon weight before degumming - cocoon weight after 10 minutes of degumming) × 100%. In the formula, the numerator "cocoon weight before degumming - cocoon weight after 1 minute of degumming" represents the weight of sericin detached in the initial 1 minute, and the denominator "cocoon weight before degumming - cocoon weight after 10 minutes of degumming" represents the total weight of sericin contained in the cocoon. The ratio of the two multiplied by 100% is the proportion of sericin detached in the initial 1 minute, which directly reflects the ease of hydrolysis of sericin protein—the higher the ratio, the easier it is for sericin to hydrolyze and detach, and vice versa.
[0036] Taking a sample of feed cocoons as an example, the weight of the cocoons before degumming is 0.2881 grams, the weight after 1 minute of degumming is 0.2326 grams, and the weight after 10 minutes of degumming is 0.2135 grams. Substituting these values into the formula, we can calculate: Numerator = 0.2881 - 0.2326 = 0.0555 grams, Denominator = 0.2881 - 0.2135 = 0.0746 grams, Degumming rate = (0.0555 / 0.0746) × 100% ≈ 74.40%. Four significant figures should be retained during the calculation, and the final result should be rounded to two decimal places to ensure data accuracy.
[0037] For mulberry leaf cocoon samples, the same formula is used for calculation. For example, if a mulberry leaf cocoon sample weighs 0.3859 grams before degumming, 0.3198 grams after 1 minute of degumming, and 0.2760 grams after 10 minutes of degumming, the numerator = 0.3859 - 0.3198 = 0.0661 grams, the denominator = 0.3859 - 0.2760 = 0.1099 grams, and the degumming rate = (0.0661 / 0.1099) × 100% ≈ 60.15%. After calculation, the results need to be verified for reasonableness. The degumming rate should be within the range of 0%-100%. If it exceeds this range, it indicates an error in the weighing or operation process, and the experiment needs to be repeated. Simultaneously, the average degumming rate and relative standard deviation (RSD) of parallel experiments need to be calculated for each sample. The RSD should be ≤5% to ensure the repeatability of the experimental results.
[0038] The calculated degumming rate is compared with a preset rate threshold, and a preliminary judgment on the origin of the silkworm cocoons is output based on the comparison results.
[0039] This step is the application stage of degumming rate measurement. It requires determining a reasonable discrimination threshold based on a large amount of experimental data. By comparing the thresholds, a preliminary distinction can be made between silkworm cocoons from different feeding sources. At the same time, the accuracy and applicable scope of the discrimination are clarified. The specific implementation method is as follows: The predetermined rate threshold was determined based on statistical analysis. By statistically analyzing the degumming rate data of 9 feed cocoon samples and 19 mulberry leaf cocoon samples, the average degumming rate of feed cocoons was found to be 66.87%, and the average degumming rate of mulberry leaf cocoons was 62.58%, showing a significant difference between the two groups (t-test, P < 0.05). Based on the intersection of the two groups and combined with accuracy verification, the predetermined rate threshold was determined to be 65%—that is, when the degumming rate is greater than 65%, the cocoons are initially identified as those from silkworms fed with artificial feed (feed cocoons); when the degumming rate is less than or equal to 65%, the cocoons are initially identified as those from silkworms fed with mulberry leaves (mulberry leaf cocoons).
[0040] For example, the degumming rate of the feed cocoon sample was 74.40%, which is greater than 65%, and it was initially identified as a feed cocoon; the degumming rate of the mulberry leaf cocoon sample was 60.15%, which is less than 65%, and it was initially identified as a mulberry leaf cocoon. To verify the discrimination effect of this threshold, all experimental samples were retrospectively verified: among the 9 feed cocoon samples, 7 samples had a degumming rate greater than 65%, with a discrimination accuracy of 77.8%; among the 19 mulberry leaf cocoon samples, 14 samples had a degumming rate less than or equal to 65%, with a discrimination accuracy of 73.7%. The overall discrimination effect was good and could meet the needs of preliminary screening.
[0041] The output of the discrimination conclusion should include the sample number, degumming rate value, threshold comparison result, and preliminary source identification, for example, "Sample number SL-03, degumming rate 60.15%, less than the threshold of 65%, preliminarily identified as silkworm cocoons fed with mulberry leaves." It should also be noted that this discrimination is a preliminary screening result and has a certain probability of false positives, which may stem from individual sample differences, weighing errors, or deviations in the degumming operation. For samples with questionable discrimination results (degumming rates close to 65% ± 1%), secondary verification using subsequent liquid chromatography-mass spectrometry (LC-MS) analysis of molecular markers is required to ensure the accuracy of the final discrimination conclusion. Furthermore, all sample degumming rate data and discrimination results should be compiled and archived to accumulate data for building a more accurate discrimination model in the future.
[0042] S202, extracting silkworm cocoon samples with organic solvents to prepare a sample solution for liquid chromatography-mass spectrometry analysis; Specifically, silkworm cocoon samples can be cut into uniform fragments to obtain pretreated silkworm cocoon materials; This step creates favorable conditions for subsequent extraction operations. The core is to increase the contact area between the silkworm cocoon material and the extraction solvent by physically breaking it down, thus promoting the full dissolution of molecular markers (such as flavonoids and isoflavones) while ensuring fragment uniformity to reduce extraction errors. The specific implementation method is as follows: The main components of silkworm cocoon samples are fibroin and sericin. At this stage, the sample is white and has a tough, filamentous structure. Direct extraction will result in insufficient dissolution of markers due to the small contact area. Therefore, sterile medical scissors must be used to cut the cocoons into small pieces. The scissors must be disinfected with 75% ethanol beforehand to avoid introducing external impurities (such as oil or microorganisms) that could affect subsequent testing. During the cutting process, the fragment size must be controlled to ensure that all fragments are short fiber segments of 1-2 mm, with a length deviation not exceeding 0.5 mm. This size ensures that the solvent can fully penetrate the fiber and facilitates subsequent transfer and mixing. If the fragments are too long (more than 3 mm), the fibers will tangle during extraction, preventing uniform solvent contact; if they are too short (less than 0.5 mm), fine fiber debris will easily form, increasing the difficulty of subsequent filtration.
[0043] The shredding process must be performed on a clean glass workbench lined with disposable sterile filter paper to prevent direct sample contact and contamination. The shredded fragments must be sieved through an 80-mesh nylon sieve to remove any excessively long or short fragments, ensuring the uniformity of the pretreated cocoon material. Discard any excessively long fragments (typically ≤5%) that remain on the sieve and any small debris (≤3%) that falls through, retaining only 1-2 mm fragments as the final pretreatment material. The pretreated material must be immediately transferred to a clean weighing bottle and temporarily stored in a desiccator to prevent prolonged exposure to air for moisture absorption or dust adsorption, which could affect subsequent weighing accuracy and extraction results.
[0044] Accurately weigh 0.5 grams of pretreated silkworm cocoon material and add 5 ml of 75% ethanol aqueous solution at a weight-to-volume ratio of 1:10; This step is crucial for constructing the extraction system, requiring precise control of material ratios and solvent concentrations to ensure the dissolution efficiency of molecular markers while avoiding fluctuations in extraction yield due to ratio deviations. The specific implementation method is as follows: Weighing the pretreated silkworm cocoon material requires an analytical balance with an accuracy of 0.1 mg. Before weighing, a three-point calibration must be performed using standard weights (100 g, 50 g, and 10 g) to ensure that the weighing error is ≤ ±0.1 mg. During weighing, slowly add the pretreated material to a clean 50 ml centrifuge tube. Record the weight after the balance reading stabilizes for 3 seconds. If the deviation of a single weighing exceeds ±0.005 g (i.e., 0.5000 ± 0.0050 g), it must be weighed again until the accuracy requirement is met—this accuracy ensures the accuracy of subsequent volume proportioning.
[0045] A 1:10 weight-to-volume ratio (w / v) means that 1 gram of pretreated material corresponds to 10 milliliters of extraction solvent. Therefore, 0.5 grams of material requires 0.5 × 10 = 5 milliliters of solvent. This ratio was determined through experiments: when the ratio is lower than 1:8, insufficient solvent leads to incomplete dissolution of the biomarker; when the ratio is higher than 1:12, excessive solvent will dilute the biomarker concentration and increase the difficulty of subsequent detection. At a 1:10 ratio, the extraction efficiency can reach over 92%, and the biomarker concentration is suitable for detection.
[0046] The choice of a 75% ethanol aqueous solution is based on the solubility characteristics of molecular markers: flavonoids (such as quercetin and isoquercitrin) and isoflavones (such as daidzein and genistein) have the highest solubility in 75% ethanol. This avoids the shrinkage of silk fibers caused by anhydrous ethanol (affecting the dissolution of internal components) and prevents low-concentration ethanol (such as below 50%) from dissolving too many water-soluble impurities (such as small-molecule sugars). The 75% ethanol aqueous solution needs to be prepared in advance by mixing anhydrous ethanol and purified water at a volume ratio of 75:25. For example, measure 75 ml of anhydrous ethanol and 25 ml of purified water, pour them into a 100 ml volumetric flask, and shake well. Accurate measurement using a pipette is required during the preparation process to ensure a concentration deviation of ≤±1%.
[0047] When adding the solvent, use a 5 mL single-label pipette. The pipette should be rinsed three times with 75% ethanol (2 mL each time) beforehand to avoid residual moisture on the inner wall of the pipette diluting the solvent concentration. Slowly release the solvent by holding the tip of the pipette close to the inner wall of the centrifuge tube, avoiding impact on the sample and causing fragment splashing. After adding the solvent, tighten the cap of the centrifuge tube and gently invert and shake five times to ensure that the pretreated material and solvent come into full contact and form a homogeneous mixture.
[0048] The mixed solution was subjected to ultrasonic-assisted extraction for 30 minutes to generate a primary extract; This step enhances the extraction effect through ultrasound technology. The core principle is to utilize the ultrasonic cavitation effect to disrupt the microstructure of silk fibroin fibers, promoting the dissolution of molecular markers from the fiber interior into the solvent. The specific implementation method is as follows: The principle of ultrasound-assisted extraction is that when ultrasound waves propagate in a liquid, they generate periodic pressure changes, forming a large number of tiny bubbles (cavitation bubbles). The energy released when these bubbles burst can disrupt the dense structure of silk fibroin fibers, making it easier for the solvent to penetrate into the fiber interior. Simultaneously, it accelerates the diffusion of marker molecules, improving dissolution efficiency. Experimental verification shows that compared to traditional oscillation extraction, ultrasonic extraction can increase the amount of markers extracted by 30%-40% and shorten the extraction time by more than 60%.
[0049] The equipment parameters for ultrasonic extraction must be strictly controlled: the ultrasonic power should be set to 300-500W (too low a power will result in an insignificant cavitation effect, while too high a power will easily lead to solvent overheating), the ultrasonic frequency should be 40kHz (at this frequency, the cavitation effect and penetrating power are balanced, and it can act on all mixed solutions in the centrifuge tube), and the extraction time should be precisely controlled to 30 minutes – set using the equipment's built-in timer, and observe the state of the mixed solution every 10 minutes after starting the ultrasound to avoid time deviations due to equipment malfunction.
[0050] The mixed solution should be placed in a 50 mL polypropylene centrifuge tube for sonication. The centrifuge tube must be fixed on the sample rack of the sonicator to ensure that the solution inside the tube is completely within the ultrasonic field (avoid tilting the centrifuge tube so that some solution is not subjected to ultrasound). The solution temperature needs to be controlled during sonication. Since ultrasonic energy is converted into heat energy, the water bath temperature of the sonicator should be set to 25°C, and the temperature of the mixed solution should be maintained between 25-35°C by circulating water cooling. Too high a temperature (above 40°C) may cause degradation of some heat-sensitive markers (such as isoquercitrin), while too low a temperature will reduce the extraction efficiency.
[0051] After ultrasound treatment, the mixed solution is the primary extract. At this stage, the solution is pale yellow and slightly turbid. The turbidity originates from a small amount of incompletely dissolved silk fibroin fragments and tiny impurities. Most of the target molecular markers (such as quercetin in mulberry leaf cocoons and daidzein in feed cocoons) have been dissolved. The primary extract must be removed from the ultrasound machine immediately to avoid prolonged standing, which could cause impurities to settle and clog subsequent processing equipment.
[0052] The primary extract was heated at 60°C for 4 hours, and then allowed to stand overnight to form a layered solution. This step complements ultrasonic extraction and aims to promote the full dissolution of poorly soluble molecular markers. Simultaneously, it allows for the separation of impurities from the extract through settling. The specific implementation method is as follows: The core function of heat extraction is to utilize heat energy to increase the mobility of solvent molecules, breaking down the binding forces between poorly soluble markers (such as quercetin) and silk fibers, allowing them to fully dissolve in the ethanol solution. The heating temperature of 60℃ was experimentally determined: at this temperature, the solubility of 75% ethanol is significantly improved, and the stability of all target markers (flavonoids and isoflavones) is good, with a degradation rate of ≤2% after 4 hours of heating. If the temperature is higher than 70℃, it will lead to increased ethanol evaporation (change in solvent concentration) and may also trigger the oxidation of some markers; if the temperature is lower than 50℃, the poorly soluble markers will not dissolve sufficiently, and the extraction efficiency will be less than 85%.
[0053] Heating is performed using a constant-temperature water bath. The temperature fluctuation of the water bath must be controlled within ±1℃ to avoid localized overheating. Place 50 ml centrifuge tubes containing the primary extract into the water bath, ensuring the liquid level in the tube is below the water bath level (to prevent uneven heating). Slightly loosen the caps of the centrifuge tubes (leaving a small gap to balance the pressure inside and outside the tube, preventing excessive pressure from causing overflow). During heating, gently shake the centrifuge tubes once every hour to ensure even heating and promote dissolution of components. Avoid splashing or introducing air bubbles when shaking.
[0054] After 4 hours of heating and extraction, turn off the water bath, remove the centrifuge tubes, tighten the caps, and let them stand overnight at room temperature (25℃±2℃). The standing time is typically 12 hours (from 18:00 on the same day to 6:00 the next day). Sufficient standing time allows impurities in the solution (such as undissolved silk fibroin fragments and tiny fiber particles) to settle fully under gravity, forming a two-layered solution. The resulting solution exhibits a distinct two-layer structure: the upper layer is a pale yellow, clear liquid containing dissolved molecular markers; the lower layer is a white or pale yellow precipitate, mainly composed of undissolved silk fibroin residue and a small amount of insoluble impurities. The precipitate thickness is typically 0.1-0.2 cm (accounting for 1%-2% of the centrifuge tube volume). This stratified state provides a clear operational basis for subsequent separation of the supernatant.
[0055] The supernatant solution was filtered through a 0.22-micron filter membrane to obtain a clear and transparent sample solution.
[0056] This step is crucial for removing impurities from the extract, aiming to obtain a pure extract and prevent impurities from clogging the liquid chromatography column or interfering with mass spectrometry detection. The specific implementation method is as follows: The selection of a 0.22-micron filter membrane is based on impurity particle size and column protection requirements: Tiny impurities (such as silk fibroin fragments with a particle size of 0.5-1 micron) that may be present in the supernatant of the layered solution, if they enter the chromatographic system, will clog the porous packing material of the column (pore size is typically 1.7-5 microns), affecting separation efficiency. A 0.22-micron filter membrane can effectively retain all impurities larger than this particle size, while allowing small molecular markers (particle size much smaller than 0.22 microns) to pass through smoothly. The filter membrane material is Nylon 66, which has good compatibility with 75% ethanol (organic phase), will not swell or release harmful substances, and avoids contamination of the extract.
[0057] The filtration operation uses a syringe filter. The specific procedure is as follows: First, take a 10 ml disposable syringe and draw 5 ml of the supernatant of the layered solution (when drawing, the syringe needle should be close to the inner wall of the centrifuge tube to avoid inserting into the lower precipitate; if precipitate is accidentally aspirated, it must be drawn again); put the 0.22 micron nylon filter membrane into the filter and tighten the connection between the filter and the syringe; slowly push the syringe plunger so that the supernatant passes through the filter membrane at a rate of 0.5-1 ml / min. The pushing process should be evenly applied to avoid the filter membrane breaking or generating air bubbles due to excessive speed (air bubbles will affect the baseline stability of subsequent chromatographic detection).
[0058] The filtered solution is the test sample solution. At this stage, the solution is pale yellow, clear, and transparent, without any visible impurities or turbidity. If the solution is still turbid, it indicates that the filter membrane may be damaged or contain too many impurities, and a new filter membrane needs to be replaced and filtered again. The test sample solution must be immediately transferred to a 2 mL injection bottle. The injection bottle must be sterilized by drying at 120°C and cooled to room temperature in advance to avoid contamination. To prevent degradation of the marker, the test sample solution must be stored at 4°C and liquid chromatography-mass spectrometry analysis must be completed within 24 hours to ensure the accuracy of the detection results.
[0059] S203, the sample solution to be tested is separated and detected by liquid chromatography-mass spectrometry to obtain the chromatographic-mass spectrometry data of the silkworm cocoon sample; Specifically, the sample solution to be tested can be injected into an ultra-high performance liquid chromatography system, equipped with a C18 reversed-phase column to establish a separation channel; This step is fundamental to achieving efficient separation of complex components in silkworm cocoon samples. The core principle is to utilize the high resolution of ultra-high performance liquid chromatography (UPLC) and the selectivity of the C18 reversed-phase column to provide clear single-component eluates for subsequent mass spectrometry analysis, avoiding interference from overlapping components in qualitative and quantitative analysis. The specific implementation method is as follows: The core advantages of ultra-high performance liquid chromatography (UHPLC) systems lie in their high column efficiency and rapid separation. Their pumps must possess high-pressure output capabilities (maximum pressure ≥100 MPa) to ensure efficient transfer of the mobile phase within the fine-particle column. The injection of the sample solution is accomplished using an autosampler. Before injection, the syringe needle undergoes a triple cleaning: first, it is rinsed three times with 500 μL of 75% ethanol aqueous solution, then rinsed three times with the sample solution (20 μL each time) to prevent cross-contamination from residual sample or solvent. The injection volume is set to 2 μL. This volume ensures sufficient detection sensitivity (meeting the detection requirements of trace molecular markers such as quercetin) without causing peak broadening due to excessive volume, which would affect the separation effect.
[0060] The selection of a C18 reversed-phase column should match the separation requirements. A specification of 1.7 μm particle size, 50 mm column length, and 2.1 mm inner diameter is recommended. The 1.7 μm fine particle size packing provides a higher theoretical plate number (≥10,000 plates / meter), enhancing the separation ability of structurally similar components (such as quercetin and astragaloside, both with a mass-to-charge ratio of 447); the short 50 mm column length shortens the separation time, suitable for the rapid analysis requirements of ultra-high performance liquid chromatography; and the 2.1 mm inner diameter reduces mobile phase consumption and lowers experimental costs. Before use, the column should be equilibrated for 30 minutes with the initial mobile phase (80% aqueous solution containing 0.1% formic acid + 20% acetonitrile). During equilibration, monitor column pressure fluctuations. When the column pressure stabilizes within ±0.5 MPa, the separation channel is confirmed to be established, and sample analysis can begin.
[0061] The separation principle of reversed-phase chromatography is based on the "difference in hydrophobicity". The stationary phase of the C18 column is octadecylsilane-bonded silica gel, which is highly hydrophobic. The molecular markers (flavonoids and isoflavones) in the sample contain phenolic hydroxyl groups and have a certain degree of hydrophilicity. Under the action of the mobile phase (polar aqueous phase + weakly polar acetonitrile), the hydrophilic components (such as isoquercetin, which contains multiple hydroxyl groups) elute first, and the hydrophobic components (such as quercetin, which has fewer hydroxyl groups) elute later, thus achieving the sequential separation of different markers.
[0062] Start the gradient elution program with mobile phase A being an aqueous solution containing 0.1% formic acid and mobile phase B being acetonitrile, and maintain a constant flow rate of 0.400 mL / min. This step is crucial for achieving complete separation of complex components. By dynamically adjusting the composition of the mobile phase to change the elution intensity, it ensures that both hydrophilic and hydrophobic components elute at the appropriate time. Simultaneously, it optimizes the mobile phase additives and flow rate to balance separation efficiency and peak quality. The specific implementation method is as follows: The choice of mobile phase needs to balance solubility and separation selectivity: Mobile phase A is an aqueous solution containing 0.1% formic acid. The role of the 0.1% formic acid is to adjust the pH of the mobile phase to 2.5-3.0, form hydrogen bonds with the phenolic hydroxyl groups of molecular markers (flavonoids), inhibit the tailing adsorption of components on the stationary phase of the chromatographic column, and improve peak shape (tailing factor controlled between 0.9-1.1). At the same time, formic acid can enhance the efficiency of negative ion electrospray ionization (ESI-), providing a more stable ion signal for subsequent mass spectrometry detection. Mobile phase B is chromatographically pure acetonitrile. Its weak polarity is the core of achieving gradient elution. By changing the acetonitrile ratio, the elution intensity can be dynamically adjusted to adapt to the separation requirements of different hydrophobic components.
[0063] The gradient elution program is precisely controlled in time segments, with a constant flow rate of 0.400 mL / min throughout. This flow rate has been experimentally verified to ensure stable laminar flow of the mobile phase in a 1.7 μm particle size column (avoiding peak broadening caused by eddies) while completing a full analysis within 12 minutes, balancing efficiency and resolution. The program consists of five stages: The first stage (0-1.0 min) is initial isocratic elution, with the mobile phase composition maintained at 80% phase A + 20% phase B. This allows strongly hydrophilic impurities in the sample (such as residual ethanol and small-molecule sugars) to elute first, while simultaneously ensuring sufficient equilibration between the column and the mobile phase. The second stage (1.0-7.5 min) is linear gradient elution, with the proportion of phase B increasing uniformly from 20% to 100%. The elution intensity increases with the increase in the proportion of acetonitrile, effectively removing moderately hydrophobic flavonoids (such as isoquercitrin, retention time 4.28 min) and isoflavones (such as daidzein, retention time 4.85 min). The first stage (7.5-10.0 minutes) involves sequential elution; the second stage (7.5-10.0 minutes) is isocratic elution, maintaining 100% B phase, used to elute the most hydrophobic components (such as quercetin, retention time 4.99 minutes) and any remaining strongly adsorbed components, avoiding column contamination; the third stage (10.0-10.1 minutes) is rapid equilibration, with the B phase ratio dropping sharply from 100% to 20%, quickly restoring the initial mobile phase composition; the fourth stage (10.1-12.0 minutes) is reequilibration, maintaining 80% A phase + 20% B phase, allowing the column to return to its initial state, preparing for the next sample analysis.
[0064] The gradient program is automatically controlled by the software of the ultra-high performance liquid chromatography system. The mobile phase is mixed in a low-pressure mixing mode to ensure that the A and B phases in different proportions are mixed evenly. During the mixing process, the ultraviolet absorption value (254 nm) of the mobile phase is monitored. When the absorption value fluctuation is ≤0.001AU, the mixing is confirmed to be uniform and elution can be carried out normally.
[0065] The ultraviolet absorption chromatogram was obtained by scanning at a wavelength of 254 nm using a photodiode array detector; This step is crucial for simultaneously acquiring the ultraviolet characteristics of the components. Utilizing the wide-wavelength scanning capability of a photodiode array (PDA) detector, it captures the absorption signal at the target wavelength for quantification and records full-wavelength information for component peak purity verification. The specific implementation method is as follows: The core advantage of photodiode array detectors lies in "simultaneous detection of multiple wavelengths." Its detection unit consists of hundreds of photodiodes, capable of simultaneously acquiring signals within a wavelength range of 200-600 nanometers, obtaining a complete ultraviolet spectrum of each eluting component without switching wavelengths. 254 nanometers was chosen as the primary detection wavelength because the target molecular markers (flavonoids and isoflavones) in silkworm cocoons all contain benzene ring structures, exhibiting strong ultraviolet absorption at 254 nanometers (molar absorptivity ≥ 10^4 L / (mol・cm)), ensuring high detection sensitivity. Even trace components (such as quercetin in mulberry leaf cocoons, with a content as low as 0.00886 μmol / L) can be effectively detected.
[0066] The scanning parameter settings need to balance sensitivity and data volume: the scanning frequency is set to 10 Hz (10 data points are collected per second) to ensure that each inflection point of the chromatographic peak can be captured, and to avoid peak distortion due to too low sampling frequency; the spectral resolution is set to 1 nanometer to clearly distinguish the differences in the ultraviolet spectra of different markers (such as quercetin having an additional absorption peak at 370 nanometers, while daidzein does not have this peak), providing a basis for subsequent peak purity verification.
[0067] The generation process of the ultraviolet absorption chromatogram is as follows: When the component flowing out of the chromatographic column enters the detector flow cell, the component absorbs light at a wavelength of 254 nm. The absorption intensity is directly proportional to the component concentration (Lambert-Beer Law). The detector converts the light signal into an electrical signal, which is amplified and recorded by the data system as a time-absorption intensity curve, i.e., the ultraviolet absorption chromatogram. Each chromatographic peak in the figure corresponds to a component. The retention time of the peak is used for preliminary qualitative identification (e.g., the peak at 4.28 minutes may be isoquercitrin), and the peak area can be used for subsequent quantitative calculations. At the same time, the system automatically stores the full-wavelength ultraviolet spectrum of each chromatographic peak. By comparing the ultraviolet spectrum of the standard (e.g., quercetin standard has two absorption peaks at 254 nm and 370 nm), the purity of the sample peak can be verified. If the spectrum of the sample peak is completely consistent with that of the standard, it indicates that there is no overlap with other components, and the separation effect is qualified.
[0068] The chromatographic effluent was introduced into a single quadrupole mass spectrometer and scanned in negative ion electrospray ionization mode in the mass-to-charge ratio range of 50-1000 Da. This step is crucial for the accurate qualitative analysis of molecular markers. Utilizing the ionization and mass analysis capabilities of a mass spectrometer, it obtains the mass-to-charge ratio (m / z) and fragment information of the components, providing key evidence for distinguishing structurally similar components (such as quercetin and astragaloside, both with a m / z of 447). The specific implementation method is as follows: The chromatographic effluent is introduced using a split-flow mode, splitting the post-column effluent of the ultra-high performance liquid chromatography (UHPLC) system at a 1:1 ratio. One portion enters the photodiode array detector, while the other portion is introduced into the ion source of the mass spectrometer via a heated capillary. Precise control of the split ratio is crucial to avoid fluctuations in mass spectrometry signal intensity due to uneven splitting. The temperature of the heated capillary is set to 350℃, which allows for rapid evaporation of the mobile phase (water and acetonitrile) while preventing degradation of the target marker at high temperatures (e.g., isoquercitrin exhibits good stability at 350℃ with a degradation rate ≤1%).
[0069] The ion source employs electrospray ionization (ESI) technology. The negative ion mode was chosen because the target molecular markers (flavonoids and isoflavones) all contain multiple phenolic hydroxyl groups (-OH), which readily lose a proton (H+) during ionization. + This forms negatively charged molecular ions ([MH]). - For example, the molecular ion peak of quercetin is m / z 301 (molecular formula C). 15 H 10 O7, molecular weight 302, [MH] - =302-1=301), the molecular ion peak of daidzein is m / z 253 (molecular formula C). 15 H 10 O4, molecular weight 254, [MH] - =254-1=253), the signal intensity of these molecular ion peaks is significantly higher in negative ion mode than in positive ion mode, and the detection sensitivity is improved by 5-10 times.
[0070] The scanning parameters of a single quadrupole mass spectrometer need to cover the mass-to-charge ratio range of the target marker: the scan range should be set to 50-1000 Da. This range includes the molecular ion peaks (253-463 Da) of all target markers, while excluding interference from low-mass (<50 Da) mobile phase impurities (such as formate ions m / z 45) and high-mass (>1000 Da) macromolecular impurities (such as incompletely dissolved silk fibroin fragments). The scan rate should be set to 1000 Da / sec to ensure that at least 10 scans are completed within the elution time of each chromatographic peak (approximately 0.1 minutes) to acquire a continuous mass spectrum signal and avoid missing mass spectrometry information due to excessively slow scan rates. In addition, appropriate spray voltage (-3.5 kV) and sheath gas pressure (30 psi) need to be set. The spray voltage controls the ionization efficiency, while the sheath gas helps atomize the droplets. Together, they maximize the abundance of the molecular ion peak, providing a clear mass spectrum signal for subsequent qualitative analysis.
[0071] Simultaneously record total ion current chromatograms and mass spectrometry fragment information to generate a complete chromatogram-mass spectrometry dataset.
[0072] This step is crucial for integrating separation and detection information. By simultaneously recording total ion current and mass spectrometry debris, a three-dimensional data system of "retention time-mass-charge ratio-signal intensity" is constructed, providing complete data support for subsequent deep learning model analysis. The specific implementation method is as follows: Total ion current chromatogram (TIC) is recorded based on the full scan signal of the mass spectrometer. The principle is to sum the ion signal intensities of all mass-to-charge ratios (50-1000 Da) at each time point (0.01 minutes apart) to form a time-to-total ion intensity curve. The purpose of the TIC chromatogram is to visually present the elution of all ionizable components in the sample. The retention time of each peak in the chromatogram is compared with that of the UV absorption chromatogram. Figure 1 A one-to-one correspondence (e.g., the peak at 4.28 minutes also exists in the TIC plot) can verify the consistency of the components by comparing the peak shapes of the two plots. If a peak exists in the UV plot but is missing in the TIC plot, it means that the component cannot be ionized (not a target marker). If there is a peak in the TIC plot but no peak in the UV plot, it may be a non-UV absorbing component (such as a small molecule impurity), which needs to be further investigated.
[0073] Mass spectrometry fragmentation information is recorded at the peak of each chromatographic peak. When a peak in the TIC chromatogram reaches its maximum intensity, the mass spectrometer's fragmentation scan mode (MS / MS) is triggered. An inert gas (such as argon) is introduced through the collision cell, causing molecular ions ([MH) to... -When flavonoids collide and break apart, they produce characteristic fragment ions. For example, the molecular ion peak of quercetin is 301. After collision, it produces characteristic fragments such as m / z 151 (A-ring fragment) and m / z 179 (B-ring fragment). The relative abundance ratio of these fragments (e.g., m / z 151:m / z 179 ≈ 1:2) is a key basis for qualitative analysis and can distinguish structurally similar flavonoid components (such as quercetin and quercetin glycoside, the latter having a molecular ion peak of m / z 447 and a fragment peak containing m / z 301, indicating that it generates quercetin after hydrolysis).
[0074] A complete chromatography-mass spectrometry (GC-MS) dataset must include four core types of information: First, total ion current chromatogram data, including retention times (0-12 minutes, 0.01-minute intervals) and corresponding total ion intensities; second, ultraviolet absorption chromatogram data, including retention times and absorption intensities at 254 nm wavelength; third, molecular ion peak data, including retention times, molecular ion mass-to-charge ratios, and peak areas for each peak; and fourth, mass spectrometry fragment data, including the characteristic fragment mass-to-charge ratios and relative abundances for each molecular ion peak. The dataset is stored in a standard format to ensure direct readability by subsequent deep learning models, and includes experimental parameter records (such as column type, mobile phase ratio, and mass spectrometry ionization parameters) to provide a basis for data traceability and replication validation.
[0075] S204, a deep learning model is used to extract features and learn similarity from the chromatographic-mass spectrometry data. When the feature similarity score is higher than the preset reliability threshold, the target compound is confirmed, and the qualitative analysis of the molecular marker is completed. Specifically, a standard database can be loaded, containing retention time, mass-to-charge ratio, and characteristic fragment ion information of the target molecule markers; This step is the benchmark establishment stage for qualitative analysis of molecular markers. Its core is to construct a standard information database covering all target components, providing a precise reference for subsequent sample feature comparisons and ensuring the accuracy and specificity of the qualitative results. The specific implementation method is as follows: The standard database must be constructed based on experimentally validated target molecular markers, covering the core components for identifying the source and evaluating the quality of silkworm cocoons, including flavonoids specific to mulberry leaf cocoons (quercetin, isoquercitrin, quercitrin, astragalin) and isoflavones specific to feed cocoons (daidzein, genistein). Each standard entry in the database must include three key parameters, and all parameters must be determined by averaging multiple repeated experiments (≥5 times) to ensure data reliability.
[0076] The first parameter is retention time (RT), which refers to the time it takes for a standard to elute from the column under specific chromatographic conditions. It must be accurate to three decimal places. For example, the chromatographic retention time of isoquercitrin is 4.227 minutes, while the mass spectrometry retention time (the mass spectrometry detection time synchronized with the chromatographic elution) is 4.281 minutes. This slight difference arises from the response delays of the chromatographic and mass spectrometry detectors and must be recorded separately to avoid comparison bias. Retention time determination must be performed under chromatographic conditions identical to those used for sample analysis (C18 column, gradient elution program, flow rate 0.400 mL / min) to ensure that the influence of environmental factors (such as column temperature and mobile phase ratio) on retention time is eliminated.
[0077] The second parameter is the mass-to-charge ratio (m / z), which is the ratio of the mass to the charge of the molecular ion peak of the standard. All target markers are generated using negative ion electrospray ionization (ESI) mode [MH]. - Molecular ions, such as quercetin, have the molecular formula C1. 15 H 10 O7, with a molecular weight of 302, has an m / z of 301.23 after losing a proton; daidzein has the molecular formula C. 15 H 10 O4, molecular weight 254, m / z 253.23. The mass-to-charge ratio needs to be accurate to two decimal places, and the relative abundance of the molecular ion peak (i.e., the intensity percentage of the peak among all mass spectrometry peaks) needs to be recorded. For example, the abundance of the molecular ion peak of genistein is 92%, ensuring that subsequent judgment can be aided by the abundance.
[0078] The third parameter is the characteristic fragment ion information, which refers to the characteristic ions produced after the molecular ion is fragmented by collision and their relative abundance ratios. This is crucial for distinguishing structurally similar components. For example, quercetin (m / z 447.24) produces m / z 301.23 (quercetin aglycone fragment, abundance 75%) and m / z 146.05 (glycosyl fragment, abundance 25%) after collision, with an abundance ratio of approximately 3:1. While the structurally similar astragaloside (m / z 447.24) has the same molecular ion peak, its fragment ions are m / z 301.23 (abundance 60%) and m / z 146.05 (abundance 40%), with an abundance ratio of 1.5:1. This difference allows for precise differentiation between the two. The database should fully record 3-5 major fragment ions and their abundance ratios for each standard to ensure specificity during comparison.
[0079] When loading the database, its completeness must be verified through data validation. If any standard information is missing (e.g., fragment ions are not recorded), the system will automatically prompt for its supplementation to avoid qualitative misjudgments due to incomplete information. Simultaneously, the database supports dynamic updates, allowing the addition of experimentally validated novel biomarkers (such as characteristic components under specific breeding conditions) to ensure coverage of more application scenarios.
[0080] The chromatography-mass spectrometry dataset is input into a pre-trained deep learning model for multi-dimensional feature extraction and dimensionality reduction. This step is crucial for transforming raw detection data into comparable features. Leveraging the strong fitting capabilities of deep learning models, highly discriminative features are extracted from complex chromatographic-mass spectrometry data. Dimensionality reduction simplifies the calculations, laying the foundation for subsequent similarity comparisons. The specific implementation method is as follows: The pre-trained deep learning model employs a hybrid architecture of "Convolutional Neural Network (CNN) + Recurrent Neural Network (RNN)," which can simultaneously process time-series chromatographic data and discrete feature data from mass spectrometry. The model training process requires a large amount of labeled data (including 500 sets of standard data and 1000 sets of silkworm cocoon sample data from known sources). The training objective is to minimize the feature extraction error (error ≤ 5%). The model input is a complete chromatographic-mass spectrometry dataset, including peak shape parameters (peak height, peak width, peak area) of the total ion current chromatogram, molecular ion abundance of the mass spectrometer, fragment ion abundance ratio, and other raw data. The data must first undergo standardization (mapping all parameters to the 0-1 range) to avoid the influence of magnitude differences on feature extraction.
[0081] The multi-dimensional feature extraction process is carried out in two steps: The first step is to extract the morphological features of the chromatographic peaks through a CNN layer. The convolution kernel of the CNN (3×3) can capture details such as the rising slope, falling slope, and peak width uniformity of the peak. For example, the features "rising slope 0.8 AU / min, falling slope 0.6 AU / min, peak width 0.12 min" can be extracted from the chromatographic peak of isoquercitrin. These parameters can reflect the symmetry and separation purity of the peak. The second step is to extract the sequence features of the mass spectrometry through an RNN layer (using an LSTM structure). The LSTM unit can remember the abundance correlation between molecular ions and fragment ions. For example, the feature sequence "molecular ion abundance 90%, fragment m / z 235 abundance 15%, fragment m / z 153 abundance 5%" can be extracted from the mass spectrometry data of daidzein. This sequence can uniquely identify the mass spectrometry features of daidzein.
[0082] The feature dimensionality reduction process employs Principal Component Analysis (PCA) algorithm. Its core purpose is to reduce feature dimensionality, eliminate redundant information, and avoid model overfitting. The original extracted features were approximately 200-dimensional (containing 50 chromatographic features and 150 mass spectrometry features). PCA reduced these dimensions to 50, while preserving over 95% of the original data variance (ensuring no loss of key features). For example, after PCA processing, the 200-dimensional original features of isoquercitrin were compressed into a 50-dimensional feature vector containing core information such as peak width, molecular ion abundance, and fragment m / z 301 abundance ratio. The first three principal components explain 80% of the variance, simplifying subsequent calculations while preserving feature distinctiveness.
[0083] The dimensionality-reduced feature vectors need to be validated for effectiveness. By comparing the differences between the feature vectors of the standard and the feature vectors of known samples, it can be ensured that the dimensionality reduction has not resulted in the loss of key information. For example, if the dimensionality-reduced vector of quercetin in the standard and the dimensionality-reduced vector of quercetin in the mulberry leaf cocoon sample have a difference of ≤3% in the core dimensions (such as m / z301 abundance and peak height), it proves that the dimensionality reduction process is effective.
[0084] Calculate the cosine similarity between the feature vector of the sample and the feature vector of the standard, and generate a feature similarity score; This step is the core computational step for component matching. It uses cosine similarity to quantify the similarity between the sample and the standard. A higher score indicates a higher degree of feature matching, providing a quantitative basis for the identification of the target compound. The specific implementation method is as follows: The calculation logic of cosine similarity is based on the vector space model. Its core is to measure the cosine value of the angle between two feature vectors, which ranges from 0 to 1. The closer the value is to 1, the more consistent the directions of the two vectors (the more similar the features); the closer the value is to 0, the greater the difference in features. The calculation process requires first clarifying the composition of the sample feature vector and the standard feature vector—both are 50-dimensional vectors after dimensionality reduction. Each dimension corresponds to a core feature (such as dimensional 1 being the chromatographic peak width, dimensional 2 being the molecular ion abundance, and dimensional 3 being the fragment m / z301 abundance ratio, etc.), and the values of all dimensions have been standardized to the 0-1 range to ensure that the contribution of each dimension to the similarity is balanced.
[0085] Taking the qualitative analysis of an unknown component in a mulberry leaf cocoon sample as an example, the 50-dimensional feature vector V_sample of this component is first extracted, and then the standard feature vector V_standard of isoquercitrin is retrieved from the standard database. During calculation, the dot product of the two vectors is first calculated (i.e., the sum of the product of the corresponding dimension values), then the magnitudes of the two vectors are calculated separately (i.e., the square root of the sum of the squares of the values in each dimension). Finally, the dot product is divided by the product of the two magnitudes to obtain the cosine similarity score. For example, the dot product of V_sample and V_standard is 45.2, the magnitude of V_sample is 7.2, and the magnitude of V_standard is 6.8. The similarity score = 45.2 / (7.2×6.8)≈45.2 / 48.96≈0.923, indicating that the sample component is highly similar to the characteristics of isoquercitrin.
[0086] To ensure calculation accuracy, the influence of outliers on similarity scores must be eliminated. For example, if the value of a certain dimension deviates from the normal range (e.g., the abundance of a fragment ion is 0 due to instrument fluctuations), the value of that dimension needs to be corrected using the median filling method before similarity calculation. Simultaneously, the average similarity of multiple detection results for the same compound should be taken. For instance, if quercetin in a sample is detected three times, yielding similarity scores of 0.93, 0.95, and 0.94, the average of 0.94 should be used as the final score to reduce the impact of random errors.
[0087] The similarity scores of different types of markers are specific. For example, the similarity of the feature vectors of quercetin, which is unique to mulberry leaf cocoons, and daidzein, which is unique to feed cocoons, is usually ≤0.3, which can effectively avoid cross-class misclassification. However, the similarity scores of quercetin and astragaloside, which are structurally similar, are about 0.7-0.8. They need to be further distinguished by combining the fragment ion abundance ratio to ensure the uniqueness of the qualitative results.
[0088] The feature similarity score is compared with a preset reliability threshold to output the compound confirmation result; This step is the decision-making stage for qualitative judgment. Valid matching results are filtered by setting a pre-set confidence threshold, eliminating interfering components and low-similarity matches to ensure the accuracy of target compound identification. The specific implementation method is as follows: The determination of the confidence threshold requires statistical analysis of a large amount of experimental data. The core principle is to "maximize the detection rate of the target compound while ensuring an extremely low false positive rate (≤1%)". By statistically analyzing the similarity scores of 200 sets of standards (covering all target biomarkers) and 300 sets of interfering samples (including common impurities and other plant components), ROC curves (Receiver Operating Characteristic curves) were plotted, and the point with the largest area under the curve and the lowest false positive rate was selected as the threshold. Experimental results show that when the threshold is set to 0.95, the detection rate of the target compound reaches 98%, and the false positive rate is only 0.8%. This threshold balances accuracy and detection efficiency and is suitable for the qualitative analysis of silkworm cocoon molecular biomarkers.
[0089] The comparison process must be executed according to the logic of "comparing one by one - classifying results": For each detected chromatographic peak in the sample, its feature vector is extracted and its similarity score is calculated with the vectors of all target markers in the standard database. Then, the score is compared with a threshold of 0.95, and three types of results are output. The first type is "confirmed detection". When the similarity score of a component with a standard is ≥0.95, and the retention time of the component deviates from the retention time of the standard by ≤±0.05 minutes (excluding deviations caused by column drift), then the component is confirmed as the target compound. For example, if a component in the sample has a similarity of 0.96 with the quercetin standard and a retention time deviation of 0.03 minutes, the output is "confirmed detection of quercetin". The second category is "below the detection limit". When a component has a similarity score <0.95 with all standards and its peak area is lower than the instrument's minimum detection peak area (e.g., 1000 AU·s, determined through standard dilution experiments), the component is considered to be below the detection limit. For example, if a component in the sample has a similarity of 0.75 with daidzein and a peak area of 800 AU·s, the output will be "Daidzein below the detection limit". The third category is "suspected interference". When a component has a similarity score between 0.8 and 0.9 with a standard and a retention time deviation >0.05 minutes, it should be marked as a suspected interfering component. Further verification through supplementary experiments (e.g., changing the chromatographic column) is recommended. For example, if a component has a similarity of 0.85 with isoquercitrin and a retention time deviation of 0.08 minutes, the output will be "Suspected isoquercitrin interfering component, further verification required".
[0090] The output results should include key supporting evidence, such as similarity score, retention time deviation, and peak area value, to facilitate subsequent traceability and verification. For example, the output result for a certain component in a feed cocoon sample is "Daidzein confirmed detection, similarity 0.96, retention time 4.85 minutes (standard 4.850 minutes, deviation 0 minutes), peak area 50000 AU·s", clearly presenting the core information of the qualitative judgment.
[0091] Integrate all confirmation results to generate a qualitative analysis report of molecular markers.
[0092] This step is the summary and output stage of the qualitative analysis. By systematically integrating the confirmation results of various markers and combining them with the correlation characteristics between silkworm cocoon feeding sources and quality, a complete and interpretable analysis report is generated, providing a clear direction for subsequent quantitative analysis and quality evaluation. The specific implementation method is as follows: The analysis report must include four core modules to ensure comprehensive information and clear logic. The first module is basic sample information, including sample number, sampling date, pretreatment method (e.g., degumming time 10 minutes, extraction solvent 75% ethanol), detection date, and a summary of instrument parameters (e.g., column type, mass spectrometry ionization mode), such as "Sample number SL-05, sampling date 2024-05-10, degumming treatment 10 minutes, ultrasonic extraction with 75% ethanol, detection date 2024-05-12, column C18 (1.7μm, 2.1×50mm), mass spectrometry ESI-mode," providing a foundation for the report's traceability.
[0093] The second module is a summary of molecular marker confirmation results, presented in the following logic: "marker category - name - retention time - similarity score - confirmation status - remarks". The marker categories are divided into "mulberry leaf cocoon characteristic markers" and "feed cocoon characteristic markers", clearly distinguishing the specific components from different feeding sources. For example, the summary results for mulberry leaf cocoon samples are as follows: "Mulberry leaf cocoon characteristic markers - quercetin: retention time 4.99 minutes, similarity 0.96, confirmed detection; isoquercitrin: retention time 4.28 minutes, similarity 0.95, confirmed detection; quercetin: retention time 4.46 minutes, similarity 0.97, confirmed detection; astragaloside: retention time 4.44 minutes, similarity 0.88, suspected interference, needs verification; feed cocoon characteristic markers - daidzein: retention time 4.85 minutes, similarity 0.3, below the detection limit; genistein: retention time 5.26 minutes, similarity 0.25, below the detection limit," clearly showing the presence of each marker in the samples, and consistent with the component characteristics of mulberry leaf feed sources.
[0094] The third module is the qualitative conclusion and basis, which determines the feeding source of silkworm cocoons based on the biomarker confirmation results and explains the basis for the judgment. For example, "Qualitative conclusion: This sample is a silkworm cocoon fed with mulberry leaves (mulberry leaf cocoon); Basis for judgment: 1. Quercetin, isoquercitrin, and quercetin, which are unique to mulberry leaf cocoons, were detected, with similarity ≥0.95; 2. Daidzein and genistein, which are unique to feed cocoons, were not detected, with similarity <0.3; 3. Suspected astragalin interfering components do not affect the judgment of feeding source, as they are not the core biomarkers that distinguish between mulberry leaf and feed cocoons." Combining the correlation between silkworm cocoon components and feeding methods, the conclusion has a scientific basis.
[0095] The fourth module contains recommendations and explanations, including key biomarkers for subsequent quantitative analysis (e.g., prioritizing the quantification of quercetin and isoquercitrin to assess the quality of mulberry leaf cocoons), verification protocols for suspected components (e.g., changing the mobile phase ratio to verify astragaloside), and the scope of the report (e.g., this report is only responsible for the qualitative analysis of molecular biomarkers in this sample and does not include sensory quality evaluation), providing clear guidance for subsequent work. For example, the recommendations section states: "For subsequent quantitative analysis, it is recommended to prioritize the determination of quercetin and isoquercitrin content to assess the abundance of flavonoids in mulberry leaf cocoons; suspected astragaloside interference can be re-verified by adjusting the formic acid concentration in mobile phase A to 0.2%; the qualitative results in this report are based on current experimental conditions, and re-analysis is required if the detection parameters are changed."
[0096] After the report is generated, it must undergo three levels of review to ensure the accuracy of the information. The review includes the consistency between the marker confirmation results and the original data, the correctness of the similarity score calculation, and the logic of the conclusions and basis. After the review is passed, a special testing seal is affixed, forming a formal qualitative analysis report of molecular markers, which serves as the core basis for subsequent quantitative analysis and comprehensive quality evaluation.
[0097] S205, based on qualitative analysis results, establishes a standard curve using the external standard method, calculates the content of the target compound in silkworm cocoons based on the peak area, and realizes quantitative analysis of molecular markers and comprehensive evaluation of silkworm cocoon quality.
[0098] Specifically, a series of standard solutions with varying concentrations can be prepared based on the target molecular markers identified in the qualitative analysis report. This step is fundamental to external standard quantification. Its core is to construct a concentration gradient covering the linear response range of the target molecular biomarker, ensuring the accuracy of subsequent sample concentration calculations. It requires combining qualitative analysis results to screen key biomarkers and strictly controlling the preparation precision of the standard solutions. The specific implementation method is as follows: First, based on the qualitative analysis report, identify the target molecular markers. If the report confirms that the sample is silkworm cocoon fed with mulberry leaves (mulberry leaf cocoon), focus on the flavonoid markers unique to mulberry leaf cocoons (quercetin, isoquercitrin, quercetin); if it is silkworm cocoon fed with artificial feed (feed cocoon), focus on isoflavone markers (daidzein, genistein), while also taking into account other markers that may coexist (such as astragaloside), to ensure that the quantitative analysis covers all core components related to source identification and quality.
[0099] The preparation of standard solutions must follow the process of "stock solution preparation - gradient dilution - filtration and volume adjustment," and all operations must be performed in a clean laboratory to avoid contamination. Taking quercetin standard as an example, firstly, prepare the stock solution: Select quercetin standard with a purity ≥98%, accurately weigh 0.0076 g (quercetin molecular weight 302 g / mol, 0.0076 g ÷ 302 g / mol ≈ 0.025 mol / L, i.e., 25 mmol / L), dissolve it in 75% ethanol aqueous solution, transfer it to a 100 mL volumetric flask, and adjust the volume. After shaking well, a 25 mmol / L stock solution is obtained. This stock solution should be stored at 4℃ protected from light and has a shelf life of 7 days.
[0100] Subsequently, a serial dilution was performed. Based on the linear range determined in the preliminary experiments (0.01-1.0 μmol / L, which covers the common content of quercetin in silkworm cocoon samples, and the peak area shows good linearity with concentration), the stock solution was gradually diluted to five concentration gradients: 0.01 μmol / L, 0.05 μmol / L, 0.1 μmol / L, 0.5 μmol / L, and 1.0 μmol / L. Accurate measurements using pipettes were required during the dilution process. For example, when preparing a 0.1 μmol / L solution, 0.4 mL of the 25 mmol / L stock solution was transferred to a 100 mL volumetric flask and brought to volume with 75% ethanol (25 mmol / L × 0.4 mL = 10 μmol, 10 μmol ÷ 100 mL = 0.1 μmol / L). Each concentration gradient required triplicate preparation to ensure the repeatability of subsequent analyses.
[0101] All standard solutions of varying concentrations must be filtered through a 0.22-micron filter membrane (in the same manner as sample solutions) to remove any minute impurities that may be present, preventing column clogging or interference with peak area determination. The filtered solutions are then transferred to vials, labeled with the concentration, preparation date, and operator to ensure traceability. Analysis must be completed within 24 hours to prevent standard degradation (e.g., isoflavones are easily oxidized under light and must be stored away from light).
[0102] The standard solution was analyzed by liquid chromatography-mass spectrometry, and the chromatographic peak area data corresponding to each concentration was recorded. This step is crucial for establishing the correlation between concentration and peak area. It requires strictly adhering to the same LC-MS conditions as the sample analysis to ensure that the experimental environment and instrument parameters have the same impact on peak area as the sample analysis. Simultaneously, multiple repeated injections are used to reduce random errors. The specific implementation method is as follows: The LC-MS / MS conditions must be absolutely consistent with those of the sample analysis, including chromatographic and mass spectrometric conditions: chromatographic column: C18 reversed-phase column (1.7 μm, 2.1 × 50 mm); mobile phase A: aqueous solution containing 0.1% formic acid; mobile phase B: acetonitrile; gradient elution program: 0-1.0 min 80% A + 20% B, 1.0-7.5 min B phase up to 100%, 7.5-10.0 min maintain 100% B, 10.0-12.0 min restore initial ratio; flow rate: 0.400 mL / min; column temperature: 30 °C; mass spectrometry conditions: negative ion electrospray ionization (ESI); scan range: 50-1000 Da; spray voltage: -3.5 kV; heating capillary temperature: 350 °C; collision energy: 15 eV (for fragment ion verification).
[0103] During sample injection analysis, inject samples in ascending order of concentration (to avoid high-concentration residues affecting low-concentration determinations). Inject the standard solution for each concentration gradient three times, with an injection volume of 2 μL (consistent with the sample injection volume). Peak area recording should focus on the characteristic peaks of the target biomarker. For example, the characteristic peak of quercetin corresponds to a retention time of 4.99 minutes and a mass-to-charge ratio of 301.23, while that of daidzein corresponds to a retention time of 4.85 minutes and a mass-to-charge ratio of 253.23. Selecting the peak area of the target ion from the extracted ion chromatogram (EIC) of the mass spectrometer, rather than the peak area of the total ion chromatogram, can effectively eliminate the influence of interfering peaks.
[0104] Peak area statistics require the removal of outliers. For example, if three injections of a certain concentration yield peak areas of 44,800 AU·s, 45,200 AU·s, and 50,000 AU·s, and the 50,000 AU·s peak area deviates from the average by more than 5%, it is considered an outlier (possibly due to injection error). After removing the outlier, the average of the first two, 45,000 AU·s, is taken as the final peak area for that concentration. Simultaneously, the retention time deviation for each concentration must be recorded, ensuring the deviation is ≤ ±0.05 minutes. If the deviation is too large (e.g., exceeding 0.1 minutes), the column must be rebalanced before re-injection to avoid inaccurate peak areas due to column drift.
[0105] Taking specific concentrations as an example, the average peak areas corresponding to different concentrations of quercetin are as follows: 0.01 μmol / L is 4200 AU·s, 0.05 μmol / L is 21500 AU·s, 0.1 μmol / L is 45000 AU·s, 0.5 μmol / L is 224000 AU·s, and 1.0 μmol / L is 448000 AU·s. These data need to be recorded one by one with the concentration to provide a basis for subsequent standard curve fitting.
[0106] A standard curve is generated by fitting a linear regression equation with the concentration of the standard as the x-axis and the chromatographic peak area as the y-axis. This step forms the mathematical basis for converting peak area into concentration. A quantitative relationship between the two is established through linear regression. The least squares method must be used to ensure fitting accuracy, and the linear correlation must be verified. The specific implementation method is as follows: First, clarify the definition and units of the coordinate axes: the horizontal axis (x-axis) represents the standard concentration, in micromoles per liter (μmol / L), with values of 0.01, 0.05, 0.1, 0.5, and 1.0; the vertical axis (y-axis) represents the average peak area of the corresponding concentration, in area units per second (AU·s), with values of 4200, 21500, 45000, 224000, and 448000. The core of linear regression is finding the best-fitting line y = ax + b, where a is the slope (sensitivity coefficient) and b is the intercept (background signal). a and b are calculated using the least squares method, i.e., minimizing the sum of squared residuals between each data point and the line.
[0107] The calculation process needs to be explained in detail. For example, for quercetin data: First, calculate the average value of x ((0.01+0.05+0.1+0.5+1.0) / 5=0.332μmol / L), and the average value of y ((4200+21500+45000+224000+448000) / 5=148540AU・s); then calculate the slope a=[nΣ(xy)-ΣxΣy] / [nΣ(x 2 )-(Σx) 2 ], where n=5, Σ(xy)=0.01×4200+0.05×21500+0.1×45000+0.5×224000+1.0× 448000=0+1075+4500+112000+448000=565575, Σx=1.66, Σy=742700, Σ(x 2 )=0.01 2 +0.05 2 +0.1 2 +0.5 2 +1.0 2 =0.0001+0.0025+0.01+0.25+1=1.2626; Substituting, we get a=[5×565575-1.66×742700] / [5×1.2626-(1.66)] 2 The linear regression equation for quercetin is y = 448360x - 315. The intercept is b = y_avg - ax_avg = 148540 - 448360 × 0.332 ≈ -315AU・s.
[0108] The linear correlation was verified by calculating the correlation coefficient R. 2 Implementation, R 2 The closer the value is to 1, the better the linear relationship. The calculation formula is R.2 =1-[Σ(y-ŷ) 2 ] / [Σ(y-y_avg) 2 ], where ŷ is the theoretical peak area calculated by the regression equation. Substituting the quercetin data, we obtain R. 2 ≈0.9998, which meets the requirements for quantitative analysis (R0). 2 (≥0.999), indicating that the concentration and peak area have a very strong linear relationship in the range of 0.01-1.0 μmol / L.
[0109] The standard curve needs to be labeled with key parameters, including the regression equation, correlation coefficient, linear range, limit of detection (LOD, usually the concentration corresponding to 3 times the signal-to-noise ratio, such as quercetin LOD=0.003μmol / L) and limit of quantitation (LOQ, the concentration corresponding to 10 times the signal-to-noise ratio, such as 0.01μmol / L). These parameters directly reflect the sensitivity and applicability of the quantitative method. For example, a low LOD indicates that it can detect lower concentrations of biomarkers and is suitable for trace analysis.
[0110] Substitute the chromatographic peak area of the target compound in the silkworm cocoon sample into the standard curve equation to calculate its concentration value in the actual sample. This step is the core calculation in quantitative analysis. It requires accurately extracting the peak area of the target biomarker in the sample, substituting it into the corresponding standard curve equation, and taking into account the dilution factor of the sample pretreatment to calculate the actual content of the biomarker in the silkworm cocoon sample. The specific implementation method is as follows: The extraction of peak areas in samples must be consistent with the analysis of standards. This means selecting the peak area of the target ion using extracted ion chromatograms (EIC), and confirming that the retention time matches the standard (deviation ≤ ±0.05 minutes). For example, in a mulberry leaf cocoon sample, the retention time of quercetin is 4.98 minutes (0.01 minutes different from the standard's 4.99 minutes). The extracted peak area of the ion with a mass-to-charge ratio of 301.23 is 223,000 AU·s. This peak area needs to be subtracted from the peak area of the blank sample (the blank is a 75% ethanol solution with a peak area of approximately 200 AU·s) to obtain a net peak area of 222,800 AU·s (after deducting background interference).
[0111] Substituting the net peak area into the standard curve equation of quercetin, y = 448360x - 315, we can solve for the concentration: x = (y + 315) / 448360 = (222800 + 315) / 448360 ≈ 223115 / 448360 ≈ 0.497 μmol / L. Therefore, the concentration of quercetin in the sample solution is approximately 0.50 μmol / L (retain two significant figures).
[0112] The actual content of the biomarker in the sample needs to take into account the dilution factor of the pretreatment. The pretreatment procedure for silkworm cocoon samples is as follows: 0.5 g of pretreated silkworm cocoon material is extracted with 5 mL of 75% ethanol without additional dilution, so the dilution factor is 1 (i.e., the concentration of the extract is consistent with the concentration in the sample). The formula for calculating the content is: Biomarker content (μmol / g) = Sample solution concentration (μmol / L) × Extraction volume (L) ÷ Silkworm cocoon sample mass (g). Substituting the data: Extraction volume = 5 mL = 0.005 L, Silkworm cocoon mass = 0.5 g, Therefore, quercetin content = 0.50 μmol / L × 0.005 L ÷ 0.5 g = 0.0025 μmol ÷ 0.5 g = 0.005 μmol / g (i.e., 5 nmol / g).
[0113] Parallel experiments and error control must be conducted simultaneously. Three test solutions should be prepared for each sample, and the content should be calculated separately for each. The average value should be taken, and the relative standard deviation (RSD) should be calculated. An RSD ≤ 5% indicates good repeatability. For example, if the quercetin content of a sample is 0.0048 μmol / g, 0.0052 μmol / g, and 0.0050 μmol / g in three measurements, with an average of 0.0050 μmol / g, the RSD is (0.0002 / 0.0050) × 100% = 4%, which meets the requirements.
[0114] For undetected markers (such as daidzein in mulberry leaf cocoons), it should be recorded as "below the limit of quantitation (<0.01 μmol / L)" rather than "not detected" to ensure precise description. For example, if the peak area of daidzein in a mulberry leaf cocoon sample is 800 AU·s, substituting it into the standard curve equation of daidzein y=380000x-250, we can calculate x=(800+250) / 380000≈0.0027μmol / L<0.01μmol / L, and record it as "daidzein content <0.01μmol / L (below LOQ)".
[0115] By combining the quantitative results of various molecular markers with principal component analysis to perform multidimensional data fusion, a comprehensive evaluation report on silkworm cocoon quality is generated.
[0116] This step is the summary and application stage of quantitative analysis. It requires integrating the content data of all biomarkers, using principal component analysis to achieve dimensionality reduction and source identification of multidimensional data, and finally generating a comprehensive evaluation report that includes source identification, quality grading, and supporting evidence. The specific implementation method is as follows: First, summarize the quantitative results of each molecular marker, organizing them according to the logic of "marker category - name - content (μmol / g) - detection status". For example, the summary results of a mulberry leaf cocoon sample are as follows: Flavonoid markers (characteristic of mulberry leaf cocoons) - quercetin 0.0050 μmol / g, isoquercitrin 0.0062 μmol / g, quercetin 0.0045 μmol / g, astragaloside 0.0038 μmol / g; Isoflavone markers (characteristic of feed cocoons) - daidzein <0.01 μmol / L (sample solution concentration, corresponding to cocoon content <0.0001 μmol / g), genistein <0.01 μmol / L. At the same time, calculate the total amount of characteristic markers, such as the total amount of flavonoids in mulberry leaf cocoons = 0.0050 + 0.0062 + 0.0045 + 0.0038 = 0.0195 μmol / g.
[0117] Principal component analysis (PCA) is used for multidimensional data fusion and source identification. Its core principle is to reduce the dimensionality of multiple biomarker content data to at least a few principal components (PCs), and then classify the samples through clustering in the principal component space. Taking 10 mulberry leaf cocoon samples and 8 feed cocoon samples as an example, PCA analysis was performed using the content of six biomarkers from all samples as variables: the first principal component (PC1) explained 57% of the total variance, mainly contributed by flavonoid content (positive loading) and isoflavone content (loading); the second principal component (PC2) explained 22.5% of the total variance, mainly contributed by the difference in content between quercetin and astragaloside. In the PC1-PC2 score plot, mulberry leaf cocoon samples are concentrated on the positive half-axis of PC1 (high flavonoid content), while feed cocoon samples are concentrated on the negative half-axis of PC1 (high isoflavone content). The two groups of samples form completely separate clusters with no overlapping areas. Combined with the position of a sample to be evaluated in the score plot (PC1=2.3, PC2=0.5, located within the mulberry leaf cocoon cluster), it can be helpful to confirm that the sample source is mulberry leaf feeding.
[0118] The comprehensive quality evaluation report should include four core modules: First, the conclusion on the identification of the sample source, based on the types of biomarkers and PCA clustering, such as "This sample is a silkworm cocoon fed on mulberry leaves (mulberry leaf cocoon), based on the detection of flavonoid biomarkers specific to mulberry leaf cocoons (total 0.0195 μmol / g), the absence of isoflavone biomarkers specific to feed cocoons, and PCA analysis showing that the sample is located in the mulberry leaf cocoon cluster region"; Second, quality grading, establishing grading standards based on the total amount of characteristic biomarkers, such as a total flavonoid content ≥0.02 μmol / g is "high quality", 0.01-0.02 μmol / g is "high quality", and 0.01-0.02 μmol / g is "high quality". The grade is as follows: 0.01 μmol / g is "good", <0.01 μmol / g is "fair". The total flavonoid content of this sample is 0.0195 μmol / g, and it is graded as "good". The third part is the analysis of key biomarkers, explaining the significance of each biomarker, such as "isoquercetin has the highest content (0.0062 μmol / g). This component is an antioxidant unique to mulberry leaves. A high content indicates that the sample has good nutritional value". The fourth part is the precautions, such as "the quantitative results are based on 75% ethanol extraction. If water-soluble biomarkers need to be evaluated, the extraction solvent needs to be adjusted. The sample needs to be stored in the dark and refrigerated to prevent biomarker degradation".
[0119] The report must be accompanied by key supporting data, including detailed content of each biomarker, PCA score plot (with textual description of clustering), and standard curve parameters, to ensure the scientific validity and traceability of the evaluation conclusions. It must also indicate the scope of application of the method, such as "This evaluation is applicable to dried silkworm cocoon samples; fresh cocoons must be dried to constant weight before analysis."
[0120] It is evident that alkaline degumming of silkworm cocoon samples allows for preliminary differentiation of cocoons from different feeding sources based on differences in degumming rates. Extraction of the cocoon samples with organic solvents yields a sample solution for liquid chromatography-mass spectrometry (LC-MS). LC-MS is then used to separate and detect the sample solution, obtaining LC-MS data of the cocoon samples. A deep learning model is employed to extract features and learn similarity from the LC-MS data, enabling qualitative analysis of molecular markers. Based on the qualitative analysis results, a standard curve is established using the external standard method, and the content of the target compound in the cocoon is calculated based on its peak area. This allows for precise traceability of the cocoon's origin and objective, quantitative evaluation of its intrinsic quality.
[0121] Another embodiment of the present invention provides a silkworm cocoon origin identification and quality evaluation system based on molecular marker analysis, see [link to relevant documentation]. Figure 3 The system may include: The degumming module 301 is used to perform alkaline degumming treatment on silkworm cocoon samples, measure the degumming rate in the initial degumming stage, and make preliminary distinctions between silkworm cocoons from different feeding sources based on the difference in degumming rate. Preparation module 302 is used to extract silkworm cocoon samples with organic solvents to prepare a sample solution for liquid chromatography-mass spectrometry analysis; The detection module 303 is used to separate and detect the sample solution to be tested using liquid chromatography-mass spectrometry to obtain chromatographic-mass spectrometry data of the silkworm cocoon sample; The qualitative module 304 is used to perform feature extraction and similarity learning on the chromatographic-mass spectrometry data using a deep learning model. When the feature similarity score is higher than the preset reliability threshold, the target compound is confirmed, and the qualitative analysis of the molecular marker is completed. The quantitative module 305 is used to establish a standard curve based on the qualitative analysis results using the external standard method, and calculate the content of the target compound in the silkworm cocoon according to the peak area, so as to realize the quantitative analysis of molecular markers and the comprehensive evaluation of silkworm cocoon quality.
[0122] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.
[0123] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0124] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.
[0125] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.
Claims
1. A method for determining the origin and evaluating the quality of cocoon based on molecular marker analysis, characterized by, The method includes: Alkaline degumming treatment was performed on silkworm cocoon samples, and the degumming rate in the initial degumming stage was measured. Based on the difference in degumming rate, a preliminary distinction was made between silkworm cocoons from different feeding sources. Silkworm cocoon samples were extracted with organic solvents to prepare a sample solution for liquid chromatography-mass spectrometry analysis. The sample solution to be tested was separated and detected using liquid chromatography-mass spectrometry to obtain chromatographic-mass spectrometry data of the silkworm cocoon sample; A deep learning model is used to extract features and learn similarity from the chromatographic-mass spectrometry data. When the feature similarity score is higher than the preset reliability threshold, the target compound is confirmed, and the qualitative analysis of the molecular marker is completed. Based on the qualitative analysis results, a standard curve was established using the external standard method. The content of the target compound in the silkworm cocoon was calculated according to the peak area, thus realizing the quantitative analysis of molecular markers and the comprehensive evaluation of silkworm cocoon quality.
2. The method of claim 1, wherein, The process of alkaline degumming of silkworm cocoon samples, measuring the degumming rate in the initial degumming stage, and using the difference in degumming rate to preliminarily distinguish silkworm cocoons from different feeding sources includes: Accurately weigh the initial weight of the dried silkworm cocoon sample and record it as the cocoon weight data before degumming; The weighed silkworm cocoon sample was placed in a boiling 0.02 mol / L sodium carbonate solution and subjected to the first degumming treatment for 1 minute at a weight-to-volume ratio of 1:
400. Take out the silkworm cocoon sample after the first degumming, wash it three times with pure water, and dry it in a 60℃ oven. Record its weight as the cocoon weight data after 1 minute of degumming. The same silkworm cocoon sample was degummed for another 10 minutes, and after washing and drying again, the weight was recorded as the cocoon weight data after 10 minutes of degumming. Based on the cocoon weight data before degumming, the cocoon weight data 1 minute after degumming, and the cocoon weight data 10 minutes after degumming, calculate the degumming rate value within the initial 1 minute. The calculated degumming rate is compared with a preset rate threshold, and a preliminary judgment on the origin of the silkworm cocoons is output based on the comparison results.
3. The method of claim 2, wherein, The step of extracting silkworm cocoon samples with organic solvents to prepare a sample solution for liquid chromatography-mass spectrometry analysis includes: The silkworm cocoon sample was cut into uniform fragments to obtain pretreated silkworm cocoon material; Accurately weigh 0.5 grams of pretreated silkworm cocoon material and add 5 ml of 75% ethanol aqueous solution at a weight-to-volume ratio of 1:10; The mixed solution was subjected to ultrasonic-assisted extraction for 30 minutes to generate a primary extract; The primary extract was heated at 60°C for 4 hours, and then allowed to stand overnight to form a layered solution. The supernatant solution was filtered through a 0.22-micron filter membrane to obtain a clear and transparent sample solution.
4. The method of claim 3, wherein, The method of using liquid chromatography-mass spectrometry (LC-MS) to separate and detect the sample solution to obtain chromatographic-mass spectrometric data of the silkworm cocoon sample includes: The sample solution to be tested is injected into an ultra-high performance liquid chromatography system, and a separation channel is established by equipping a C18 reversed-phase column. Start the gradient elution program with mobile phase A being an aqueous solution containing 0.1% formic acid and mobile phase B being acetonitrile, and maintain a constant flow rate of 0.400 mL / min. The ultraviolet absorption chromatogram was obtained by scanning at a wavelength of 254 nm using a photodiode array detector; The chromatographic effluent is introduced into a single quadrupole mass spectrometer, and scanned in the negative ion electrospray ionization mode in the mass-to-charge ratio range of 50-1000 Da; The total ion current chromatogram and mass spectrum fragment information are recorded synchronously to generate a complete chromatography-mass spectrometry data set.
5. The method of claim 4, wherein, The deep learning model is used for feature extraction and similarity learning on the chromatography-mass spectrometry data, and when the feature similarity score is higher than the pre-set confidence threshold, the target compound is confirmed, and the qualitative analysis of the molecular marker is completed, including: A standard sample database is loaded, including the retention time, mass-to-charge ratio and characteristic fragment ion information of the target molecular marker; The chromatography-mass spectrometry data set is input into the pre-trained deep learning model for multi-dimensional feature extraction and dimension reduction processing; The cosine similarity between the sample feature vector and the standard sample feature vector is calculated to generate a feature similarity score; The feature similarity score is compared with the pre-set confidence threshold, and the compound confirmation result is output; All confirmation results are integrated to generate a molecular marker qualitative analysis report.
6. The method of claim 5, wherein, Based on the qualitative analysis result, a standard curve is established by the external standard method, the content of the target compound in the cocoon is calculated according to the peak area, the quantitative analysis of the molecular marker is realized, and the comprehensive evaluation of the cocoon quality is realized, including: According to the target molecular marker determined in the qualitative analysis report, a series of standard sample solutions with concentration gradients are prepared; The standard sample solutions are analyzed by liquid chromatography-mass spectrometry to record the chromatographic peak area data corresponding to each concentration; The standard curve is generated by fitting the linear regression equation with the standard sample concentration as the abscissa and the chromatographic peak area as the ordinate; The chromatographic peak area of the target compound in the cocoon sample is substituted into the standard curve equation to calculate its concentration value in the actual sample; The quantitative results of each molecular marker are integrated, and multi-dimensional data fusion is performed combined with principal component analysis to output a cocoon quality comprehensive evaluation report.
7. A cocoon origin discrimination and quality evaluation system based on molecular marker analysis, characterized by comprising: a cocoon quality evaluation device according to any one of claims 1 to 6; and a cocoon origin discrimination device according to any one of claims 1 to 6. The system comprises: A degumming module for alkali degumming treatment of the cocoon sample to determine the degumming rate in the initial degumming stage, and realizing preliminary discrimination of cocoon samples from different feeding sources based on the degumming rate difference; A preparation module for preparing a sample solution for liquid chromatography-mass spectrometry analysis by extracting the cocoon sample with an organic solvent; A detection module for separating and detecting the sample solution by liquid chromatography-mass spectrometry to obtain the chromatography-mass spectrometry data of the cocoon sample; A qualitative module for using a deep learning model to extract features and learn similarities from the chromatography-mass spectrometry data, and confirming the target compound when the feature similarity score is higher than the pre-set confidence threshold, to complete the qualitative analysis of the molecular marker; A quantitative module for establishing a standard curve by the external standard method based on the qualitative analysis result, calculating the content of the target compound in the cocoon according to the peak area, realizing the quantitative analysis of the molecular marker, and comprehensively evaluating the quality of the cocoon.
8. The system of claim 7, wherein, The degumming module is specifically used for: Accurately weighing the initial weight of the dried cocoon sample and recording it as the cocoon weight data before degumming; Placing the weighed cocoon sample in a boiling 0.02 mol / L sodium carbonate solution for the first degumming treatment at a weight-to-volume ratio of 1:400 for 1 minute; Take out the cocoon sample after the first degumming, wash it with pure water three times, and then dry it in an oven at 60°C. Record the weight as the cocoon weight data after 1 minute of degumming; Continue to degum the same cocoon sample for 10 minutes, wash and dry it again, and then record the weight as the cocoon weight data after 10 minutes of degumming; According to the cocoon weight data before degumming, the cocoon weight data after 1 minute of degumming, and the cocoon weight data after 10 minutes of degumming, calculate the degumming rate value within the initial 1 minute; Compare the calculated degumming rate value with the preset rate threshold, and output the preliminary discrimination conclusion of the cocoon source according to the comparison result.
9. A storage medium, characterized by The storage medium has a computer program stored therein, wherein the computer program is configured to execute the method of any one of claims 1-6 when running.
10. An electronic device comprising a memory and a processor, characterized in that, The memory has a computer program stored therein, and the processor is configured to execute the computer program to execute the method of any one of claims 1-6.
Citation Information
Cited By
A system for traceability analysis of a natural flavor of sea bass and a synthetic flavor
CN122193484A