Method for discriminating raw milk and liquid milk based on volatile substance mass spectrum fingerprint
Patent Information
- Application Number
- CN202610692776.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]基于蛋白和 DNA 的物种鉴别,易受加热、均质和发酵等加工的干扰,且不能鉴定非物种异常
[0010]This invention collects raw milk and liquid milk samples from different species, origins, and varieties. It gathers spectral fingerprint data for volatile matter, low-to-medium volatile matter, and medium-to-low volatile matter from each sample. A discrimination model is established using training data and validated and optimized using validation data. This overcomes the limitations of using a single indicator for sample discrimination, achieving global feature representation driven by big data and overcoming the insufficient robustness of traditional single-indicator or few-indicator discrimination. Through deep analysis of the fingerprint data, it achieves digital correlation and global representation of complex milk sample attributes, significantly improving discrimination precision. It can simultaneously identify species, feeding methods, geographical origin, and product category/brand. It enables milk sample evaluation, accurate authenticity determination, and efficient traceability.
Smart Images

Figure CN122731007A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for distinguishing between raw milk and liquid milk based on volatile matter spectral fingerprinting. Background Technology
[0002] my country has a vast territory, with its core milk source coming from the "golden milk belt" stretching across Northeast, North, and Northwest my country. Large-scale dairy farming also exists in the Central Plains region, and coastal areas also have advantageous cattle-raising areas. Milk from Chinese Yellow Cattle grazing on grasslands, Mongolian cattle, and highland cattle has become a specialized, refined, unique, and innovative dairy product, enriching the structure of my country's liquid milk products.
[0003] The quality and sensory properties of raw milk depend not only on the species and strain of dairy-producing livestock, but also on feeding methods, forage, and natural geographical and climatic conditions. Specifically regarding cow's milk, Holstein milk is the mainstream, considered the bulk milk, while Jersey milk is also a globally recognized dairy category. Different raw milks have different nutritional qualities and sensory characteristics. Specialty raw milks and newly developed, specialized milks (such as organic milk, pasture-raised milk, and DHA milk) are relatively scarce, have unique flavors, and high nutritional value, leading to supply shortages and higher prices than ordinary cow's and goat's milk. They are easily adulterated or counterfeited by lower-priced cow's and goat's milk or other raw materials.
[0004] Species identification based on proteins and DNA is susceptible to interference from processing methods such as heating, homogenization, and fermentation, and cannot identify non-species anomalies. While research has been conducted on the authenticity determination and traceability of raw and liquid milk based on nutrient fingerprints / nutrient sequences (fatty acids, amino acids, triglycerides, minerals, stable isotopes, carotenoids, etc.), existing methods suffer from the following problems: First, the dimensionality and precision of the determination are limited, relying on only a single or a few physicochemical indicators, leading to poor robustness and weak specificity. It is difficult to accurately distinguish between species origin, feeding methods (organic, grass-fed, grazing, etc.), and geographical origin. Second, there are issues with the stability of characteristic signals during industrial processing. Existing methods based on protein or DNA detection are easily affected by physical interference from liquid milk homogenization and ultra-high temperature sterilization, leading to identification failure. Third, existing methods lack a global discrimination logic, making it impossible to accurately identify illegally adulterated low-priced raw milk in high-priced dairy products.
[0005] The purpose of this invention is to provide a method for distinguishing raw milk and liquid milk based on volatile matter spectral fingerprinting that can reduce the limitations of single-index discrimination, construct a multi-dimensional integrated solution, and fill the gap in systematic traceability of the entire industrial chain. Summary of the Invention
[0006] To address the aforementioned problems, this invention provides a method for distinguishing between raw milk and liquid milk based on volatile matter spectral fingerprinting, comprising: a sample collection step S10, which involves collecting raw milk and liquid milk samples from different species / strains, regions / origins, and brands / categories; a fingerprint data collection step S20, which involves collecting fingerprint data from each raw milk and liquid milk sample and performing detection using two orthogonal detection systems to obtain volatile matter spectral fingerprint data, volatile substance fingerprint data, low-to-medium volatile matter spectral fingerprint data, and low-to-medium volatile matter fingerprint data; and a dataset partitioning step S30, which involves unidirectionally integrating the volatile matter spectral fingerprint data, volatile substance fingerprint data, low-to-medium volatile matter spectral fingerprint data, and low-to-medium volatile matter fingerprint data from each raw milk and liquid milk sample into a fingerprint database within a machine learning module, and then partitioning the fingerprint data in the fingerprint database to obtain training set data and validation set data.
[0007] Model building step S50 involves importing the training set data into the model via a solid line to establish validation and optimization nodes, selecting the modeling method, and setting and debugging parameters. Data structure analysis is then performed using software to obtain the discriminant model. Discriminant model accuracy validation step S60 involves running the validation set data in the discriminant model and comparing the model's discrimination results with the authenticity of each known sample in the validation set data to determine the accuracy of the discriminant model. Model optimization step S70 involves optimizing the discriminant model based on the discrimination results of the validation set data to obtain the optimized discriminant model. Authenticity traceability step S80 involves using the optimized discriminant model to trace the authenticity of the raw milk and liquid milk to be tested.
[0008] Preferably, in the dataset partitioning step S30, 90% to 99% of the samples are partitioned from the total sample set in each round to obtain the training set data, and the remaining data is used as an independent dataset to obtain the validation set data.
[0009] Preferably, the method further includes a data processing step S40, which performs data structure analysis on the training set data and validation set data, and performs scaling, noise reduction and other processing to unify the data units and remove interference noise.
[0010] This invention collects raw milk and liquid milk samples from different species, origins, and varieties. It gathers spectral fingerprint data for volatile matter, low-to-medium volatile matter, and medium-to-low volatile matter from each sample. A discrimination model is established using training data and validated and optimized using validation data. This overcomes the limitations of using a single indicator for sample discrimination, achieving global feature representation driven by big data and overcoming the insufficient robustness of traditional single-indicator or few-indicator discrimination. Through deep analysis of the fingerprint data, it achieves digital correlation and global representation of complex milk sample attributes, significantly improving discrimination precision. It can simultaneously identify species, feeding methods, geographical origin, and product category / brand. It enables milk sample evaluation, accurate authenticity determination, and efficient traceability. Attached Figure Description
[0011] Figure 1 This is a flowchart of a method for distinguishing between raw milk and liquid milk based on volatile mass spectrometry fingerprinting; Figure 2 This is a schematic diagram of the technical route for a method to distinguish between raw milk and liquid milk based on volatile matter spectral fingerprinting. Figure 3A Tables were created for the fingerprint database of 30%, 60%, and 80% yak milk; Figure 3B Tables were created for training sets containing 30%, 60%, and 80% yak milk. Figure 3C Tables were created for the validation sets of 30%, 60%, and 80% yak milk. Figure 4A OPLS-DA species discrimination model diagram for spectral fingerprints of volatile substances; Figure 4B OPLS-DA bovine discrimination model for spectral fingerprinting of volatile substances; Figure 5A OPLS-DA model for distinguishing UHT milk from North and South regions based on spectral fingerprints of volatile substances; Figure 5B OPLS-DA model for distinguishing between domestic and imported UHT milk based on spectral fingerprints of volatile substances; Figure 6A OPLS-DA origin discrimination model diagram for the fingerprint of volatile substances in mare's milk; Figure 6B Contribution plot of fingerprint characteristic variables for volatile components of horse milk; Figure 6C OPLS-DA origin discrimination model diagram for the fingerprint of volatile substances in Bactrian camel milk; Figure 6D Contribution plot of volatile component fingerprint characteristics variables for Bactrian camel milk from three production areas; Figure 7A OPLS-DA species discrimination model diagram for the spectral fingerprint of low volatile matter in three types of yak milk; Figure 7B Hierarchical cluster analysis (HCA) plot of low volatile matter in yak milk at three different contents; Figure 8A OPLS-DA model for discriminant model of feeding patterns for medium and low volatile matter spectral fingerprints; Figure 8B HCA plot of the mass spectrometry fingerprint of medium and low volatile matter; Figure 9 This is a schematic diagram illustrating the verification process of a method for distinguishing between raw milk and liquid milk based on volatile matter spectral fingerprinting. Detailed Implementation
[0012] First, the technical terms will be explained. Volatile substances group / spectrum: refers to volatile components (including short-chain fatty acids, aldehydes, ketones, alcohols, etc.) that are released through headspace equilibrium at room temperature or slight heating (40℃~45℃) and have significant sensory activity. These substances can be directly and non-destructively collected in situ by PTR-TOF-MS, eliminating the need for complex pretreatment.
[0013] Volatile matter spectrum: refers to the digital feature matrix converted from the raw mass spectrometry signal (containing only mass-to-charge ratio and response intensity dimensions) directly acquired by PTR-TOF-MS without component identification.
[0014] Medium and low volatile substances group / spectrum: refers to characteristic substances that require targeted extraction with specific polar solvents (single-phase or composite reagents), followed by centrifugation and freeze enrichment, and then volatilize under high-temperature desorption conditions in gas chromatography.
[0015] Medium and low volatile matter spectra: refers to the multidimensional digital feature matrix composed of raw mass spectrometry signal fragments acquired by GC-MS without qualitative analysis of components.
[0016] Raw milk: Milk secreted by domesticated livestock during lactation, exhibiting significant species specificity, and its processed liquid form.
[0017] Training sample set: or training set for short, is a database of representative raw milk or liquid milk samples with clear background information, which is used to train and build feature evaluation and authenticity discrimination models.
[0018] Validation sample set: also known as the validation set, is a set of known samples (including real target samples and fake samples) that are independent of the training set and are used specifically to evaluate, correct and optimize model performance.
[0019] The sample to be tested is an unknown dairy product sample to be jointly judged from multiple dimensions, such as species origin, feeding mode, origin traceability, or brand authenticity.
[0020] Raw milk: also known as unprocessed milk, refers to fresh milk produced by dairy livestock such as cows, horses, goats, camels, sheep, buffalo, and yaks without any processing or treatment.
[0021] Liquid milk: refers to fresh milk products that have undergone various heat sterilization processes and are then packaged and sold. Low-temperature pasteurized liquid milk usually requires cold chain transportation and storage, and is also known as pasteurized milk or low-temperature milk. Liquid fresh milk products packaged and sold after ultra-high temperature (UHT) sterilization can be stored at room temperature for a long time, and are also known as UHT milk or room-temperature milk.
[0022] Figure 1 This is a flowchart of the method for distinguishing raw milk and liquid milk based on volatile mass spectrometry fingerprinting according to the present invention.
[0023] like Figure 1 As shown, the method for distinguishing raw milk and liquid milk based on volatile matter spectral fingerprinting includes: sample collection step S10, which fully collects raw milk and liquid milk samples from different species / strains, different regions / origins and different brands / categories.
[0024] A wide variety of representative milk samples (such as raw milk and liquid milk covering multiple species, categories, origins, and different feeding methods) were collected to serve as the physical basis for constructing the fingerprint database. In this embodiment, the collected samples include: ① by species: mare milk, camel milk, goat milk, Holstein cow milk, yak milk, buffalo milk, etc.; ② by origin: two brands from the north (Inner Mongolia), two brands from the south (Fujian), imported brands, mare milk from three banners and counties in Inner Mongolia, Bactrian camel milk from three regions, etc.; ③ by category: organic, ordinary, DHA, etc.; ④ by content: 30%, 60%, and 80% yak milk.
[0025] In the fingerprint data acquisition step S20, fingerprint data is acquired from each of the raw milk and liquid milk samples. The fingerprint data is then obtained through two orthogonal detection systems to obtain fingerprint data of volatile substances, fingerprint data of volatile substances, fingerprint data of medium and low volatile substances, and fingerprint data of medium and low volatile substances.
[0026] Representative breast samples, once collected, proceed through the fingerprint data acquisition module. At this stage, to meet the high-dimensional analysis requirements of machine learning, two orthogonal detection systems must be used to transform the physical samples into four core digital matrix features: ① Fingerprints of volatile substances (non-destructive direct sampling based on PTR-MS / PTR-TOF-MS): Accurately weigh 2.0 g of milk sample without any pretreatment, place it in a 20 mL headspace sampling vial and seal it; place the sealed headspace sampling vial in a constant temperature water bath at 45°C and let it stand for 10 min to promote the volatile substances in the upper gas phase space of the headspace sampling vial to reach a gas-liquid equilibrium state. Subsequently, the injection needle of PTR-MS / PTR-TOF-MS is used to quickly puncture the sealing membrane of the headspace sampling vial and insert into the upper gas phase space in the vial to detect volatile substances.
[0027] Setting of detection parameters: the full-spectrum acquisition rate is 1 spectrum per second, and the cumulative detection duration for a single sample is 60 s; all samples adopt a random detection order to eliminate the influence of systematic errors on the detection results; each sample is scanned continuously for 3 times, and 3 sets of mass spectrometry detection data within the detection duration interval of 40-60 s are selected for average calculation to obtain the average detection value of the sample; at the same time, under the completely same experimental conditions as sample detection, blank air is scanned and detected to obtain the average blank detection value.
[0028] Difference calculation is performed between the average sample detection value and the average blank detection value to obtain effective PTR-TOF-MS mass spectrometry data that can be used for subsequent statistical analysis, and 3 groups of parallel experiments are set for each sample to ensure data reliability. IoniTOF 4.0 software matched with PTR-TOF-MS is used to complete data acquisition and mass spectrogram generation, and mass spectrograms and each molecular ion peak are analyzed with the assistance of PTR-MS Viewer 3.2 software in combination with the NIST online mass spectrometry database. After blank subtraction of the original data from direct headspace injection of milk samples, 495 characteristic mass peaks are obtained, preliminary qualitative identification of the corresponding molecular ions of each peak is carried out by combining NIST library retrieval results and related literatures, and the types of volatile compounds are determined. The original data of relative content of each volatile substance is exported through the Export Traces function of PTR-MS Viewer 3.2, the average value of 3 sets of original data within the detection interval of 40-60 s is calculated as the average sample content, and after subtracting the average blank content, the final relative content of volatile compounds (unit: μg / L) is obtained.
[0029] ② Medium and low volatility substance fingerprinting (targeted extraction based on GC-MS): In the present patent, medium and low volatility substances refer to milk components that are extracted with a specific solvent (i.e., acetonitrile) and gasified under gas chromatography conditions.
[0030] Extraction solvent: a suitable moderately polar single or composite organic solvent that can dissolve and extract characteristic flavor molecules in milk to the greatest extent. Other specific available solvents include methanol, isopropanol, acetonitrile, tetrahydrofuran, ethyl acetate, methanol, ethanol, propanol, benzene, etc., which can be used for extraction alone or in combination. For solvent preparation, the richness of milk extract should be considered first, and secondly, the selection should also take into account the chromatographic column suitable for matching with the solvent.
[0031] The specific steps are as follows: Frozen milk samples are thawed naturally at room temperature and then pre-frozen. Finished milk is pre-frozen directly, then lyophilized and ground into a fine powder after 48 hours. Accurately weigh 200 mg of the lyophilized powder into a 10 mL sterile centrifuge tube, add 200 mg of trisodium citrate and 1.5 mL of acetonitrile, vortex to mix, and then incubate in a 45°C water bath for 10 min (shaking once every 2-3 min). Subsequently, centrifuge at 2500 rpm. -1 Centrifuge for 4 min, incubate at 4℃ for 15 min, and transfer the supernatant to a new centrifuge tube; freeze at -20℃ for 30 min, then centrifuge at 5000 rpm. -1 Centrifuge for 30 seconds, then transfer the supernatant to a vial and seal. Refrigerate at 4°C until analysis. Perform three parallel experiments for each milk sample.
[0032] GC-MS detection parameter settings: Programmed temperature ramp: initial column temperature 60℃, hold for 6 min, then increase at 8℃·min. -1 The temperature was raised to 250°C and held for 15 minutes; the carrier gas was high-purity helium, and the column flow rate was 1.05 mL / min. -1 Total flow rate 2 mL·min -1 Splitless injection, injection port temperature 250℃, injection volume 1μL. Mass spectrometry parameters: electron impact ionization (EI), electron energy 70eV; ion source and interface temperature both 250℃, full scan mode covering the characteristic mass-to-charge ratio range of the target compound.
[0033] Data processing: Data was acquired and processed using a Shimadzu GC-MS solution workstation to generate mass spectra, which were then matched and searched using the NIST mass spectrometry database. Compounds with a retention index (SI) ≥ 80 were selected as effective target compounds, and the percentage quantification was performed using the peak area normalization method to obtain the relative content of each compound.
[0034] In the dataset partitioning step S30, the volatile matter spectral fingerprint data, volatile substance fingerprint data, medium- and low-volatile matter spectral fingerprint data, and medium- and low-volatile substance fingerprint data of each of the raw milk and liquid milk samples are unidirectionally merged into the fingerprint database within the machine learning module along the solid line (this patent uses Simca 14.1 software and Pirouette 4.5). The fingerprint data in the fingerprint database undergoes dataset partitioning and transfer to obtain training set data and validation set data.
[0035] Training set data partitioning: Data flows to training set data nodes through the solid line main channel. The rule is that 90% to 99% of the samples are partitioned from the total sample set in each round for the initial pattern learning of the subsequent model.
[0036] Validation set data partitioning: The remaining data flows to the validation set data nodes through the dashed path, serving as an independent dataset for subsequent cross-evaluation of model performance.
[0037] In data processing step S40, the training set data and validation set data are combined and fed into the data preprocessing module. The system will perform the following low-level operations on the high-dimensional fingerprint data: first, analyze the data structure and perform necessary data truncation (such as truncation of the effective stable range of the mass spectrometry signal); then, conduct pattern exploration and implement necessary data transformations / conversions (such as scaling, noise reduction, etc.) to unify the data dimensions and remove interference noise.
[0038] In model building step S50, the training set data is imported into the model through solid lines to establish validation and optimization nodes, and the modeling method is selected and the parameters are set for debugging. The data structure is analyzed by software to obtain the discriminant model.
[0039] In the discrimination model accuracy verification step S60, the validation set data is run on the discrimination model, and the discrimination results of the model are compared with the authenticity of each known sample in the validation set data to determine the discrimination accuracy of the discrimination model.
[0040] In the model optimization step S70, the discrimination model is optimized based on the discrimination results of the validation set data to obtain the optimized discrimination model.
[0041] In the authenticity traceability step S80, the optimized discrimination model is used to trace the authenticity of the raw milk and liquid milk to be tested.
[0042] After the discrimination model is constructed, testing is performed on unknown test samples: First, unknown test samples are collected, and volatile matter spectral fingerprint data is acquired using the same method as described above, generating independent fingerprint data for the test sample. This fingerprint data is then imported into a unified data preprocessing module along the application process. After format standardization conversion, it is directly input into the finalized and validated model module for intelligent computation. Finally, the trained model accurately outputs the quality evaluation, authenticity determination, and origin traceability results of the test sample along the application process. The output results fully support the practical application of this invention in volatile matter characteristic evaluation, sample authenticity determination, origin and brand traceability, and visualization of test data results.
[0043] The discriminant model can be used to identify and trace the origin of species (cattle, yaks, buffalo, horses, donkeys, camels, goats, sheep), feeding methods (organic and non-organic, grass-fed and grain-fed, grazing and stall-fed, etc.), place of origin (region, province, prefecture (league) city, domestic and foreign), and manufacturer brand.
[0044] Figure 2 This is a technical roadmap diagram illustrating the method for distinguishing between raw milk and liquid milk based on volatile matter spectral fingerprinting. It includes the following aspects: Sample Collection 11: Collect a wide range of representative milk samples. The collection should be comprehensive, including raw and liquid milk from multiple species, varieties, origins, and feeding methods. This will serve as the physical basis for constructing the fingerprint database. The samples collected in this patent include: ① By species / breed: mare milk, camel milk, goat milk, Holstein cow milk, yak milk, buffalo milk, etc.; ② By region / origin: two brands from northern (Inner Mongolia), two brands from southern (Fujian), imported brands, mare milk from three banners / counties in Inner Mongolia, and Bactrian camel milk from three regions; ③ By category: organic, regular, DHA, etc.; ④ By content: 30%, 60%, and 80% yak milk.
[0045] The sampling volume should be sufficient and representative, covering both raw and liquid milk, and different breeds, regions, seasons, and feeding or grazing methods. It is not advisable to repeatedly collect too many homogeneous samples from a single animal, a single pasture, or a single season.
[0046] The samples are divided into "training set data" and "validation set data." The training set dataset is used for model building, and the validation set dataset is used for model validation; the two can be used interchangeably as appropriate. After model validation, if the model accuracy is low and correction or optimization is needed, this portion of the validation set data is added back to the validation set, the model is recalculated or the model parameters are modified, and then the model is validated again using other validation set data for further optimization. Even if the model is very accurate and robust, validation samples should still be included in the training set, allowing the training set to expand indefinitely.
[0047] It is important to note that for both validation and training set data, a thorough understanding of the sample information is essential, including known species or classifications, absolute true and false results, and a clear quality level or adulteration ratio. Incorrect or abnormal samples can interfere with the modeling process; if not removed, they will reduce the model's accuracy and robustness.
[0048] Fingerprint Data Acquisition 12: The collected representative breast samples are then fed into the fingerprint data acquisition module via a solid-line process. In this stage, to meet the high-dimensional analysis requirements of machine learning, two orthogonal detection systems must be used to transform the physical samples into four types of core digital matrix features: volatile matter spectral fingerprint data, volatile matter fingerprint data, low-to-medium volatile matter spectral fingerprint data, and low-to-medium volatile matter fingerprint data.
[0049] Fingerprint database 13 establishment: The volatile matter spectral fingerprint data, volatile matter fingerprint data, medium and low volatile matter spectral fingerprint data, and medium and low volatile matter fingerprint data of each milk sample collected are imported into the software's database to obtain fingerprint database 13. Dataset partitioning: The fingerprint data of each milk sample entered into the fingerprint database are partitioned and transferred within the fingerprint database to obtain training set data 14 and validation set data 15.
[0050] Data Preprocessing 16: First, analyze the fingerprint data structure and perform necessary data truncation (such as truncation of the effective stable range of the mass spectrometry signal); then conduct pattern exploration and implement necessary data transformations / conversions (such as scaling, noise reduction, etc.) to unify the data dimensions and remove interference noise.
[0051] Model Building and Optimization 17: Import the training set data into the model via solid lines to establish validation and optimization nodes, select modeling methods, and adjust and set parameters. Perform data structure analysis using software to obtain the discriminant model. Run the validation set data in the discriminant model and compare the model's discrimination results with the authenticity of each known sample in the validation set data to determine the accuracy of the discriminant model. Optimize the discriminant model based on the discrimination results of the validation set data to obtain the optimized discriminant model.
[0052] The established discrimination models can be: a general model for identifying multiple species and abnormal milk quality; a two-species discrimination model, such as mare's milk vs. cow's milk, camel's milk vs. cow's milk, camel's milk vs. goat's milk, goat's milk vs. cow's milk, cow's milk vs. soy milk or plant materials, cow's milk vs. margarine, etc.; a two-substance discrimination model can also establish a quantitative analysis model for identifying adulteration.
[0053] The fingerprint data of volatile substances, medium and low volatile substances, and medium and low volatile substances in the validation set data are imported into the fingerprint database of the software. Then, the established discrimination model is run to determine the authenticity of each known sample in the validation set data. The model's true or false judgment, or the result of logical grouping, is compared with the actual situation to evaluate the accuracy of the model's discrimination.
[0054] After model validation, if correction or optimization is needed, these validation samples, especially those that were misclassified or not accurately categorized, should be added to the training sample set for recalculation or modification of model parameters. Then, other validation sample sets should be collected and used to validate the optimized model again. Even if the model is very accurate and robust, representative validation samples should still be included in the training sample set to continuously expand the sample size and increase the number of representative samples.
[0055] In this embodiment, the main machine learning module uses Simca 14.1 software and Pirouette 4.5. This invention is a learning and growth system; as the amount of practical samples and the fingerprint database increase, the model becomes more accurate and robust, and the applicability and functionality of the model or model group also increase.
[0056] The optimized discrimination model can trace the authenticity of the raw milk and liquid milk to be tested.
[0057] Figure 9This is a schematic diagram illustrating the verification process of a method for distinguishing between raw milk and liquid milk based on volatile matter spectral fingerprinting. (Example) Figure 9 As shown, First, sample collection 21 is performed, which involves collecting unknown samples to be tested. Then, fingerprint data of the samples to be tested is collected 22. The collected samples are then fed into the fingerprint data acquisition module via a solid line process to collect fingerprint data of volatile substances, volatile substances, medium and low volatile substances, and medium and low volatile substances.
[0058] Data preprocessing 24 involves importing the fingerprint data of the sample to be tested into a unified data preprocessing module along the application process, and then performing format standardization conversion. The processed fingerprint data is then directly input into the finalized, verified, and optimized discrimination model 25 for intelligent computation.
[0059] The evaluation and identification of the traceability results are 26. Finally, the trained model accurately outputs the quality evaluation, authenticity identification and origin traceability results of the test sample along the application process.
[0060] The following examples provide further illustration. Example 1 (1) Sample collection Volatile substance group and / or its mass spectrometry fingerprint sample information:
[0061] Information on medium and low volatile compounds and / or their mass spectra:
[0062] (2) Establishment of fingerprint database, validation set and training set Import the data obtained from the machine test (i.e., fingerprint data collection) into an Excel spreadsheet and integrate them to form a unified fingerprint database; record all fingerprint data as the training sample set (training set), and extract 1% to 10% of the samples from it as the verification sample set (verification set).
[0063] Given the large sample size, the data on low and medium volatile components in yak milk with contents of 30%, 60%, and 80% are used as examples for illustration, as shown in Tables 3A, 3B, and 3C.
[0064] Figure 3A Table 3B shows the establishment of fingerprint databases for 30%, 60%, and 80% yak milk. The horizontal axis header represents the response time, and column F and thereafter represents the signal strength corresponding to the sample. Table 3C shows the establishment of the training set for 30%, 60%, and 80% yak milk.
[0065] like Figure 3AAs shown, a summary table of medium and low volatile matter spectral data obtained by GC-MS detection is called the fingerprint database; the horizontal axis header represents the response time, and column F and thereafter represents the signal intensity corresponding to the sample.
[0066] like Figure 3B As shown, using all volatile substance fingerprint data as the training sample set, we conducted classification and accuracy analysis of the training set itself. The horizontal axis header represents the response time, and the F column and thereafter represent the signal intensity corresponding to the sample.
[0067] like Figure 3C As shown, 10% of the training sample set is randomly allocated as a validation set, and the remaining 90% is used for model training. The trained model is used to predict the category of the validation set samples. By comparing the prediction results with the actual categories, the classification accuracy is calculated, thereby verifying whether the model has the ability to accurately classify samples that were not included in the training. The horizontal axis header represents the response time, and the F column and thereafter represent the signal strength corresponding to the sample.
[0068] (3) Data preprocessing The optimal data preprocessing method was selected. This patent uses Log10 combined with center scaling, which yields the best clustering results. (These results were automatically generated by the software and require no manual intervention.) (4) Model building, validation and optimization (OPLS-DA analysis) After optimization of the preprocessing method and OPLS-DA analysis (using Simca 14.1 software and Pirouette 4.5 in this patent), the results shown in Figures 4-8 were obtained.
[0069] As shown in Figure 4-8, there are significant species differences in the volatile matter groups and / or their mass spectrometric fingerprints of different species and the medium and low volatile matter groups and / or their mass spectrometric fingerprints, which are projected into different regions in the OPLS-DA score map (i.e., different types of clustering have better results). Using the OPLS-DA method to identify the tested samples, if a sample point enters any species cluster, it indicates that the sample belongs to that species; if it falls on the line connecting any two species clusters, it can be determined as milk from a mixture of these two species to varying degrees; for sample points not in the above regions, this method does not perform further discrimination.
[0070] OPLS-DA analysis is a machine learning analysis module that can also be used for discrimination. In practice, an OPLS-DA discriminant analysis module based on OPLS-DA analysis can be used, which integrates multiple OPLS-DA analyses and machine learning.
[0071] This study established an OPLS-DA model (see Figure 4-8) using spectral fingerprint data of volatile substances and fingerprint data of medium and low volatile substances from various raw milks and different types of cow's milk, and conducted internal and external validation (validation results are shown in Table 1-7).
[0072] The OPLS-DA model is well-suited for multi-objective, multi-possibility analysis of completely unknown samples.
[0073] (5) Evaluate and judge the source tracing results by Figure 4B For example (other results are shown in the figure captions), the four types of samples—Holstein milk (yellow), yak milk (blue), goat milk (red), and buffalo milk (purple)—formed their own independent cluster distributions in the model. Yak milk samples were concentrated in the upper left quadrant, while Holstein milk samples were concentrated in the lower left quadrant. The mass spectrometry fingerprint characteristics of the two samples differed significantly from those of goat milk and buffalo milk samples. Goat milk and buffalo milk samples were distributed on the right side of the model. There was some overlap in the clusters between groups, but overall, they could still be effectively distinguished. This indicates that the model can effectively identify and distinguish the above four types of milk.
[0074] (6) Fingerprint data collection and fingerprint database of the tested samples Once the model is established, the results exported after the sample to be tested is processed on the machine (i.e., fingerprint data acquisition) are called fingerprint data.
[0075] (7) Preprocessing of sample data The optimal data preprocessing method was selected to achieve the best clustering results. (This result was automatically generated by the software and requires no manual intervention.) (8) Evaluate and judge the source tracing results Significant species differences were observed in the chromatographic fingerprints of different types of volatile compounds and the fingerprints of medium and low volatile compounds, which were projected into different regions in the OPLS-DA score map. Using the OPLS-DA method to identify the tested sample, if the sample spot falls into any species cluster, the sample belongs to that species; if it falls on the line connecting any two species clusters, it can be determined as milk from a mixture of these two species to varying degrees; if the sample spot does not fall on the line connecting any two species, it belongs to other species or is other abnormal milk.
[0076] Figure 4A and Figure 4B This is a visualization of the OPLS-DA species discrimination model based on volatile mass spectrometry fingerprinting. Figure 4A This is a diagram of the OPLS-DA species discrimination model for spectral fingerprints of volatile substances. Figure 4B This is a diagram of the OPLS-DA bovine discrimination model for spectral fingerprints of volatile substances.
[0077] As shown in Figure 4A, an OPLS-DA model was established for 42 mare milk samples, 42 camel milk samples, 15 goat milk samples, 30 Holstein cow milk samples, 15 yak milk samples, and 15 buffalo milk samples. The preprocessing was Log10 combined with center scaling. Holstein cow milk (yellow), yak milk (light blue), goat milk (red), and buffalo milk (purple) belong to the same Bovidae family and are clustered on the same side (left side) of the figure. Mare milk (green) and camel milk (dark blue) showed significant differences in volatile matter spectral fingerprint characteristics compared to other Bovidae.
[0078] like Figure 4B As shown, Figure 4A The bovine animals in the lower left corner were plotted separately (15 goat milk samples, 30 Holstein milk samples, 15 yak milk samples, and 15 buffalo milk samples). The preprocessing was Log10 combined with center scaling. Yak milk and buffalo milk were relatively close, belonging to different regions from goat milk and Holstein milk. Holstein milk (yellow) and yak milk (blue) were concentrated on the left, while goat milk (red) and buffalo milk (purple) were distributed on the right and partially overlapped. They were clustered along the main axis, indicating that the model can effectively distinguish the volatile component characteristics of milk from different animal species.
[0079] Figure 5A and Figure 5B This is a visualization of the OPLS-DA origin discrimination model based on volatile mass spectrometry fingerprinting. Figure 5A The OPLS-DA model for distinguishing between North and South UHT milk based on the spectral fingerprint of volatile substances is shown in the figure. Figure 5B OPLS-DA model for distinguishing between domestic and imported UHT milk based on the spectral fingerprint of volatile substances.
[0080] like Figure 5A As shown, an OPLS-DA model was established for 5 brands each of northern brands (Inner Mongolia) A and B, and 5 brands each of southern brands (Fujian) C and D. The preprocessing was Log10 combined with Centerscaling. Green represents northern brands from Inner Mongolia and blue represents southern brands from Fujian. The samples were clustered separately along the main axis (horizontal axis), which can effectively distinguish the volatile characteristics of milk from the north and south.
[0081] like Figure 5B As shown, for 20 samples of domestic brands and 5 samples of imported brands, the preprocessing was Log10 combined with center scaling. The two types of samples achieved obvious grouping and clustering along the principal component axis. The samples of domestic brands could be effectively separated along the vertical axis, which can characterize the volatile substance characteristics of each sample and can be used to realize the differentiation of milk brands and the traceability analysis of origin.
[0082] Figure 6A ,6B Visualization diagrams and feature variable contribution diagrams of the OPLS-DA model for determining the origin of mare's milk and Bactrian camel milk based on volatile substance fingerprints, 6C and 6D. Figure 6A OPLS-DA origin discrimination model diagram for the fingerprint of volatile substances in mare's milk. Figure 6B Contribution plot of fingerprint characteristic variables for volatile components of horse milk. Figure 6C OPLS-DA origin discrimination model diagram for the volatile matter fingerprint of Bactrian camel milk. Figure 6D Contribution plot of volatile component fingerprint characteristics variables for Bactrian camel milk from three production areas.
[0083] like Figure 6A As shown, the OPLS-DA model visualization effect is constructed for 10 mare's milk samples from Evenk Banner, 10 mare's milk samples from Sunite Left Banner, and 10 mare's milk samples from Uxin Banner. The preprocessing method is Center Scaling. Factor 1 (contribution rate 38.0%) on the horizontal axis, Factor 2 (contribution rate 19.3%) and Factor 3 (contribution rate 16.3%) on the vertical axis together represent the clustering distribution and degree of difference of mare's milk from different production areas. The red, blue and yellow dots correspond to mare's milk from Evenk Banner, Sunite Left Banner and Uxin Banner, respectively.
[0084] As shown in Figure 6B, for... Figure 6A The corresponding OPLS-DA feature variable contribution plot shows that each point represents the key volatile substance characteristic components of mare's milk from different origins, clearly distinguishing the core difference markers of mare's milk from different origins.
[0085] As shown in Figure 6C, the OPLS-DA model visualization effect is constructed for 10 Bactrian camel milk samples from Hulunbuir, 10 Bactrian camel milk samples from Alashan, and 10 Bactrian camel milk samples from Xinjiang. The preprocessing method is Center Scaling. Factor 1 (contribution rate 47.6%) on the horizontal axis, Factor 2 (contribution rate 21.5%) and Factor 3 (contribution rate 8.9%) on the vertical axis together represent the clustering distribution and degree of difference of Bactrian camel milk from different origins. Yellow, red and brown dots correspond to Bactrian camel milk from Hulunbuir, Xinjiang and Alashan, respectively.
[0086] Figure 6D To and Figure 6C The corresponding OPLS-DA feature variable contribution plot shows that each point represents the key volatile characteristic components of Bactrian camel milk from different origins, clearly distinguishing the core differential markers of Bactrian camel milk from different origins.
[0087] Figure 7A , Figure 7BThe image shows a visualization of the OPLS-DA species discrimination model based on the spectral fingerprint of low-to-medium volatile matter and a hierarchical clustering analysis (HCA) diagram. Figure 7A OPLS-DA species discrimination model diagram for the spectral fingerprint of low volatile matter in three types of yak milk. Figure 7B The diagram shows the hierarchical clustering analysis (HCA) of low volatile matter in three types of yak milk.
[0088] like Figure 7A As shown, an OPLS-DA model was established for 12 samples of 30% yak milk, 12 samples of 60% yak milk, and 12 samples of 80% yak milk. The preprocessing was Log10 combined with Centerscaling. Circles, squares, and triangles represent yak milk with different contents of 30%, 60%, and 80%, respectively, and they are distributed on the right, left, and top sides, respectively. The samples showed tight aggregation within the groups and obvious separation between the groups, demonstrating excellent ability to distinguish the addition ratio.
[0089] Figure 7B In order to be in Figure 7A HCA plots were generated based on the established OPLS-DA model. The cluster boundaries of the three types of samples were clear, and the differences between groups were significant, demonstrating the effective ability to distinguish between low and medium volatile matter fingerprints.
[0090] Figure 8A and Figure 8B The image shows a visualization of the OPLS-DA feeding pattern discrimination model based on low-to-medium volatile matter spectral fingerprinting, along with an HCA diagram. Figure 8A This is a model diagram of OPLS-DA for discriminating feeding patterns based on the spectral fingerprint of low-to-medium volatile matter. Figure 8B HCA diagram of medium- and low-volatile matter spectral fingerprint.
[0091] like Figure 8A As shown, the OPLS-DA model was established for 12 organic milk samples, 12 non-organic milk samples, and 12 DHA-fortified milk samples. The preprocessing was Log10 combined with Centerscaling. Green, blue, and red represent non-organic, organic, and DHA-fortified milk, respectively. The three types of samples showed significant separation between groups and tight aggregation within groups in the projection space, which verifies that the method of the present invention can accurately identify milk type based on the spectral fingerprint of medium and low volatile matter, providing key support for its identification and traceability.
[0092] Figure 8B In order to be in Figure 8A HCA plots were generated based on the established OPLS-DA model. The cluster boundaries of the three types of samples were clear, and the differences between groups were significant, demonstrating the effective distinguishing ability of the spectral fingerprints of medium and low volatile substances.
[0093] OPLS-DA model validation: Table 1. Validation results of the OPLS-DA species discrimination model based on volatile mass spectrometry fingerprinting.
[0094] Table 2. Validation results of the OPLS-DA discriminant model for UHT milk from North and South China based on volatile mass spectrometry fingerprinting.
[0095] Table 3. Validation results of the OPLS-DA discrimination model for domestic and fast-acting UHT milk based on volatile mass spectrometry fingerprinting.
[0096] Table 4. Validation results of the OPLS-DA origin discrimination model for Bactrian camel milk based on volatile component fingerprints.
[0097] Table 5. Validation results of the OPLS-DA origin discrimination model for mare milk based on volatile compound fingerprints.
[0098] Table 6. Validation results of the OPLS-DA species discrimination model based on medium- and low-volatile matter spectral fingerprinting
[0099] Table 7. Validation results of the OPLS-DA feeding pattern discrimination model based on low-to-medium volatile matter spectral fingerprinting
[0100] The embodiments of the present invention have been described above. It can be seen that the present invention, by fully collecting raw milk and liquid milk samples from different species, origins, and varieties, and collecting volatile matter fingerprint data, volatile substance fingerprint data, low- and medium-volatile matter fingerprint data, and medium- and low-volatile substance fingerprint data for each sample, establishes a discrimination model using training set data and verifies and optimizes the discrimination model using validation set data. This overcomes the limitations of using a single indicator to discriminate samples, achieving global feature representation driven by big data, and overcoming the insufficient robustness caused by traditional single-indicator or few-indicator discrimination. Through in-depth analysis of fingerprint data, digital association and global representation of complex attributes of milk samples are achieved, significantly improving the precision of discrimination. It can simultaneously identify species, determine feeding patterns, trace geographical origins, and identify product categories and brands. It can achieve milk sample evaluation, accurate authenticity determination, and efficient traceability.
[0101] This innovative method employs advanced detection technologies such as PTR-MS / PTR-TOF-MS and GC-MS to accurately capture volatile substances, low-to-medium volatile substances, and / or their fingerprint spectra in samples. Combined with machine learning analysis and modeling, it enables the evaluation of dairy samples, accurate identification of authenticity, and efficient traceability. It can rapidly and accurately identify trace adulteration, multi-component mixed adulteration, and variety adulteration with high sensitivity and specificity, thereby improving the accuracy of identification and the reliability of traceability.
Claims
1. A method for distinguishing between raw milk and liquid milk based on volatile matter spectral fingerprinting, characterized in that, include: Sample collection step (S10): Collect raw milk and liquid milk samples from different species / strains, regions / origins and brands / categories. The fingerprint data acquisition step (S20) involves acquiring fingerprint data from each of the raw milk and liquid milk samples, and then detecting the data using two orthogonal detection systems to obtain fingerprint data for volatile substances, fingerprint data for volatile substances, fingerprint data for medium and low volatile substances, and fingerprint data for medium and low volatile substances. In the dataset partitioning step (S30), the volatile matter spectral fingerprint data, volatile matter fingerprint data, medium and low volatile matter spectral fingerprint data, and medium and low volatile matter fingerprint data of each of the raw milk and liquid milk samples are respectively unidirectionally merged into the fingerprint database in the machine learning module along the solid line. The fingerprint data in the fingerprint database is partitioned and transferred to obtain training set data and validation set data. The model building step (S50) involves importing the training set data into the model through solid lines to establish validation and optimization nodes, selecting the modeling method and setting the parameters for debugging, and performing data structure analysis through software to obtain the discriminant model. The accuracy verification step of the discrimination model (S60) involves running the validation set data in the discrimination model and comparing the discrimination results of the model with the authenticity of each known sample in the validation set data to determine the accuracy of the discrimination model. The model optimization step (S70) optimizes the discrimination model based on the discrimination results of the validation set data to obtain the optimized discrimination model; The authenticity traceability step (S80) uses an optimized discrimination model to trace the authenticity of the raw milk and liquid milk to be tested.
2. The method for distinguishing raw milk and liquid milk based on volatile mass spectrometry fingerprinting according to claim 1, characterized in that, In the dataset partitioning step (S30), 90% to 99% of the samples are partitioned from the total sample set in each round to obtain the training set data, and the remaining data is used as an independent dataset to obtain the validation set data.
3. The method for distinguishing raw milk and liquid milk based on volatile mass spectrometry fingerprinting according to claim 2, characterized in that, It also includes a data processing step (S40), which performs data structure analysis on the training set data and validation set data, and performs scaling, noise reduction and other processing to unify the data units and remove interference noise.
4. The method for distinguishing raw milk and liquid milk based on volatile mass spectrometry fingerprinting according to claim 3, characterized in that, The raw milk is collected from cattle, yaks, buffalo, horses, donkeys, camels, goats, and sheep.