Method and kit for identifying freshness of apple juice

Through high-resolution mass spectrometry combined with stoichiometry, fingerprint maps and machine learning algorithms, the subjectivity and accuracy of apple juice freshness detection in the existing technology are solved, and fast and accurate identification of apple juice freshness is achieved.

CN120559129APending Publication Date: 2025-08-29CHINESE ACAD OF INSPECTION & QUARANTINE

Patent Information

Application Number
CN202510728114.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing juice freshness detection methods are easily subjectively affected and are difficult to fully reflect freshness. The number of flavor substances screened by the existing technology is limited, so it is impossible to accurately quantify the freshness of apple juice.

Method used

High-resolution mass spectrometry combined with stoichiometry, the overall changes of small and medium-sized metabolites of apple juice were analyzed through fingerprint map, and the apple juice was clustered using K-mean clustering analysis and orthogonal partial least squares discriminant analysis, and a supervised machine learning algorithm model was constructed, and the identification was performed with a linear support vector machine.

Benefits of technology

It has achieved rapid and accurate identification of apple juice freshness, fast detection speed, high identification accuracy, and stable model, providing an important industry regulatory basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120559129A_ABST
    Figure CN120559129A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of analytical chemistry, and particularly relates to a method and a kit for identifying freshness of apple juice, and the method specifically comprises the following steps: S1, selecting apple juice prepared from known apples with different freshness, and analyzing and detecting the apple juice by using a chromatography-mass spectrometry system to obtain a corresponding fingerprint spectrum; s2, clustering and defining the fingerprints by using K-means clustering analysis and orthogonal partial least squares discriminant analysis based on peak information of the fingerprints corresponding to the apple juices with different freshness in S1 as characteristic variables, and defining the fingerprints as a fresh group or a rotten group to obtain a data set; s3, based on the data set in the S2, constructing an identification model by adopting a supervised machine learning algorithm; s4, to-be-detected unknown apple juice is taken, the fingerprint spectrum of the unknown apple juice is obtained under the same detection conditions in the step S1, and the freshness of the unknown apple juice is identified based on the identification model in the step S3. The method is high in identification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of analytical chemistry, and in particular relates to a method for identifying the freshness of apple juice and a kit thereof. Background Art

[0002] At present, the detection methods used to detect the freshness of fruit juice mainly adopt sensory evaluation and marker evaluation. Sensory evaluation is easily affected by subjective factors, and marker evaluation relies on specific markers, which makes it difficult to fully reflect the freshness.

[0003] For example, prior art CN 105717227 B discloses a method for determining the flavor quality of concentrated apple juice. This method uses headspace solid-phase microextraction gas chromatography-mass spectrometry, ultraviolet-high-performance liquid chromatography, pre-column derivatization-ultraviolet-high-performance liquid chromatography, and differential refractive index-high-performance liquid chromatography to analyze the apple juice. The flavor components in the juice are then classified using principal component analysis and factor analysis for comprehensive scoring. The main flavor components are then screened for flavor quality in the concentrated apple juice. However, this method only screens a limited number of flavor components and cannot accurately quantify the freshness of apples.

[0004] Therefore, there is an urgent need to develop an accurate, convenient and universal method to identify the freshness of freshly squeezed juice. Summary of the Invention

[0005] In light of this, the present invention aims to develop a method for identifying the freshness of apple juice. This method, combining high-resolution mass spectrometry with chemometrics, provides a more precise tool for distinguishing apple juices of varying freshness. Based on fingerprint analysis, this method comprehensively captures the overall changes in small molecule metabolites in apple juice. It does not focus on the analysis of a single indicator or specific component, making it less susceptible to the influence of individual components. It offers rapid detection, high identification accuracy, and a stable model, providing an important technical basis for relevant industry regulation.

[0006] The technical solutions adopted in the present invention include: A first aspect of the present invention provides a method for identifying the freshness of apple juice, comprising: S1. Selecting apples of known different freshness, including fresh apples and rotten apples, and preparing apple juices of different freshness from the apples of different freshness, and analyzing and detecting the apple juices using a chromatography-mass spectrometry system to obtain corresponding fingerprints; S2. Based on the peak information of the fingerprints corresponding to the apple juices of different freshness described in S1 as feature variables, the fingerprints are clustered and defined using K-means cluster analysis and orthogonal partial least squares discriminant analysis, and the fingerprints are defined as fresh group or spoiled group to obtain a data set; S3. Based on the data set described in S2, a supervised machine learning algorithm is used to build an identification model; S4. Take unknown apple juice to be tested, use the same detection conditions as S1 to obtain the fingerprint of the unknown apple juice, and identify the freshness of the unknown apple juice based on the identification model described in S3.

[0007] In some embodiments, the step S2 further includes: based on the fingerprint map described in S1, using orthogonal partial least squares discriminant analysis to extract difference variables from the apple juice data of different freshness described in S1, combining S-plot analysis and VIP score to screen out markers with significant differences, and based on the markers, assisting in verifying the results of the identification model.

[0008] In some embodiments, the apple juices of different freshness in step S1 are subjected to solvent extraction to obtain an extract, and the extract is analyzed and detected using a chromatography-mass spectrometry system, wherein the solvent extraction includes ultrasonic-assisted extraction, and the solvent is a mixed solution including one or more of methanol, ethanol, propanol, ethyl acetate, acetonitrile, methanol or water.

[0009] In some embodiments, the solvent is composed of a mixed solution of acetonitrile, methanol and water, and the volume ratio of acetonitrile, methanol and water is (0.5-2):(0.5-2):2.

[0010] In some embodiments, the step S2 of clustering and defining the fingerprint using K-means cluster analysis and orthogonal partial least squares discriminant analysis includes: S2-1. Perform unsupervised grouping of the fingerprint data by K-means cluster analysis, calculate the silhouette coefficient and Davis-Boulding index when the number of clusters is 2-8, comprehensively refer to the silhouette coefficient and Davis-Boulding index, give priority to the group with a silhouette coefficient higher than 0.1, and select the cluster number corresponding to the group with the lowest Davis-Boulding index among the groups with a silhouette coefficient higher than 0.1 as the optimal cluster number; S2-2 is based on the K-means clustering results, and the orthogonal partial least squares discriminant analysis is used for analysis. The significance of the model is verified by cross-validation variance analysis, and the model with the largest F statistic is selected. p The group division scheme with the smallest value defines apple juices of different storage degrees as fresh group or spoiled group.

[0011] In some embodiments, the S4 step includes: S4-1, inputting the fingerprint of the unknown apple juice to be tested into the identification model, and calculating its classification probability using the linear decision boundary built into the identification model; S4-2. Determine whether the data in S4-1 is a fresh group or a rotten group based on a preset probability threshold. If the probability threshold is ≥0.8, it is determined to be a fresh group, and if the probability threshold is ≤0.2, it is determined to be a rotten group. If the probability value is between 0.2-0.8, perform secondary verification by repeating the test or introducing auxiliary indicators.

[0012] In some embodiments, the chromatography-mass spectrometry system uses the following chromatographic detection conditions: Chromatographic column: WATERS ACQUITY UPLC BEH C18 column, specifications: 2.1×100 mm, 1.7 μm; mobile phase A: methanol; mobile phase B: 0.1% formic acid-water solution; flow rate: 0.2 mL / min; injection volume: 2 μL; column temperature: 40 °C, The chromatography adopts gradient elution, and the gradient elution conditions include: starting, mobile phase A: 2%; 0-10 minutes, mobile phase A: 2-50%; 10-14 minutes, mobile phase A: 50-98%; 14-17 minutes, mobile phase A: 98%; 17-17.1 minutes, mobile phase A: 98-2%; 17.1-20 minutes, mobile phase A: 2%; And / or, the mass spectrometry detection conditions of the chromatography-mass spectrometry system are: Scan mode: full scan + secondary fragmentation, positive and negative ion switching; resolution: primary 70,000, secondary 15,000; capillary temperature: 350 °C; auxiliary gas temperature: 320 °C; sheath gas flow rate: 40 arb; auxiliary gas flow rate: 10 arb; S-lens RF voltage: 55 V; spray voltage: +3.2 kV or -3.0 kV; scan range: m / z 70-1050 Da.

[0013] In some embodiments, the supervised machine learning algorithm is selected from any one of the following algorithms: random forest, k-nearest neighbor, support vector machine, linear discriminant analysis, orthogonal partial least squares-discriminant analysis, or linear support vector machine.

[0014] In some embodiments, the markers include: 4-guanidinobutyraldehyde, cytosine, choline sulfate, maleate, coniferyl alcohol, sulfoacetone, acetyl-L-carnitine, panthenol, 4-(2-aminophenyl)-2,4-dioxobutyrate, L-hypoglycine, 4-methylene-L-glutamate, adenosine, 2,3-dihydro-3-hydroxy-2-oxo-1H-indole-3-acetic acid methyl ester, 3-(3,4-dihydroxyphenyl) lactic acid, and caffeic acid 3-glucoside.

[0015] A second aspect of the present invention discloses a kit for identifying the freshness of apple juice, comprising (a) a data analysis module, wherein the (a) data analysis module comprises a data analysis USB flash drive, wherein the data analysis USB flash drive comprises the following (a-1) to (a-3): (a-1) an identification model constructed based on the supervised machine learning algorithm described in the first aspect of the present invention; (a-2) Data preprocessing R scripts that automatically perform data correction, binning, and normalization; (a-3) Python that automatically calls the identification model, performs unknown sample identification, and outputs the results; or, the kit includes the (a) data analysis module and the following (b) to (e) modules: (b) a pre-treatment consumables module, comprising a filter membrane and a high-speed centrifuge tube; (c) a solvent module, comprising a solvent, the solvent being consistent with the extraction solvent in S1 described in the first aspect of the present invention, for obtaining an extract of the unknown apple juice sample, and performing analysis and detection based on the extract using a chromatography-mass spectrometry system, wherein the solvent is a mixed solution of acetonitrile, methanol, and water in a volume ratio of 0.5-2:0.5-2:2; (d) an internal standard module, wherein the internal standard module comprises 2-chloro-L-phenylalanine; (e) Operation guide module, including standardized testing procedures and troubleshooting solutions.

[0016] The present invention has at least the following beneficial effects: (1) The present invention uses high-resolution mass spectrometry combined with multivariate statistical analysis methods to analyze the changes in the chemical composition of apple juice during the spoilage process as the storage time increases based on the fingerprint spectrum basic system. It can comprehensively and unbiasedly capture complex chemical components and reveal subtle differences in apple juice samples at different levels of spoilage without pre-setting assumptions. It has high sensitivity and wide applicability. Based on the fingerprint spectrum basic system, K-means clustering analysis (K-means clustering) and orthogonal partial least squares discriminant analysis (OPLS-DA) are used to cluster apple juices of different storage times and define freshness and spoilage. A linear support vector machine (Linear SVM) is used to establish an identification model. This model can successfully distinguish between fresh and spoiled freshly squeezed apple juice. Therefore, this method has fast detection speed, high identification accuracy, and a stable model, which improves the efficiency and accuracy of freshness identification of freshly squeezed apple juice and provides an important technical basis for relevant industry supervision.

[0017] (2) The present invention provides a kit for identifying the freshness of apple juice. Using the kit to identify the freshness of apple juice is simple to operate, requires a short analysis time, requires only a small amount of sample, requires no or minimal pretreatment, minimizes sample loss, and has a fast detection speed, high sensitivity, and high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is a PCA score graph of freshly squeezed apple juice samples of different freshness provided by one embodiment of the present invention; Figure 2a A schematic diagram of the results of partial least squares discriminant analysis (OPLS-DA) of fresh and spoiled freshly squeezed apple juice samples provided by one embodiment of the present invention; Figure 2b A schematic diagram of verifying the reliability of OPLS-DA using 200 permutation tests on fresh and spoiled apple juice samples according to one embodiment of the present invention; Figure 3 A schematic diagram of the S-plot analysis results provided by one embodiment of the present invention; Figure 4 A schematic diagram of VIP analysis results provided by one embodiment of the present invention; Figure 5a A schematic diagram of the distribution of the marker 4-guanidinobutyraldehyde in different groups provided in one embodiment of the present invention; Figure 5b A schematic diagram of the distribution of the marker cytosine in different groups provided in one embodiment of the present invention; Figure 5c A schematic diagram of the distribution of the marker choline sulfate in different groups provided in one embodiment of the present invention; Figure 5d A schematic diagram of the content distribution of the marker maleic acid amide salt in different groups provided in one embodiment of the present invention; Figure 5e A schematic diagram of the content distribution of coniferyl alcohol, a marker, in different groups according to one embodiment of the present invention; Figure 5f A schematic diagram of the content distribution of the marker styraclostrobin in different groups provided in one embodiment of the present invention; Figure 5gA schematic diagram of the content distribution of the marker acetyl-L-carnitine in different groups provided in one embodiment of the present invention; Figure 5h A schematic diagram of the distribution of the marker ubiquinol in different groups according to one embodiment of the present invention; Figure 5i A schematic diagram of the content distribution of the marker 4-(2-aminophenyl)-2,4-dioxobutyrate in different groups provided in one embodiment of the present invention; Figure 5j A schematic diagram of the content distribution of the marker L-hypoglycine in different groups provided in one embodiment of the present invention; Figure 5k A schematic diagram of the distribution of the marker 4-methylene-L-glutamate in different groups provided in one embodiment of the present invention; Figure 5l A schematic diagram of the distribution of the marker adenosine in different groups according to one embodiment of the present invention; Figure 5m A schematic diagram of the content distribution of the marker methyl 2,3-dihydro-3-hydroxy-2-oxo-1H-indole-3-acetate in different groups provided in one embodiment of the present invention; Figure 5n A schematic diagram of the content distribution of the marker 3-(3,4-dihydroxyphenyl) lactate in different groups provided in one embodiment of the present invention; Figure 5o A schematic diagram showing the distribution of the content of the marker caffeic acid 3-glucoside in different groups according to one embodiment of the present invention; Figure 6a In one embodiment of the present invention, random forests are used to predict the classification accuracy of fresh and spoiled freshly squeezed apple juice samples; Figure 6b According to one embodiment of the present invention, the classification accuracy of fresh and spoiled apple juice samples was predicted using k-nearest neighbor. Figure 6c In one embodiment of the present invention, a support vector machine is used to predict the classification accuracy of fresh and spoiled freshly squeezed apple juice samples; Figure 6d According to one embodiment of the present invention, linear discriminant analysis is used to predict the classification accuracy of fresh and spoiled freshly squeezed apple juice samples; Figure 6e According to one embodiment of the present invention, orthogonal partial least squares-discriminant analysis is used to predict the classification accuracy of fresh and spoiled freshly squeezed apple juice samples; Figure 6f According to one embodiment of the present invention, a linear support vector machine is used to predict the classification accuracy of fresh and spoiled freshly squeezed apple juice samples; Figure 7A Venn diagram of compounds extracted by four pretreatment methods provided in one embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0021] Furthermore, in the description of the present invention, unless otherwise specified, “plurality” in the present invention means two or more.

[0022] In some embodiments of the present invention, the fresh apples and rotten apples are completely fresh and rotten apples visible to the naked eye.

[0023] According to one aspect of the present invention, the present invention provides a method for identifying the freshness of apple juice, comprising: S1. Selecting apples of known different freshness, including fresh apples and rotten apples, and preparing apple juices of different freshness from the apples of different freshness, and analyzing and detecting the apple juices using a chromatography-mass spectrometry system to obtain corresponding fingerprints; S2. Based on the peak information of the fingerprints corresponding to the apple juices of different freshness described in S1 as feature variables, the fingerprints are clustered and defined using K-means cluster analysis and orthogonal partial least squares discriminant analysis, and the fingerprints are defined as fresh group or spoiled group to obtain a data set; S3. Based on the data set described in S2, a supervised machine learning algorithm is used to build an identification model; S4. Take unknown apple juice to be tested, use the same detection conditions as S1 to obtain the fingerprint of the unknown apple juice, and identify the freshness of the unknown apple juice based on the identification model described in S3.

[0024] The method of the present invention utilizes high-resolution mass spectrometry combined with multivariate statistical analysis to systematically analyze the chemical composition changes of apple juice during its spoilage process based on fingerprints. This method comprehensively and unbiasedly captures complex chemical compositions and reveals subtle differences between samples at different degrees of spoilage. Without requiring pre-defined assumptions, it exhibits high sensitivity and broad applicability. K-means clustering and orthogonal partial least squares discriminant analysis (OPLS-DA) are used to cluster and define known apple juice samples of varying freshness. A discrimination model is established using a linear support vector machine (Linear SVM), successfully distinguishing between fresh and spoiled apple juice. This method boasts rapid detection, high discrimination accuracy, and a stable model, improving the efficiency and accuracy of freshness identification for freshly squeezed apple juice and providing an important technical basis for relevant industry regulation.

[0025] In some embodiments of the present invention, the fingerprints are analyzed and processed using K-means clustering and orthogonal partial least squares discriminant analysis (OPLS-DA) to cluster apple juices of varying storage duration and define freshness and spoilage. This is because, in Example 1, using freshly squeezed apple juice as an example, samples were divided into eight natural groups based on storage time. However, the composition of juices stored for three days and freshly picked apples likely did not change significantly, and from a microscopic perspective, they should both be classified as fresh. Consequently, the primary problem is the lack of clear data to define the boundary between freshness and spoilage. Currently, mainstream PCA analysis primarily assesses clustering based on the researcher's subjective perception of the sample distribution in principal component space. Due to the lack of an objective, unified quantitative standard, different researchers' perception and understanding of data distribution vary, resulting in a high degree of subjectivity and uncertainty in the sample groupings derived from PCA results. To address this issue, the present invention first employed K-means cluster analysis, using the Silhouette Score and Davies-Bouldin Index as evaluation metrics to determine the optimal number of statistical groupings for apple juice samples. Furthermore, the OPLS-DA method, leveraging its powerful feature information extraction capabilities, maximized inter-group differences and minimized intra-group differences to identify the most appropriate experimental groups to be grouped together. This allowed for a clear distinction between fresh and spoiled apple juice groups, providing a more scientific and accurate basis for freshness classification of apple juice.

[0026] In some embodiments, the unknown apple juice to be tested in step S4 is freshly squeezed apple juice, and the apple variety used is the same as the variety of apples of different freshness in step S1.

[0027] In some embodiments, the freshly squeezed apple juice is prepared by homogenizing and sieving the edible portion of apples. The apples are of varying freshness (e.g., from visibly fresh to completely rotten). The varying freshness refers to the number of days the apples have been stored under the same conditions after being picked. This ensures that changes in the composition of the apple juice are solely related to storage time and are not affected by conditions such as temperature and humidity.

[0028] In some embodiments, the step S2 further includes: based on the fingerprint map described in S1, using orthogonal partial least squares discriminant analysis to extract difference variables from the data of apple juice of different freshness described in S1, combining S-plot analysis with VIP score to screen out markers with significant differences, and based on the markers, assisting in verifying the results of the identification model.

[0029] In some embodiments, the apple juices of different freshness in step S1 are subjected to solvent extraction to obtain an extract, and the extract is analyzed and detected using a chromatography-mass spectrometry system, wherein the solvent extraction includes ultrasonic-assisted extraction, and the solvent is a mixed solution including one or more of methanol, ethanol, propanol, ethyl acetate, acetonitrile, methanol or water.

[0030] In some embodiments, the solvent is composed of a mixed solution of acetonitrile, methanol, and water, wherein the volume ratio of acetonitrile, methanol, and water is (0.5-2):(0.5-2):2. Extraction using this extraction solvent has high extraction efficiency, the most complete extraction, good fiber removal effect, low instrument contamination, and high data stability and quality, good fitting effect, high significance, and high prediction accuracy.

[0031] In some embodiments, the extraction comprises the following steps: (1) mixing the apple juices of different freshness with the solvent in a volume ratio of 1:2, wherein the extraction system comprises methanol, acetonitrile, and water in a volume ratio of 1:1:2, to obtain a mixed solution; (2) subjecting the mixed solution to vortexing and ultrasonic dispersion treatment in sequence; (3) centrifuging at a speed of not less than 10,000 r / min at 4°C for 8 to 12 minutes; (4) separating the supernatant and filtering it through a hydrophobic filter membrane to obtain a filtrate; and (5) storing the filtrate at -20°C until analysis and detection. This method yields the largest number and variety of compounds.

[0032] In some embodiments, the step S2 of clustering and defining the fingerprint using K-means cluster analysis and orthogonal partial least squares discriminant analysis includes: S2-1. Perform unsupervised grouping of the fingerprint data by K-means cluster analysis, calculate the silhouette coefficient and Davis-Boulding index when the number of clusters is 2-8, comprehensively refer to the silhouette coefficient and Davis-Boulding index, give priority to the group with a silhouette coefficient higher than 0.1, and select the cluster number corresponding to the group with the lowest Davis-Boulding index among the groups with a silhouette coefficient higher than 0.1 as the optimal cluster number; S2-2 is based on the K-means clustering results, and the orthogonal partial least squares discriminant analysis is used for analysis. The significance of the model is verified by cross-validation variance analysis, and the model with the largest F statistic is selected. p The grouping scheme with the smallest value defines apple juices of different freshness as fresh group or spoiled group.

[0033] In some embodiments, the S4 step includes: S4-1, inputting the fingerprint of the unknown apple juice to be tested into the identification model, and calculating its classification probability using the linear decision boundary built into the identification model; S4-2. Determine whether the data in S4-1 is a fresh group or a rotten group based on a preset probability threshold. If the probability threshold is ≥0.8, it is determined to be a fresh group, and if the probability threshold is ≤0.2, it is determined to be a rotten group. If the probability value is between 0.2-0.8, perform secondary verification by repeating the test or introducing auxiliary indicators.

[0034] The introduction of auxiliary indicators for secondary verification includes: (S4-2a) extracting characteristic variables with a VIP score ≥ 1.5 from the fingerprint of the unknown juice; (S4-2b) calculating the Pearson correlation coefficient between the characteristic variable and the aforementioned marker; (S4-2c) When the number of variables with an absolute value of the correlation coefficient ≥ 0.6 exceeds 50% of the total number of variables, they are classified into corresponding groups based on the sign of the correlation coefficient.

[0035] In some embodiments, the chromatographic detection conditions used in the chromatography-mass spectrometry system are as follows: chromatographic column: WATERS ACQUITY UPLC BEH C18 chromatographic column, specifications: 2.1×100 mm, 1.7 μm; mobile phase A: methanol; mobile phase B: 0.1% formic acid-water solution; flow rate: 0.2 mL / min; injection volume: 2 μL; column temperature: 40°C; The above-mentioned chromatographic conditions can effectively improve the separation of each component in the sample and reduce cross-interference between components, thereby increasing the detection sensitivity and resolution of the target compound. Through precise mobile phase control and optimized chromatographic conditions, different components can be effectively separated within a shorter analysis time, thereby improving the accuracy and repeatability of mass spectrometry detection. At the same time, lower flow rates and higher column temperature settings also help maintain good separation effects, improving the stability and reliability of detection results.

[0036] The chromatography employed gradient elution: The gradient elution conditions include: starting, mobile phase A: 2%; 0-10 minutes, mobile phase A: 2-50%; 10-14 minutes, mobile phase A: 50-98%; 14-17 minutes, mobile phase A: 98%; 17-17.1 minutes, mobile phase A: 98-2%; 17.1-20 minutes, mobile phase A: 2%; adopting the above elution conditions, the separation effect of each component is good. Gradient elution adjusts the composition of the mobile phase so that compounds of different polarities can be effectively separated in different elution time periods. The above gradient elution conditions are obtained by the inventor based on the experimental exploration of the component composition and physicochemical characteristics of freshly squeezed apple juice. The above elution conditions not only improve the resolution of each component in the sample, but also optimize the service life of the chromatographic column, reduce the accumulation of pollutants, and help to improve the reproducibility of the analysis and the accuracy of the separation, thereby ensuring the accuracy of subsequent mass spectrometry analysis.

[0037] And / or, in some embodiments, the mass spectrometry detection conditions of the chromatography-mass spectrometry system are: Scan mode: Full scan + secondary fragmentation, positive and negative ion switching, Full scan + ddMS 2 Resolution: primary 70,000, secondary 15,000; capillary temperature: 350°C; auxiliary gas temperature: 320°C; sheath gas flow rate: 40 arb; auxiliary gas flow rate: 10 arb; S-lens RF voltage: 55 V; spray voltage: +3.2 kV or -3.0 kV; scan range: m / z 70-1050 Da. The mass spectrometer is a high-resolution mass spectrometer.

[0038] The mass spectrometry detection conditions described above provide a high response value for the detection of apple juice components, and the data have good accuracy and reproducibility.

[0039] In some embodiments, the fingerprint is a full-scan fingerprint. The method of the present invention is based on fingerprint data for analysis, and can comprehensively and systematically capture information on all chemical components in the sample without pre-selecting target compounds. It has high sensitivity and high resolution and is particularly suitable for multi-component analysis of complex samples. In addition, the fingerprint can intuitively display the differences in the chemical composition of the samples, providing an efficient and reliable basis for identifying the freshness of apple juice. The mass spectrometer is a high-resolution mass spectrometer. Therefore, the accuracy of compound identification is high. The full scan can cover all detectable chemical components in the sample. By widely collecting mass spectrometry data across the entire mass range, it provides more comprehensive information for subsequent qualitative analysis. Compared with the targeted scan of specific target substances, the full scan can capture new components that may exist in the sample, thereby improving the comprehensiveness and sensitivity of the analysis. In addition, the full scan data can generate more representative and reliable fingerprints, so that the complexity and diversity of the sample can be more fully analyzed and characterized.

[0040] According to an embodiment of the present invention, the fingerprint data is data on the changes in the marker components over the storage time of the apple. Therefore, through the correlation between the changes in the marker and the storage time of the apple, the freshness of the apple juice can be predicted and reversely verified based on the marker.

[0041] In some embodiments, the supervised machine learning algorithm is selected from any one of the following: random forest (RF), k-nearest neighbor (KNN), support vector machine (SVM), linear discriminant analysis (LDA), orthogonal partial least squares-discriminant analysis (OPLS-DA), or linear support vector machine (Linear SVM). These algorithms can fully exploit the characteristic information of apple juice samples from multiple dimensions when constructing a discriminant model, avoiding the limitation of relying solely on a single feature to judge apple juice freshness. By comprehensively analyzing various features, the model has greater adaptability and robustness, thereby providing more accurate and reliable assurance for apple juice quality monitoring.

[0042] In some embodiments, the markers include: 4-Guanidinobutanal, Cytosine, Cholinesulfate, Maleamate, Coniferylalcohol, Sulcatone, Acetyl-L-carnitine, Pantothenol, 4-(2-Aminophenyl)-2,4-dioxobutanol, One or more of the following compounds should be considered: methyl 2,3-dihydro-3-hydroxy-2-oxo-1H-indole-3-acetate, methyl 3-(3,4-dihydroxyphenyl) lactate, and caffeic acid 2-glucoside. The levels of these compounds change with the storage time of apple juice and can be used to assess the freshness of apple juice.

[0043] In some embodiments, when extracting apple juice samples, an internal standard, 2-chloro-L-phenylalanine, is added to the extraction solvent, and its concentration in the extraction solvent is 1 μg / mL. The addition of the internal standard can, on the one hand, correct experimental errors and ensure the reproducibility of results. Metabolite responses may be lost or changed during sample extraction and instrument detection (such as mass spectrometry signal fluctuations). The internal standard can quantify these errors and compensate for them, eliminating the impact of differences in experimental conditions (such as instrument stability differences between multiple batches of data, sample size differences caused by human factors, etc.) on the results. On the other hand, it can correct the deviation of the mass spectrometry mass axis, so that the mass number accuracy of the metabolite qualitative determination is controlled within △5ppm, ensuring the accuracy of the results.

[0044] A second aspect of the present invention provides a kit for identifying the freshness of apple juice, comprising (a) a data analysis module, wherein the (a) data analysis module comprises a data analysis USB flash drive, wherein the data analysis USB flash drive comprises the following (a-1) to (a-3): (a-1) an identification model constructed based on the supervised machine learning algorithm described in the first aspect of the present invention; (a-2) Data preprocessing R scripts that automatically perform data correction, binning, and normalization; (a-3) Python that automatically calls the identification model, performs unknown sample identification, and outputs the results; or, the kit includes the (a) data analysis module and the following (b) to (e) modules: (b) a pre-treatment consumables module, comprising a filter membrane and a high-speed centrifuge tube; (c) a solvent module, comprising a solvent, the solvent being consistent with the extraction solvent in S1 described in the first aspect of the present invention, for obtaining an extract of the unknown apple juice sample, and performing analysis and detection based on the extract using a chromatography-mass spectrometry system, wherein the solvent is a mixed solution of acetonitrile, methanol, and water in a volume ratio of 0.5-2:0.5-2:2; (d) an internal standard module, wherein the internal standard module comprises 2-chloro-L-phenylalanine; (e) Operation guide module, including standardized testing procedures and troubleshooting solutions.

[0045] The present invention utilizes the kit to extract unknown small molecule metabolites in apple juice, which has the advantages of simple operation, short analysis time, only a trace amount of sample required for detection, no or only little pretreatment required, little sample loss, fast detection speed, high sensitivity and accuracy.

[0046] Below in conjunction with embodiment, scheme of the present invention will be explained.It will be appreciated by those skilled in the art that the following examples are merely for illustration of the present invention and should not be considered as limiting the scope of the present invention.Unindicated specific technology or condition in the embodiment, according to the technology or condition described in the document in this area or according to product specification sheet, carry out.Reagents used therein or instrument do not indicate manufacturer, are all conventional products that can be obtained by commercial means, for example, can be purchased from Sigma company.

[0047] Taking apple juice as an example, let's establish a method for identifying the freshness of apple juice: Example 1 Using the method of the embodiment of the present invention, apple juice with a known storage time was tested, clustered, and classified into fresh and spoiled. A model for identifying apple juice freshness was constructed and its accuracy was verified, as follows: 1. Instruments and Reagents Ultra-high-performance liquid chromatography-orbitrap mass spectrometry (Dionex Ultimate 3000 Series / QExactive) (Thermo Fisher Scientific); electronic balance (precision 0.0001 g, sensitivity 0.01 g); vortex mixer; high-speed centrifuge (minimum speed not less than 8000 rpm); L-2-chlorophenylalanine (purity ≥98%, Aladdin, China); methanol (chromatographically pure, Fisher Scientific, USA); acetonitrile (chromatographically pure, Fisher Scientific, USA); formic acid (chromatographically pure, J&K, China); and Milli-Q water purifier (Millipore, Germany).

[0048] 2. Experimental Methods 2.1 Sample Collection and Processing: Fuji apples were collected from apple picking orchards in Beijing. Equidistant sampling was used to ensure representative sampling. Fresh, ripe, and undamaged apples were selected from various locations on each tree, including the top, bottom, left, right, inside, and outside. The apples were divided into eight groups, each containing seven apples. These apples were then stored in a constant temperature and humidity chamber with controlled temperature (23°C) and relative humidity (65%) to simulate the post-harvest decay process. Starting from the first day after harvest, samples were collected from each group (seven apples) every three days and labeled J1 (0 days), J2 (3 days), J3 (6 days), J4 (9 days), J5 (12 days), J6 (15 days), J7 (18 days), and J8 (21 days). These samples were processed to prepare apple juice representing different stages of decay.

[0049] Apple juice was squeezed from the apples using a juicer. Large debris and residue were removed through a 100-mesh filter. 1 mL of the supernatant was added to 2 mL of the extraction solvent (methanol:acetonitrile:water = 1:1:2 (v / v / v)) containing 1 μg / mL L-2-chlorophenylalanine. The mixture was sonicated for 30 minutes and centrifuged (12,000 rpm, 4°C, 10 minutes). The supernatant was filtered through a 0.22 µm PTFE membrane and stored at -20°C until analysis.

[0050] 2.2 Chromatographic Analysis and Mass Spectrometry Conditions: Chromatographic analysis was performed using a WATERS ACQUITY UPLC BEH C18 column (2.1 × 100 mm, 1.7 μm). Mobile phase A consisted of methanol and mobile phase B consisted of a 0.1% formic acid-water solution. The flow rate was set at 0.2 mL / min, the injection volume was 2 μL, and the column temperature was maintained at 40°C. A gradient elution was used, as shown in Table 1. Under these conditions, the components in apple juice were effectively separated and analyzed.

[0051] Table 1 Gradient elution program

[0052] Mass spectrometry conditions: Full scan + ddMS was used for mass spectrometry analysis 2 mode (positive and negative ion switching), primary resolution 70000, secondary resolution 15000, scan range m / z 70-1050 Da, capillary temperature: 350 °C, auxiliary gas temperature: 320 °C, sheath gas flow rate: 40 arb, auxiliary gas flow rate: 10 arb, S-lens RF: 55 V, spray voltage: +3.2 kV or -3.0 kV.

[0053] 2.3 Data Acquisition and Processing: Compound Discoverer v3.3 software was used for deconvolution, peak alignment, and library searching of high-resolution mass spectrometry data. To eliminate instrumental errors between multiple batches of data, the output data were first corrected for metabolite mass and peak area using internal standards using R v4.0.5 software. To facilitate multivariate statistical analysis, mass binning was performed at 0.2 Da intervals using R v4.0.5 software, and Pareto scaling was used for normalization, with each variable being mean-centered and divided by the square root of its standard deviation.

[0054] Multivariate statistical analysis, model building, and model evaluation, including K-means clustering, principal component analysis (PCA), random forest (RF), k-nearest neighbor (KNN), support vector machine (SVM), linear discriminant analysis (LDA), orthogonal partial least squares discriminant analysis (OPLS-DA), and linear support vector machine (Linear SVM), were implemented in multiple stages using MetaboAnalyst v6.0, Python v3.13.1, and SIMCA v14.1 software.

[0055] The marker content change graph was drawn using GraphPad Prism v9.3.1 software.

[0056] 3. Results and Discussion The freshness of freshly squeezed apple juice depends on the overall changes in the small molecule metabolites therein. The inventors used the QExactive high-resolution mass spectrometer to collect high-resolution data on apple juice samples of different freshness after different storage times, and finally obtained chemical component fingerprints covering the range of m / z 70-1050 Da. However, due to the complexity of the sample components, the overlap of signal peaks and the masking of small differences, these fingerprints are complex and contain a large number of subtle differences, which are difficult to distinguish directly by naked eye or intuitive observation. Therefore, it is necessary to further combine chemometric methods to perform multivariate statistical analysis on high-resolution mass spectrometry data to explore the potential regular differences in the data.

[0057] The inventors first used unsupervised multivariate analysis to reveal underlying patterns and compositional variations within samples, allowing them to cluster and define fresh and spoiled apple juice. They then applied supervised methods to optimize the differentiation of apple freshness and construct a model for identifying fresh and spoiled apple juice.

[0058] 3.1 Clustering and definition of fresh and spoiled apple juice: First, PCA is used to reveal the changes in the composition of apple juice over time during storage. The results are as follows: Figure 1As shown, samples gradually separated along the principal components. Early samples clustered tightly in PC1 and PC2 space. As storage age increased, samples gradually became more dispersed, but the boundary between freshness and spoilage remained difficult to quantify. Specifically, groups J1, J2, and J3 clustered closely, indicating similar chemical composition during the initial storage period. In contrast, groups J6, J7, and J8 exhibited greater dispersion and were clearly separated from samples at the initial storage period, reflecting significant compositional changes that may be caused by the accumulation of metabolic byproducts. Although an overall trend of separation between groups with different storage times was observed, some overlap still existed between some groups (such as J4 and J5). Using subjective perception to assess clustering lacks an objective, unified quantitative standard, and defining the boundary between freshness and spoilage is highly subjective and uncertain. Consequently, the greatest challenge is the lack of clear data support, necessitating further clustering and definition. To establish a quantifiable standard for defining fresh and spoiled apple juice, the inventors used K-means clustering to determine the optimal number of sample groups. The accuracy and reliability of the grouping number were evaluated using the Silhouette Score and Davies-Bouldin Index. The results, shown in Table 2, show that compared with other clustering methods (Silhouette Score < 0.1), the Silhouette Scores for the two-cluster and three-cluster solutions were significantly higher, at 0.27 and 0.24 (greater than 0.1), respectively, indicating the best inter-group separation. Furthermore, the Davies-Bouldin Index for the two-cluster solution was the lowest of the two groups, at 1.89, indicating good compactness and separation. Furthermore, based on the above data, increasing the number of clusters beyond three did not significantly improve clustering quality and resulted in insufficient inter-group separation. Because clear inter-group separation was prioritized, the two-cluster solution was selected.

[0059] Table 2 K-means clustering results

[0060] After determining the number of groups, in order to further refine the boundary between fresh and spoiled apple juice in terms of storage time and thus clearly define the fresh and spoiled groups, the inventors used OPLS-DA and tested seven models using CV-ANOVA. p The performance of each model was evaluated by the mean square (MS) and standard deviation (SD). The results are shown in Table 3. Among them, the M3 model (J1-J3 is fresh, J4-J8 is spoiled) performed best with the highest F statistic (17.36). pThe value was the smallest (2.39E-12) and the SD was relatively low (2.35). This indicates that the M3 model most statistically significantly and stably distinguished between the fresh and spoiled groups. Therefore, J1-J3 were defined as fresh apple juice, and J4-J8 as spoiled apple juice. This dataset was used for subsequent apple juice freshness identification model development and marker screening.

[0061] Table 3 CV-ANOVA test results

[0062] 3.2 Construction of freshly squeezed apple juice identification model and sample identification 3.2.1 Marker screening and regularity study: Based on high-resolution mass spectrometry data (i.e., the chemical component fingerprints in the range of m / z 70-1050 Da of fresh and spoiled apple juice samples defined in 3.1), orthogonal partial least squares discriminant analysis (OPLS-DA) was used to extract differential variables between fresh and spoiled apple juice samples. Markers with significant differences were screened out by combining S-plot analysis and VIP score. Score plot ( Figure 2a ) shows the spatial distribution of samples in the latent variable space, with fresh samples and spoiled samples forming different clusters. This indicates that the OPLS-DA model is robust in identifying variance associated with freshness. 200 permutation tests ( Figure 2b ) shows that the R² and Q² values ​​of the actual data are significantly higher than those of the permuted data set, confirming the reliability and predictive ability of the model. The results of S-plot analysis ( Figure 3 ) shows the difference variables between different apple juice sample groups, where the red-marked points are compounds with |𝑝[1]|>0.05 and |𝑝(corr)|>0.1, which have a greater contribution to the distinction between different groups and are selected as candidate markers. The difference variables were further evaluated through VIP (Variable Importance Projection) analysis. The top-ranked variables showed strong stability and importance, that is, the top-ranked variables have a higher importance for the model to distinguish between fresh and spoiled apple juice samples, reflecting their contribution to the difference between groups. The results of VIP analysis are shown in Figure 4 The results show that the top 30 variables ranked by VIP score were selected as candidate markers. Finally, 15 compounds that met both the S-plot and VIP score requirements were identified as important markers (Table 4). T-tests were performed on the marker contents, and the results showed that the contents of the 15 markers were significantly different between the fresh and spoiled groups ( p <0.05), meeting the statistical and biological relevance criteria.

[0063] Table 4 Apple juice freshness markers

[0064] 3.2.2 Correlation between changes in marker content and apple juice freshness: The distribution of marker content among groups and the trend of change with apple storage time ( Figures 5a-5o ), this example further reveals the dynamic changes in the chemical composition of apple juice at different spoilage stages.

[0065] The contents of 4-guanidinobutyraldehyde (129) and cytosine (111) gradually increased from fresh to spoiled samples, reflecting their association with protein degradation and nitrogen-related metabolic byproducts. In contrast, markers such as caffeic acid 3-glucoside (342) showed a downward trend, indicating that they were consumed or transformed during storage. Several markers showed consistent changes. For example, coniferyl alcohol (180) and acetyl-L-carnitine (203) reached peaks in the intermediate stage (J4-J6), indicating that they were involved in transitional metabolic pathways, such as oxidative stress response. Markers such as choline sulfate (183) and panthenol (205) maintained relatively stable levels in fresh samples (J1-J3), but increased significantly in spoiled samples (J6-J8), indicating that they began to accumulate at the end of spoilage. The dynamic change trend of the content of these markers provides an important chemical basis for the scientific identification of apple juice freshness, and also lays the foundation for the subsequent establishment of a more accurate identification model. These markers can assist in judging the freshness of freshly squeezed apple juice.

[0066] 3.2.3 Construction of freshly squeezed apple juice identification model: Based on the definition of fresh and rotten apples in 3.1, six algorithms were used: random forest (RF), k-nearest neighbor (KNN), support vector machine (SVM), linear discriminant analysis (LDA), orthogonal partial least squares-discriminant analysis (OPLS-DA), and linear support vector machine (Linear SVM) to construct fresh and rotten freshly squeezed apple juice identification models. The classification performance of various models in distinguishing fresh and rotten apple juice samples was evaluated through five-fold cross-validation. The results are shown in the figure below. Figure 6a to Figure 6f and shown in Table 5.

[0067] The RF model showed moderate separation between fresh and spoiled samples ( Figure 6a ), with some overlap near the decision boundary, the KNN model achieved 84.6% accuracy for the training samples and 82.6% accuracy for the test samples. While most fresh and spoiled samples were correctly classified, a few fresh samples were misclassified as spoiled, reflecting a tendency for the model to slightly overfit the training data. Compared to the RF model, the KNN model showed a clearer separation between the different sample types ( Figure 6b). KNN correctly identified 89.7% of the training samples, but the accuracy for the test samples dropped slightly to 82.6%. KNN correctly classified most of the fresh samples, but misclassified two fresh samples, indicating a slight decrease in generalization performance. The SVM showed overlapping prediction probabilities near the decision boundary ( Figure 6c ), indicating weak separation. The model correctly identified 82.05% of the training samples and 86.9% of the test samples, with reduced accuracy, especially for fresh samples. This suggests that the model may not be able to effectively capture the complex relationships in the dataset compared to other methods. LDA reasonably distinguished between fresh and corrupted samples, although there was some overlap near the decision boundary ( Figure 6d The accuracy of identification of modeled data is 89.7%, and the accuracy of identification of test samples is 82.6%, showing good performance and accuracy close to KNN. The OPLS-DA model shows strong separation ability ( Figure 6e ), most samples were correctly classified. The accuracy of the training samples was 94.8%, and the accuracy of the test samples was 82.6%, showing robust performance and becoming one of the best performing models. The linear SVM also had very good prediction accuracy, with almost no overlap between fresh and spoiled samples ( Figure 6f The model achieved 100% classification accuracy for training samples and 91.3% for test samples, making it the most accurate model in this study. The clear distinction between different classes highlights the model's robustness and versatility.

[0068] Table 5 Evaluation of the six models on the identification performance of fresh and rotten apple samples

[0069] In summary, linear SVM performed the best in this study, with OPLS-DA also showing excellent results. RF, KNN, and LDA also performed well, though somewhat less so in terms of accuracy and generalizability. SVM with a radial basis function kernel exhibited the weakest separation. The superior performance of linear SVM over nonlinear SVM and other machine learning models (such as RF and KNN) can be attributed to the strong linear separability of the data. By optimizing a linear decision boundary, linear SVM effectively strikes a balance between model simplicity and classification accuracy, and successfully avoids the risk of overfitting associated with nonlinear models when using relatively small sample sizes. Unlike complex machine learning models that require larger datasets for effective generalization, linear SVM combines the simplicity of linear models with the powerful optimization capabilities of machine learning algorithms, enabling superior generalization and classification performance. Therefore, linear SVM outperformed all other models in this study, demonstrating its suitability for distinguishing fresh from spoiled apple juice samples.

[0070] 3.2.3 Classification Threshold Optimization and Secondary Verification Strategy for the Identification Model: To determine the optimal classification probability threshold, a receiver operating characteristic (ROC) curve analysis was performed on the prediction results of the linear SVM model in 3.2.2. By calculating the true positive rate (TPR) and false positive rate (FPR) at different thresholds, the threshold combination that maximizes the Youden Index was determined to be 0.8 (fresh group) and 0.2 (spoiled group). At this threshold, high sensitivity and specificity can be achieved, while the proportion of samples in the fuzzy region (0.2-0.8) is controlled below 10%. For samples with probability values ​​between 0.2 and 0.8, repeated testing (repeatedly collecting fingerprints for the same unknown juice sample and recalculating the classification probability by averaging the multiple test results) or secondary verification is performed. When performing secondary verification, the secondary verification process is as follows: Variable screening: Extract high-contribution characteristic variables with a VIP score ≥ 1.5 to ensure that the screened variables have a significant impact on the differences between groups; Correlation verification: Calculate the Pearson correlation coefficient matrix between the screening variables and the markers in 3.2.1 (Table 4), with the threshold set at |r| ≥ 0.6 (p < 0.01); Judgment rules: When more than 50% of the high-contribution variables are strongly positively correlated with spoilage markers (such as 4-guanidinobutyraldehyde and choline sulfate) (r≥0.6), or strongly negatively correlated with freshness markers (such as caffeic acid 3-glucoside) (r≤-0.6), it is judged to be the spoilage group; conversely, if it is strongly positively correlated with freshness markers (r≥0.6) or strongly negatively correlated with spoilage markers (r≤-0.6), it is judged to be the fresh group.

[0071] It has been verified that this strategy further reduces the misclassification rate of ambiguous samples.

[0072] 4. Conclusion This study, using high-resolution mass spectrometry combined with multivariate statistical modeling, systematically reveals the chemical composition changes of freshly squeezed apple juice during storage. Key markers closely related to apple juice freshness were screened and identified, providing a scientific basis and data support for apple juice freshness identification and quality evaluation. This provides important reference and technical support for the standardized development and regulation of the industry.

[0073] Example 2 Since apple juice samples contain a rich variety of compounds with different solubilities, it is necessary to develop a pretreatment method that can simultaneously extract fat-soluble substances and water-soluble substances. The more commonly used methods are pure water dilution method and organic solvent extraction method, but the ratio of organic phase to aqueous phase is not clear. Therefore, this example optimizes the ratio of organic phase to aqueous phase, examines the number of extracted compounds, and ultimately determines the optimal extraction conditions. Four pretreatment methods were selected for optimization and comparison, as follows: Pretreatment Method 1: Take 1 mL of apple juice supernatant and add 2 mL of the extraction solvent (methanol:acetonitrile:water = 1:1:2 (v / v / v)). Vortex for 1 minute, sonicate for 30 minutes, and centrifuge at high speed (12,000 rpm, 4°C, 10 minutes). The supernatant was filtered through a 0.22 µm PTFE membrane and stored at -20°C until analysis. Pretreatment Method 2: Take 5 mL of apple juice supernatant and add 5 mL of water and 40 mL of 90% ethanol. Vortex for 1 minute, sonicate for 50 minutes, and let stand for 5.5 hours. The supernatant was dried over nitrogen at 40°C, reconstituted with the original mobile phase, and filtered through a 0.22 µm PTFE membrane and stored at -20°C until analysis. Pretreatment Method 3: Take 1 mL of apple juice supernatant and add 20 mL of ultrapure water. Vortex for 1 minute, sonicate for 30 minutes, and centrifuge at high speed (12,000 rpm, 4°C, 10 minutes). The supernatant was filtered through a 0.22 µm PTFE membrane and stored at -20°C pending analysis. Pretreatment Method 4: 1 mL of apple juice supernatant was added to 1 mL of methanol. The mixture was vortexed for 1 minute, sonicated for 30 minutes, and then centrifuged at high speed (12,000 rpm, 4°C, 10 minutes) after 30 minutes. The supernatant was filtered through a 0.22 µm PTFE membrane and stored at -20°C pending analysis.

[0074] Method 1 detected a total of 5661 compounds, method 2 detected a total of 3514 compounds, method 3 detected a total of 1437 compounds, and method 4 detected a total of 3012 compounds. Venn diagram ( Figure 7) illustrates the quantitative differences among the compounds obtained by the four methods. Further identification of the compounds obtained by each method revealed that Method 2 was significantly more effective in removing sugars, with the number and variety of carbohydrates significantly lower than those of the other methods. There were significant differences in peptides and lipids between Methods 4 and 3, likely due to the inability of pure water to extract lipid compounds. Method 1 yielded the greatest number and variety of compounds.

[0075] In summary, method 1 was determined to be the optimal condition for extracting apple juice samples in this study. This system has a high extraction efficiency and provides a scientific and reliable technical guarantee for the freshness identification of apple juice.

[0076] Example 3 Composition and Operation Procedure of a Kit for Identifying the Freshness of Apple Juice 1. Kit Components List: (a) Data analysis module, including data analysis USB flash drive: An identification model constructed using a supervised machine learning algorithm (linear SVM identification model): built-in training files (based on the data from the fresh groups J1-J3 and the spoiled groups J4-J8 in Example 1): an R script for automatically performing data preprocessing including data correction, binning, and normalization; a Python script for automatically calling the identification model, performing unknown sample identification, and outputting results; optionally, the USB flash drive also includes a table of chromatography-mass spectrometry instrument parameters for detecting apple juice.

[0077] (b) Pretreatment consumables module, including the following: 0.22 µm PTFE filter membranes (sterile, individually packaged, suitable for 1.5 mL centrifuge tubes); 10 mL high-speed centrifuge tubes (resistant to 12,000 rpm, with volume scale); (c) Solvent module, including solvents: acetonitrile, methanol, and water mixture (acetonitrile, methanol, and water in a 1:1:2 volume ratio); (d) Internal standard module, including 2-chloro-L-phenylalanine standard solution (1000 μg / mL, ≥98% purity). Dilute the solution in (c) to 1 μg / mL and use it as the extraction solvent. Prepare it immediately. (e) Operational guide module, including: Standardization process manual (including gradient elution program and mass spectrometry parameters); and troubleshooting appendix (e.g., solutions for filter blockage and model errors).

[0078] 2. Operational procedures for testing the freshness of unknown apple juice Step 1: Sample preparation: Kit components: mixed solvent, PTFE filter membrane, centrifuge tube, internal standard Operation steps: Take unknown apple juice sample: Use a pipette to accurately measure 1 mL of unknown apple juice sample, add 2 mL of extraction solvent containing internal standard (prepared by diluting the internal standard to 1 μg / mL with the mixed solvent in the kit), vortex mix for 1 minute, ultrasonicate for 30 minutes (room temperature, 40 kHz), centrifuge at 12,000 rpm and 4°C for 10 minutes, take the supernatant and filter it through a 0.22 μm PTFE membrane, and transfer the filtrate to a sample vial.

[0079] Step 2: Chromatography-Mass Spectrometry Analysis: Kit components: mobile phase parameters, column type Instrument setup: Chromatographic conditions: Column: Waters ACQUITY UPLC BEH C18 (2.1 × 100 mm, 1.7 μm). Mobile phase A: methanol; mobile phase B: 0.1% formic acid in water.

[0080] Table 6 Gradient program:

[0081] Mass spectrometry conditions: Scan mode: Full scan + ddMS 2 (positive and negative ion switching), resolution: primary 70000, secondary 15000, scan range: m / z 70-1050 Da, capillary temperature: 350℃, spray voltage: +3.2 kV or -3.0 kV.

[0082] Operation steps: Inject 2 μL of filtered sample, start the chromatography-mass spectrometry system to collect data, and export the data.

[0083] Step 3: Data preprocessing and freshness identification Kit components: Data analysis USB drive Procedure: Data preprocessing: Use an R script to load the data collected by the chromatography-mass spectrometry system in Step 2 and perform data preprocessing including data correction, binning, and normalization. Freshness determination: Run a Python script on the preprocessed data, load the pretrained linear SVM model, and output the freshness probability of the unknown sample (e.g., "probability of freshness ≥ 90%" or "probability of spoilage ≥ 85%"). The following logic is embedded in the Python script: When the identification model outputs a probability value between 0.2 and 0.8, the marker verification process is automatically triggered; the marker data pre-stored on the USB flash drive (markers are listed in Table 4) are called, and targeted normalized peak area integration and correlation calculations (e.g., Pearson correlation coefficient) are performed. Based on the preset threshold, the initial determination result is overwritten.

[0084] Step 4: Result Generation and Quality Control Generates: (1) freshness identification results of unknown samples; (2) linear SVM classification diagram: displays the decision boundary and classification results.

[0085] Quality control requirements: Coefficient of variation (CV) of the internal standard signal peak area ≤ 30%, internal standard parent ion mass accuracy ≤ 5 ppm. The kit procedure for this invention is as follows: sample extraction (solvent: methanol:acetonitrile:water = 1:1:2) → centrifugal filtration → UPLC-HRMS analysis → peak area and mass axis calibration (internal standard) → data binning (0.2 Da) → normalization (Pareto analysis) → unsupervised clustering (K-means / PCA) → supervised discrimination (linear SVM) → marker-assisted validation → report generation.

[0086] The kit of the present invention efficiently extracts small molecule metabolites from apple juice through a standardized process (capable of detecting more than 5,000 compounds), and integrates the entire process from sample pretreatment, chromatography-mass spectrometry analysis to machine learning identification, making it suitable for rapid and accurate detection of apple juice freshness.

[0087] The technical features of the embodiments of this specification may be combined arbitrarily, and as long as there is no contradiction in the combination of these technical features, they all fall within the scope of this specification. The embodiments are illustrative and do not limit the scope of protection of the present invention. Those skilled in the art may make several modifications or improvements without departing from the concept of the present invention, which all fall within the scope of protection of the present invention, and the specific scope of protection shall be subject to the claims.

Claims

1. A method for identifying the freshness of apple juice, characterized in that: include: S1. Selecting apples of known different freshness, including fresh apples and rotten apples, and preparing apple juices of different freshness from the apples of different freshness, and analyzing and detecting the apple juices using a chromatography-mass spectrometry system to obtain corresponding fingerprints; S2. Based on the peak information of the fingerprints corresponding to the apple juices of different freshness described in S1 as feature variables, the fingerprints are clustered and defined using K-means cluster analysis and orthogonal partial least squares discriminant analysis, and the fingerprints are defined as fresh group or spoiled group to obtain a data set; S3. Based on the data set described in S2, a supervised machine learning algorithm is used to build an identification model; S4. Take unknown apple juice to be tested, use the same detection conditions as S1 to obtain the fingerprint of the unknown apple juice, and identify the freshness of the unknown apple juice based on the identification model described in S3.

2. The method according to claim 1, characterized in that The step S2 also includes: based on the fingerprint map described in S1, using orthogonal partial least squares discriminant analysis to extract difference variables from the apple juice data of different freshness described in S1, combining S-plot analysis and VIP score to screen out markers with significant differences, and based on the markers, assisting in verifying the results of the identification model.

3. The method according to claim 1, characterized in that In step S1, the apple juices of different freshness are subjected to solvent extraction to obtain extracts, and the extracts are analyzed and detected using a chromatography-mass spectrometry system. The solvent extraction includes ultrasound-assisted extraction, and the solvent is a mixed solution of one or more of methanol, ethanol, propanol, ethyl acetate, acetonitrile, methanol or water.

4. The method according to claim 3, characterized in that The solvent consists of a mixed solution of acetonitrile, methanol and water, and the volume ratio of the acetonitrile, methanol and water is (0.5-2):(0.5-2):

2.

5. The method according to claim 1, wherein The step S2 of clustering and defining the fingerprint spectrum using K-means cluster analysis and orthogonal partial least squares discriminant analysis includes: S2-1. Perform unsupervised grouping of the fingerprint data by K-means cluster analysis, calculate the silhouette coefficient and Davis-Boulding index when the number of clusters is 2 to 8, comprehensively refer to the silhouette coefficient and Davis-Boulding index, give priority to the group with a silhouette coefficient higher than 0.1, and select the number of clusters corresponding to the group with the lowest Davis-Boulding index among the groups with a silhouette coefficient higher than 0.1 as the optimal number of clusters; S2-2, based on the K-means clustering results, orthogonal partial least squares discriminant analysis was used for analysis, and the significance of the model was verified by cross-validation variance analysis, and the model with the largest F statistic was selected. p The grouping scheme with the smallest value defines apple juices of different freshness as fresh group or spoiled group.

6. The method according to claim 1, characterized in that The S4 step includes: S4-1, inputting the fingerprint of the unknown apple juice to be tested into the identification model, and calculating its classification probability using the linear decision boundary built into the identification model; S4-2. Determine whether the data in S4-1 is a fresh group or a rotten group based on a preset probability threshold. If the probability threshold is ≥0.8, it is determined to be a fresh group, and if the probability threshold is ≤0.2, it is determined to be a rotten group. If the probability value is between 0.2-0.8, perform secondary verification by repeating the test or introducing auxiliary indicators.

7. The method according to claim 1, characterized in that The chromatographic detection conditions adopted by the chromatography-mass spectrometry system are: Chromatographic column: WATERS ACQUITY UPLC BEH C18 column, specifications: 2.1 × 100 mm, 1.7 μm; Mobile phase A: methanol; Mobile phase B: 0.1% formic acid-water solution; Flow rate: 0.2 mL / min; Injection volume: 2 μL; Column temperature: 40 °C, The chromatography adopts gradient elution, and the gradient elution conditions include: starting, mobile phase A: 2%; 0-10 minutes, mobile phase A: 2-50%; 10-14 minutes, mobile phase A: 50-98%; 14-17 minutes, mobile phase A: 98%; 17-17.1 minutes, mobile phase A: 98-2%; 17.1-20 minutes, mobile phase A: 2%; And / or, the mass spectrometry detection conditions of the chromatography-mass spectrometry system are: Scan mode: full scan + secondary fragmentation, positive and negative ion switching; Resolution: Level 1 70000, Level 2 15000; Capillary temperature: 350 °C; Auxiliary gas temperature: 320 ℃; Sheath gas flow: 40 arb; Auxiliary gas flow rate: 10 arb; S-lens RF voltage: 55 V; Spray voltage: +3.2 kV or -3.0 kV; Scan range: m / z 70-1050 Da.

8. The method according to claim 1, characterized in that The supervised machine learning algorithm is selected from any one of the following algorithms: random forest, k-nearest neighbor, support vector machine, linear discriminant analysis, orthogonal partial least squares-discriminant analysis or linear support vector machine.

9. The method according to claim 2, characterized in that The markers include: one or more of 4-guanidinobutyraldehyde, cytosine, choline sulfate, maleate, coniferyl alcohol, sulfoacetone, acetyl-L-carnitine, panthenol, 4-(2-aminophenyl)-2,4-dioxobutyrate, L-hypoglycine, 4-methylene-L-glutamate, adenosine, 2,3-dihydro-3-hydroxy-2-oxo-1H-indole-3-acetic acid methyl ester, 3-(3,4-dihydroxyphenyl) lactic acid, and caffeic acid 3-glucoside.

10. A kit for identifying the freshness of apple juice, characterized in that: It includes (a) a data analysis module, wherein the (a) data analysis module includes a data analysis USB flash drive, and the data analysis USB flash drive includes the following (a-1) to (a-3): (a-1) an identification model constructed based on the supervised machine learning algorithm according to any one of claims 1 to 9; (a-2) Data preprocessing R scripts that automatically perform data correction, binning, and normalization; (a-3) Python that automatically calls the identification model, performs unknown sample identification, and outputs the results; or, the kit includes the (a) data analysis module and the following (b) to (e) modules: (b) a pre-treatment consumables module, comprising a filter membrane and a high-speed centrifuge tube; (c) a solvent module, comprising a solvent, the solvent being consistent with the extraction solvent in S1 according to any one of claims 1 to 9, for obtaining an extract of an unknown apple juice sample, and performing analysis and detection based on the extract using a chromatography-mass spectrometry system, wherein the solvent is a mixed solution of acetonitrile, methanol, and water in a volume ratio of 0.5-2:0.5-2:2; (d) an internal standard module, wherein the internal standard module comprises 2-chloro-L-phenylalanine; (e) Operation guide module, including standardized testing procedures and troubleshooting solutions.

Citation Information

Patent Citations

  • A method for judging the flavor quality of concentrated apple juice and its application

    CN105717227B

Cited By

  • Method for evaluating freshness of raw milk

    CN121068819A