Method for predicting and identifying whether non-concentrated reduced fruit juice contains concentrated reduced fruit juice

The model established using Raman spectroscopy and tabular probabilistic feedforward network algorithm solves the problem of identifying non-concentrated fruit juice, achieving rapid and accurate detection of fruit juice adulteration, filling a market gap, and is applicable to the identification and quantitative analysis of various fruit juices.

CN121601078APending Publication Date: 2026-03-03SGS CSTC STANDARDS TECH SERVICES (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511785130.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and accurately identify whether non-concentrated fruit juice is adulterated with concentrated fruit juice, and the testing costs are high and the cycle is long, resulting in serious counterfeiting in the market, which affects consumer rights and the market ecosystem.

Method used

By employing Raman spectroscopy combined with the TabProbabilisticFeedforwardNetwork (TabPFN) algorithm, classification and regression models are established through spectral scanning of different fruit juice samples, allowing for direct prediction and identification of the fruit juice category and adulteration ratio from Raman spectral data.

Benefits of technology

It enables rapid and accurate identification of whether non-concentrated fruit juice is adulterated with concentrated fruit juice, and can quantify the adulteration ratio. The process is simple, low-cost, applicable to a variety of common fruit juices, and provides accurate results, meeting market demands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121601078A_ABST
    Figure CN121601078A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of analytical chemistry and mathematical modeling, and particularly relates to a method for predicting and identifying whether non-concentrated and reduced fruit juice contains concentrated and reduced fruit juice, which comprises the following steps: taking fruit juice with different types and different doped concentrated and reduced fruit juice proportions as modeling samples; performing spectrum scanning on the modeling sample by adopting a Raman spectrum technology to obtain original Raman spectrum information; creating a data set and dividing the data set; training an FC fruit juice / NFC fruit juice classification model by adopting a table probability feed-forward network TabPFN; storing parameters of the FC / NFC classification model; training an FC / NFC doping proportion regression model; predicting the FC / NFC category of a new sample; predicting the FC / NFC doping proportion of the doped sample; compared with the prior art, the method can quickly and accurately identify whether the non-concentrated and non-reduced fruit juice is adulterated with the concentrated and non-reduced fruit juice, and is simple and quick in identification process, short in analysis time, low in test cost and high in practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of analytical chemistry and mathematical modeling. Specifically, it is a method for predicting and identifying whether non-concentrated fruit juice contains concentrated fruit juice based on Raman spectroscopy and tabular probabilistic feedforward network algorithm. Background Technology

[0002] With the continuous improvement of people's health awareness, the market for 100% additive-free fruit juice has been growing year by year. It is estimated that by 2030, the market size of 100% fruit juice in China will reach approximately US$4.7375 billion, with a compound annual growth rate of 8.3% from 2025 to 2030. Among them, Not From Concentrate (NFC) juice, as a high-end category, is experiencing rapid growth. NFC technology is a juice processing technique whose core lies in not concentrating or reducing the juice. Instead, fresh fruit is washed and juiced, and then directly bottled after pasteurization or non-thermal sterilization (such as ultra-high pressure processing), thus preserving the natural flavor, color, and nutrients of the fruit to the greatest extent. NFC juice has successfully stood out in the juice market and become the daily healthy beverage choice for more and more consumers due to its higher nutritional value, taste closer to fresh fruit, quality assurance brought by advanced technology, and precise alignment with the trend of healthy consumption.

[0003] NFC (Not From Concentrate) juice, due to its complex processing and high cold chain costs, is typically priced significantly higher than FC (From Concentrate) juice. Unscrupulous merchants, seeking maximum profit, engage in rampant adulteration. Passing off FC juice as NFC juice is another common form of adulteration. While "NFC" is not a mandatory national standard term, some companies play word games on labels, such as adding a "+" after "NFC" or stating "added NFC concentrate," when the main ingredient is still NFC juice. Adulteration of NFC juice not only directly infringes on consumer rights but also has a profound negative impact on the entire market ecosystem by disrupting prices, destroying trust, attracting stricter regulation, damaging brands, and distorting competition. Although DNA or component analysis-based detection methods exist, these tests are not yet widely adopted in market supervision and corporate self-inspection due to their long testing cycles, high costs, and complex technologies. The low probability of adulteration being detected in random checks encourages unscrupulous merchants to take chances.

[0004] Raman spectroscopy is a non-destructive analytical technique that analyzes the chemical composition and molecular structure of a sample by detecting the scattered spectra produced after the sample interacts with incident light. When a laser irradiates a sample, the frequency of the vast majority of the scattered light remains unchanged (Rayleigh scattering), with only a very small number of photons experiencing a frequency shift due to energy exchange with molecular vibrations or rotations (Raman scattering). This frequency shift characteristic is directly related to molecular vibrational modes and can therefore serve as a basis for substance identification. However, in previous practical applications of food testing, the complexity of sample composition and the overlap of Raman signals from different compounds have made spectral analysis difficult, limiting the analytical efficiency of this technique in complex matrices.

[0005] In recent years, with the development of machine learning methods, Raman spectroscopy has made significant progress in the application of food detection, and its technological value has been increasingly recognized, demonstrating its potential for rapid analysis in complex food systems. Against the backdrop of technological convergence, data-driven algorithms, represented by deep learning, are gradually permeating various disciplines. However, conventional deep learning models have high requirements for data scale and computational resources, while food analysis typically faces challenges such as limited sample size and high real-time requirements, making it difficult to directly apply large-scale deep learning networks. Tabular Probabilistic Feedforward Network (TabPFN), as a deep learning framework specifically designed for tabular data, combines efficiency and small-sample adaptability, integrating the advantages of probabilistic modeling and feedforward neural networks. Raman spectroscopy data is essentially tabular data, therefore TabPFN is naturally suited for analyzing this type of data. It is important to note that Raman spectra typically have high-dimensional features (the number of wavenumber points far exceeds the number of samples), so dimensionality reduction or feature selection processing is necessary before using TabPFN. Summary of the Invention

[0006] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a method for predicting and identifying whether non-concentrated fruit juice contains concentrated fruit juice. This test method can quickly and accurately identify whether non-concentrated fruit juice is adulterated with concentrated fruit juice, and the identification process is simple and fast, the analysis time is short, the test cost is low, and the practicality is high.

[0007] To achieve the above objectives, a method for predicting and identifying whether non-concentrated fruit juice contains concentrated fruit juice is designed, comprising the following steps: 1) Use different types of fruit juice with different proportions of concentrated and reduced fruit juice as modeling samples; 2) The modeled sample obtained in step 1) is subjected to spectral scanning using Raman spectroscopy to obtain the original Raman spectral information; 3) Create and partition the dataset; 4) The TabPFN algorithm was used to train a classification model for concentrated FC juice / non-concentrated NFC juice; 5) Save the FC / NFC classification model parameters; 6) Train the FC / NFC doping ratio regression model; 7) Predict FC / NFC category for new samples; 8) For non-pure NFC juice and non-pure FC juice samples, predict the FC / NFC doping ratio.

[0008] Further, in step 1), the different types of juice include, but are not limited to, any one of orange juice, apple juice, grape juice, peach juice, mango juice, pineapple juice, watermelon juice, strawberry juice, lemon juice, tomato juice, blueberry juice, and carrot juice. Each type of juice has at least one set of NFC juice and / or one set of FC juice, where one set refers to two different brands. Randomly select one NFC and one FC juice sample and mix it with FC juice at intervals of 5% from 95% to 5% of the NFC juice to obtain juices with different FC / NFC mixing ratios.

[0009] Furthermore, in step 2), the instrument used is an Agilent Resolve handheld Raman spectrometer with an incident depth set to 1.5.

[0010] Further, in step 3), the Raman spectral data of each juice obtained in step 2) are randomly sampled from the NFC and FC datasets, with 80% of the samples forming the training set and 20% forming the prediction set, for training the FC / NFC classification model; in the dataset of each juice with different FC / NFC doping ratio, random selections are used for training the regression model of FC / NFC doping.

[0011] Further, in step 4), the Raman spectral data of the randomly selected NFC and FC juices in step 3) are used to train a classification model using the TabPFN pre-trained model. The TabPFN pre-trained model does not require hyperparameter adjustment during training. After training, the trained model is used to predict the FC / NFC ratio of the prediction set samples, and the accuracy is calculated based on the predicted FC / NFC ratio and the actual FC / NFC ratio.

[0012] Furthermore, the classification accuracy obtained in step 4) should be greater than 95%, and the area under the receiver operating characteristic (ROC) curve should be greater than 0.95. For classification results that are not ideal, try to correct the model again by increasing the amount of data or re-dividing the data.

[0013] Furthermore, in step 5), the qualified model parameters trained in step 4) are saved as initialization parameters for the subsequent FC / NFC doping ratio regression model.

[0014] Further, in step 6), the juice samples with different NFC / FC doping ratios randomly selected in step 3) are used to train a regression model in combination with the TabPFN pre-trained model in step 5); a portion of the mixed samples are randomly selected as prediction samples to predict their doping ratio; the root mean square error (RMSE) and correlation coefficient (R) are used to evaluate the error between the prediction result and the actual doping ratio; the RMSE should not be greater than 10% of the mean doping ratio, and the R should not be less than 0.95.

[0015] Further, in step 7), for a new unknown sample, after acquiring Raman spectral data using step 3), the TabPFN model trained in step 5 is used to predict its FC / NFC ratio to determine whether it is pure NFC or pure FC juice.

[0016] Further, in step 8), for the result predicted in step 7), if the unknown sample is neither pure NFC juice nor pure FC juice, the TabPFN regression model trained in step 6) is used to predict its doping ratio.

[0017] Compared with existing technologies, this invention provides a method for predicting and identifying whether non-concentrated fruit juice contains concentrated fruit juice by directly acquiring Raman spectral characteristic data without sample processing. A model established using the TabPF algorithm predicts and identifies the presence of concentrated fruit juice in non-concentrated fruit juice. This method is applicable to common fruit juices such as orange juice, apple juice, grape juice, peach juice, mango juice, pineapple juice, watermelon juice, strawberry juice, lemon juice, tomato juice, blueberry juice, and carrot juice. The identification process is simple, fast, and accurate, filling a technological gap in the market and meeting market testing needs. In summary, this invention can quickly and accurately identify whether non-concentrated fruit juice is adulterated with concentrated fruit juice. Furthermore, if non-concentrated fruit juice is adulterated with concentrated fruit juice, the adulteration ratio can be accurately quantified. The identification results are consistent with the actual situation. The identification process is simple and fast, with short analysis time, low testing cost, high accuracy, and high practicality, making it worthy of application and promotion. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the structure of the present invention; Figure 2 This is the spatial distribution map of principal component analysis (PCA) of Embodiment 1 of the present invention; Figure 3 This is the regression linear graph of Embodiment 2 of the present invention. Detailed Implementation

[0019] This invention belongs to the fields of analytical chemistry and mathematical modeling. Specifically, it is a method for predicting and identifying whether non-concentrated fruit juice contains concentrated fruit juice based on Raman spectroscopy and tabular probabilistic feedforward network algorithm. The method acquires spectra based on Raman spectroscopy, and uses tabular probabilistic feedforward network algorithm to further model and analyze the data. First, it distinguishes between concentrated and non-concentrated juice, and then uses tabular probabilistic feedforward network algorithm to predict the doping ratio of impure fruit juice.

[0020] The basic principle of this invention is as follows: Raman spectral characteristic data are directly collected from fruit juice samples, and a tabular probabilistic feedforward network algorithm is used to predict and identify whether non-concentrated fruit juice contains concentrated fruit juice, as shown in the appendix. Figure 1 As shown, the specific steps are as follows: 1) Use different types of fruit juice with different proportions of concentrated and reduced fruit juice as modeling samples; 2) The modeled sample obtained in step 1) is subjected to spectral scanning using Raman spectroscopy to obtain the original Raman spectral information; 3) Create and partition the dataset; 4) The TabPFN algorithm was used to train classification models for different proportions of concentrated FC juice and non-concentrated NFC juice. 5) Save the FC / NFC classification model parameters; 6) Train the FC / NFC doping ratio regression model; 7) Predict FC / NFC category for new samples; 8) For non-pure NFC juice and non-pure FC juice samples, predict the FC / NFC doping ratio.

[0021] In step 1), different types of juice include, but are not limited to, any one of orange juice, apple juice, grape juice, peach juice, mango juice, pineapple juice, watermelon juice, strawberry juice, lemon juice, tomato juice, blueberry juice, and carrot juice. Each type of juice has at least one set of NFC juice and one set of FC juice, where one set refers to two different brands. Randomly select one NFC and one FC juice sample and mix it with FC juice at intervals of 5% from 95% to 5% of the NFC juice to obtain juices with different FC / NFC mixing ratios.

[0022] In step 2), the instrument used is an Agilent Resolve handheld Raman spectrometer with an incident depth set to 1.5.

[0023] In step 3), the Raman spectral data of each juice obtained in step 2) are randomly sampled from the NFC and FC datasets, with 80% of the samples forming the training set and 20% forming the prediction set, which are used to train the FC / NFC classification model; in the dataset of each juice with different FC / NFC doping ratio, random selections are used to train the regression model of FC / NFC doping.

[0024] In step 4), the Raman spectral data of the randomly selected NFC and FC juices from step 3) are used to train a classification model using the TabPFN pre-trained model. The TabPFN pre-trained model does not require hyperparameter adjustment during training. After training, the trained model is used to predict the FC / NFC ratio of the prediction set samples, and the accuracy is calculated based on the predicted and actual FC / NFC ratios. Typically, a classification accuracy greater than 95% and an area under the receiver operating characteristic (ROC) curve greater than 0.95 are required. For classification results that are not ideal, the model can be revised by increasing the dataset size or re-dividing the dataset.

[0025] In step 5), the model parameters that meet the requirements after training in step 4) are saved as initialization parameters for the subsequent FC / NFC doping ratio regression model.

[0026] In step 6), the juice samples with different NFC / FC doping ratios randomly selected in step 3) are used to train a regression model in combination with the TabPFN pre-trained model in step 5); a portion of the mixed samples are randomly selected as prediction samples to predict their doping ratio; the root mean square error (RMSE) and correlation coefficient (R) are used to evaluate the error between the prediction result and the actual doping ratio; generally speaking, the RMSE should not be greater than 10% of the mean doping ratio, and the R should not be less than 0.95.

[0027] In step 7), for a new unknown sample, after acquiring Raman spectral data in step 3), the TabPFN model trained in step 5 is used to predict its FC / NFC ratio to determine whether it is pure NFC or pure FC juice.

[0028] In step 8), for the result predicted in step 7), if the unknown sample is neither pure NFC juice nor pure FC juice, the TabPFN regression model trained in step 6) is used to predict its doping ratio.

[0029] The present invention will be further described below with reference to the accompanying drawings and specific embodiments: Example

[0030] FC / NFC certification of brand A and brand B juice (1) Randomly purchased juice samples of brand A and brand B from the market will be used as test samples; (2) The instrument used was an Agilent Resolve handheld Raman spectrometer. Raman spectra of two juice samples, A and B, were collected at an incident depth of 1.5. (3) The Raman spectral data of samples A and B are input into the TabPFN prediction model established in this invention to obtain the sample prediction results. A is NFC juice and B is FC juice.

[0031] As attached Figure 2 As shown, samples A and B are distributed in the NFC and FC sample regions respectively in the same principal component analysis (PCA) space, proving that A is NFC juice and B is FC juice. The prediction results are consistent with the actual situation of the samples, proving the effectiveness of the classification method. Example

[0032] FC / NFC Regression Prediction of FC / NFC Mixed Juices (1) Mix the FC and NFC samples from Example 1 to prepare two samples C and D with FC / NFC ratios of 40% and 50%, respectively.

[0033] (2) The instrument used was an Agilent Resolve handheld Raman spectrometer. Raman spectra of two juice samples, C and D, were collected at an incident depth of 1.5.

[0034] (3) The Raman spectral data of samples C and D were input into the regression model in TabPFN to predict the doping ratio. The results showed that the FC juice content in C was 41.4% and the FC juice content in D was 51.7%, as shown in the attached figure. Figure 3 As shown, the prediction results are all within the corresponding doping ratio range, proving the effectiveness of the FC / NFC doping ratio regression model proposed in this invention.

[0035] Contents not described in detail in this specification are existing technologies known to those skilled in the art and will not be elaborated upon here. This invention is not limited to the above-described embodiments; any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of this invention should be considered equivalent substitutions and are included within the scope of protection of this invention.

Claims

1. A method for predicting and identifying whether non-concentrated fruit juice contains concentrated fruit juice, characterized in that, Includes the following steps: 1) Use different types of fruit juice with different proportions of concentrated and reduced fruit juice as modeling samples; 2) The modeled sample obtained in step 1) is subjected to spectral scanning using Raman spectroscopy to obtain the original Raman spectral information; 3) Create and partition the dataset; 4) The TabPFN algorithm was used to train a classification model for concentrated FC juice / non-concentrated NFC juice; 5) Save the FC / NFC classification model parameters; 6) Train the FC / NFC doping ratio regression model; 7) Predict FC / NFC category for new samples; 8) For non-pure NFC juice and non-pure FC juice samples, predict the FC / NFC doping ratio.

2. The method as described in claim 1, characterized in that: In step 1), different types of juice include, but are not limited to, any one of orange juice, apple juice, grape juice, peach juice, mango juice, pineapple juice, watermelon juice, strawberry juice, lemon juice, tomato juice, blueberry juice, and carrot juice. Each type of juice has at least one set of NFC juice and / or one set of FC juice, where one set refers to two different brands. Randomly select one NFC and one FC juice sample and mix it with FC juice at intervals of 5% from 95% to 5% of the NFC juice to obtain juices with different FC / NFC mixing ratios.

3. The method as described in claim 1, characterized in that: In step 2), the instrument used is an Agilent Resolve handheld Raman spectrometer with an incident depth set to 1.

5.

4. The method as described in claim 1, characterized in that: In step 3), the Raman spectral data of each juice obtained in step 2) are randomly sampled from the NFC and FC datasets, with 80% of the samples forming the training set and 20% forming the prediction set, which are used to train the FC / NFC classification model; in the dataset of each juice with different FC / NFC doping ratio, random selections are used to train the regression model of FC / NFC doping.

5. The method as described in claim 1, characterized in that: In step 4), the Raman spectral data of NFC and FC juices randomly selected in step 3) are used to train a classification model using the TabPFN pre-trained model. The TabPFN pre-trained model does not require hyperparameter adjustment during training. After training, the trained model is used to predict the FC / NFC ratio of the prediction set samples, and the accuracy is calculated based on the predicted FC / NFC ratio and the actual FC / NFC ratio.

6. The method as described in claim 5, characterized in that: Step 4) The classification accuracy should be greater than 95%, and the area under the receiver operating characteristic (ROC) curve should be greater than 0.

95. For classification results that are not ideal, try to correct the model again by increasing the amount of data or re-dividing the data.

7. The method as described in claim 1, characterized in that: In step 5), the model parameters that meet the requirements after training in step 4) are saved as initialization parameters for the subsequent FC / NFC doping ratio regression model.

8. The method as described in claim 1, characterized in that: In step 6), the juice samples with different NFC / FC doping ratios randomly selected in step 3) are used to train a regression model in combination with the TabPFN pre-trained model in step 5); a portion of the mixed samples are randomly selected as prediction samples to predict their doping ratio. The error between the predicted result and the actual doping ratio is evaluated using the root mean square error (RMSE) and the correlation coefficient (R); the RMSE should not be greater than 10% of the mean doping ratio, and the R should not be less than 0.

95.

9. The method as described in claim 1, characterized in that: In step 7), for a new unknown sample, after acquiring Raman spectral data in step 3), the TabPFN model trained in step 5 is used to predict its FC / NFC ratio to determine whether it is pure NFC or pure FC juice.

10. The method as described in claim 1, characterized in that: In step 8), for the result predicted in step 7), if the unknown sample is neither pure NFC juice nor pure FC juice, the TabPFN regression model trained in step 6) is used to predict its doping ratio.