Method for detecting protein content in soybean based on low-field nuclear magnetic resonance technology

The relaxation time of hydrogen protons in soybeans is detected by low-field nuclear magnetic resonance technology, and a protein content calculation model is established, which solves the problems of time-consuming, low accuracy and chemical reagent use of existing soybean protein content determination methods, achieving rapid, accurate and lossless measurement of protein content.

CN119936097APending Publication Date: 2025-05-06SHANGHAI NIUMAI ELECTRONICS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411940387.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing soy protein content measurement methods take a long time, have low accuracy, and require the use of chemical reagents, which poses environmental pollution and health risks.

Method used

Low-field NMR technology was used to detect the relaxation time of hydrogen protons in soybeans, and the protein content was calculated by establishing a protein content calculation model. This method requires no chemical reagents, is simple and fast in operation, and the measurement results have high accuracy and good repeatability.

Benefits of technology

It realizes rapid, accurate and non-destructive measurement of soy protein content, avoids the use of chemical reagents, reduces environmental pollution and health risks, and improves the repetition and reproducibility of measurements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119936097A_ABST
    Figure CN119936097A_ABST
Patent Text Reader

Abstract

The invention discloses a method for detecting protein content in soybean based on a low-field nuclear magnetic resonance technology, which comprises the following steps: step 1) obtaining a soybean model sample series which comprises a plurality of soybean samples with known protein content and different protein content; 2) detecting the model sample series by adopting low-field nuclear magnetic resonance, and recording peak point signal data of each sample in the model sample series; 3) establishing a protein content calculation model according to the protein content of each sample in the model sample series and the peak point signal data of each sample; 4) detecting the soybean sample to be detected by adopting low-field nuclear magnetic resonance; and calculating the protein content of the soybean sample to be detected according to the measured peak point signal data by adopting the protein content calculation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of protein detection, and in particular to a method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology. Background Art

[0002] In the fields of agricultural technology, biotechnology and analytical chemistry, the measurement of protein content is an important task. Existing solutions mainly include chemical analysis methods and physical analysis methods. Chemical analysis methods such as Kjeldahl nitrogen determination method and Coomassie brilliant blue method, although with high accuracy, are complex to operate, time-consuming, and require the use of chemical reagents, which pose environmental pollution and health risks. Physical analysis methods such as infrared spectroscopy and ultraviolet spectroscopy, although relatively simple to operate, have low measurement accuracy and cannot meet the needs of high-precision measurement. Therefore, it is necessary to develop a simple, fast, non-destructive, high-precision and repeatable protein content measurement method.

[0003] Since the T1 and T2 relaxation times of proteins in samples can reflect the differences between proteins and other substances, the magnetic resonance method can be used to measure protein content, which is simple, convenient, non-destructive, and rapid, has low requirements for test sample processing, and has good repeatability. However, there is currently no patent for using low-field nuclear magnetic resonance technology to measure soybean protein content. Summary of the invention

[0004] This method proposes a method for detecting the protein content in soybeans based on low-field nuclear magnetic resonance technology, which has solved the problems of the existing soybean protein content determination being time-consuming, low in accuracy, and requiring the use of chemical reagents. The above purpose can be achieved through the implementation of the following technical solutions:

[0005] A method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology comprises the following steps:

[0006] Step 1) obtaining a soybean model sample series, wherein the soybean model sample series comprises a plurality of soybean samples with known protein contents and different protein contents;

[0007] Step 2) using low-field nuclear magnetic resonance to detect the model sample series, and recording the peak signal data of each sample in the model sample series;

[0008] Step 3) establishing a protein content calculation model according to the protein content of each sample in the model sample series and the peak signal data of each sample;

[0009] Step 4) using low-field nuclear magnetic resonance to detect the soybean sample to be tested; using the protein content calculation model to calculate the protein content of the soybean sample to be tested based on the measured peak signal data.

[0010] Optionally, the protein content of the soybean samples in the soybean model sample series is 30-40 g / 100 g.

[0011] Optionally, the soybean samples in the soybean model sample series are all measured for protein content by Kjeldahl method.

[0012] Optionally, establishing a protein content calculation model in step 3) comprises the following steps:

[0013] a) selecting peak signals with high correlation with protein content from the peak signal data of each sample;

[0014] b) assigning weights to the screened peak signals;

[0015] c) establishing a protein content calculation model according to the weighted peak point signals and the protein content of each sample in the model sample series.

[0016] Optionally, in step a), a genetic algorithm based on partial least squares is used to screen out peak signals with a high correlation with protein content from the peak signal data of each sample.

[0017] Optionally, the method of using a genetic algorithm based on partial least squares to screen out peak point signals with a high correlation with protein content from the peak point signal data of each sample specifically includes:

[0018] Step 1) Echo data encoding: each peak signal is regarded as a gene and binary encoded. If the encoded gene is 1, the peak signal is selected, and if it is 0, it is not selected;

[0019] Step 2) Initializing the number of echo points: randomly generating n groups of encoded peak point signals, each group having different genes, representing a different combination of peak point signal data;

[0020] Step 3) using the partial least squares regression analysis model to calculate the fitness function of each gene population and sort them from small to large;

[0021] The fitness function F(x)=RMSE+1 / R 2 , where RMSE represents the root mean square error of the model prediction after modeling all peak point signals in each population, R 2 represents the prediction correlation coefficient of the model;

[0022] Step 4) according to the sorting in step 3), exclude the population with a larger total in proportion, and then perform replication, crossover and mutation operations to form a new population. After each operation, record the peak signal to confirm the selected frequency and record the frequency probability;

[0023] Step 5) Iterate according to step 3) and step 4) until the loop ends, and obtain the optimal peak signal for the final regression analysis operation, and finally screen out the peak signal with the highest frequency.

[0024] Preferably, during the screening process, the crossover probability is set to 0.8, the mutation probability is set to 0.01, the population size is set to 50, and the evolutionary generation is set to 100.

[0025] Optionally, in the step b), weights are assigned to the screened peak signals according to their frequencies, and the frequency weight parameter is set to 0.02.

[0026] Optionally, the weight of the sampling points screened in step 6) is greater than 0.2.

[0027] Optionally, the modeling algorithm is partial least squares analysis with non-negative NN restrictions.

[0028] Optionally, the protein content calculation model interactive verification correlation coefficient R 2 >0.9.

[0029] Optionally, the parameters of the low-field nuclear magnetic resonance detection are as follows: number of echoes: 100-500000; number of inversions: 10-50; MagicSE sampling time: 8us-400us, number of accumulations: 2-1024.

[0030] Optionally, the step (iv) further includes: using the protein content calculation model to obtain a predicted value of the protein content of each sample in the soybean model sample series; then using the predicted value of the protein content of each sample in the soybean model sample series and the known protein content to establish a working curve; then using low-field nuclear magnetic resonance to detect the soybean sample to be tested; and using the protein content calculation model and the working curve to calculate the protein content of the soybean sample to be tested.

[0031] The technical solution of the present invention has the following advantages:

[0032] The present invention is based on low-field nuclear magnetic resonance technology, and characterizes sample characteristics and determines protein content by detecting the relaxation time of hydrogen protons in soybeans. This method can solve the problems of complex operation and long time consumption of traditional methods, and can also avoid the use of chemical reagents, fundamentally solving the problems of environmental pollution and health risks. The low-field nuclear magnetic resonance technology of the present invention has significant advantages in measuring soybean protein content, and the measurement results have good repeatability, reproducibility and accuracy, and can be widely used in the food industry and quality control fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0034] Figure 1 Filter result graph for original sampling data and data sampling points;

[0035] Figure 2 It is the model working curve diagram;

[0036] Figure 3 This is the result diagram of model interaction verification. DETAILED DESCRIPTION

[0037] Now, various exemplary embodiments of the present invention are described in detail, and this detailed description should not be considered as a limitation of the present invention, but should be understood as a more detailed description of certain aspects, characteristics and embodiments of the present invention. It should be understood that the terms described in the present invention are only for describing specific embodiments and are not used to limit the present invention.

[0038] In addition, for the numerical range in the present invention, it is understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. The intermediate value in any stated value or stated range, and each smaller range between any other stated value or intermediate value in the range is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded in the scope.

[0039] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the invention pertains. Although only preferred methods and materials are described herein, any methods and materials similar or equivalent to those described herein may also be used in the practice or testing of the present invention.

[0040] The words “include,” “including,” “have,” “contain,” etc. used in this document are open-ended terms, meaning including but not limited to.

[0041] Example 1

[0042] This embodiment provides a method for detecting the protein content in soybeans based on low-field nuclear magnetic resonance technology:

[0043] The test equipment uses a low-field nuclear magnetic resonance instrument; magnetic field strength: 0.5T; magnet temperature: 32°C; probe diameter: 18mm; probe dead time: ≤10μs.

[0044] Step 1) Prepare modeling samples, use 17 (number of samples: >15) soybeans with known protein content as model samples, and record the corresponding sample masses; Test samples: 17 soybean samples with known protein content, record the sample masses:

[0045] Table 1 Test sample information

[0046]

[0047]

[0048] The soybean samples used are all stored soybeans with a moisture content of less than 12%. The protein content range of the above model samples covers the protein content range to be tested; the known protein content of soybeans can be measured by the Kjeldahl nitrogen method.

[0049] Step 2) The sample is kept at a constant temperature to a target temperature to eliminate the influence of sample temperature differences on the NMR signal. The target temperature is the temperature control temperature of the NMR equipment.

[0050] Step 3) Use the NMR equipment MagicSE-FSR-CPMG sequence to test the sample and record the peak signal; the main parameters are as follows: number of echoes: 500, number of inversions: 30, FID sampling time: 0.4ms, accumulation times 8;

[0051] Step 4) Screen all the recorded signal points to select the signal most relevant to the protein content. The GA-PLS (genetic algorithm based on partial least squares) selected as the screening algorithm uses Matlab 2019, Optimization toolbox, standard genetic algorithm. During the screening process, the crossover probability is set to 0.8, the mutation probability is 0.01, the population size is 50, and the evolutionary generation is 2000; the basic steps of GA-PLS are as follows:

[0052] ① Echo data encoding: each peak signal data sampling point is regarded as a gene and binary encoded. If the echo encoding gene is 1, the sampling point is selected, and if it is 0, it is not selected. The total number of sampling points in the present invention is 5500, and the gene encoding length is 5500;

[0053] ② Initialization of echo point number: randomly generate a group of 50 individuals, each group has different genes, representing a combination of different peak signal echo sampling point data;

[0054] ③Set the fitness function and define the function F(x)=RMSE+1 / R 2 , where RMSE represents the root mean square error of the model prediction after modeling all peak point signals in each population, R 2 Represents the prediction correlation coefficient of the model, the smaller the sum, the better;

[0055] The fitness function of each gene population is calculated using the set PLS (partial least squares) regression analysis model, and the populations are ranked from small to large.

[0056] Genetic Algorithm Operation:

[0057] ④Use the roulette wheel selection method. Calculate the proportion of each individual's fitness to the total fitness of the population as the probability of being selected. According to these probabilities, randomly select individuals to enter the breeding pool;

[0058] The fitness of an individual is Fi, The probability of an individual being selected is

[0059] ⑤ Pair the selected individuals in pairs and perform a crossover operation with a crossover probability of 0.8. Randomly select a crossover point on the binary code string of the paired individuals, and then exchange some genes after the crossover point;

[0060] ⑥ Perform mutation operations on individuals in the population with a mutation probability of 0.1. In binary coding, mutation means changing the 0 on the gene position to 1 or 1 to 0.

[0061] ⑦ Each time iterates, check whether the termination condition is met. It terminates when the maximum number of iterations reaches 2000, and the peak point signal is obtained. The peak point signal obtained by screening is as follows: Figure 1 shown.

[0062] Step 5) A protein content calculation model was established based on the peak signal obtained by the above screening and the protein content of each model sample obtained by actual measurement. The modeling algorithm was partial least squares (PLS) analysis with non-negative NN restrictions. When modeling, 5 potential variables were selected, and the cross-validation correlation coefficient R was ensured. 2 >0.9, the reason for selecting 5 latent variables is that when there are more than 5, the contribution score exceeds 99%;

[0063] The NMR test results of each model sample were brought into the established protein content calculation model to calculate the predicted protein content of each model sample, and then a working curve was established. The horizontal axis of the working curve was the true value of the sample protein content (signal / mass), and the vertical axis was the predicted protein content (such as Figure 2 The model interaction verification results are shown in Figure 3 shown.

[0064] Step 6) The sample to be tested is weighed and then steps 2) to 3) are repeated. The unit mass signal obtained by the measurement is brought into the protein content calculation model to obtain a predicted value. The predicted value of the protein content calculation model is then brought into the working curve to obtain the sample protein content.

[0065] The above established model was used to test the protein content in soybean samples. The test results are as follows:

[0066] Table 2 Test repeatability

[0067]

[0068] Table 3 Measurement accuracy

[0069]

[0070]

[0071] It can be seen from the above experimental data that the method of the present invention has high test accuracy, good repeatability and stability, and can meet the demand for protein content measurement in agricultural production. The test speed is fast, the operation is simple, and the production efficiency is improved. The method does not require labeling, avoiding protein denaturation and functional loss that may be caused during the labeling process. The method reduces the use of chemical reagents and is beneficial to environmental protection.

[0072] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.

Claims

1. A method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology, characterized in that: The steps include: Step 1) obtaining a soybean model sample series, wherein the soybean model sample series comprises a plurality of soybean samples with known protein contents and different protein contents; Step 2) using low-field nuclear magnetic resonance to detect the model sample series, and recording the peak signal data of each sample in the model sample series; Step 3) establishing a protein content calculation model according to the protein content of each sample in the model sample series and the peak signal data of each sample; Step 4) using low-field nuclear magnetic resonance to detect the soybean sample to be tested; using the protein content calculation model to calculate the protein content of the soybean sample to be tested based on the measured peak signal data.

2. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to claim 1, characterized in that: The protein content of the soybean samples in the soybean model sample series is 30-40 g / 100 g.

3. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to claim 1, characterized in that: The protein content of soybean samples in the soybean model sample series was determined by Kjeldahl method.

4. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to claim 1, characterized in that: The step 3) of establishing a protein content calculation model comprises the following steps: Step a) selecting peak signals with high correlation with protein content from the peak signal data of each sample; Step b) assigning weights to the screened peak signals; Step c) establishing a protein content calculation model according to the weighted peak point signals and the protein content of each sample in the model sample series.

5. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to claim 4, characterized in that: In the step a), a genetic algorithm based on partial least squares is used to screen out peak point signals with a high correlation with protein content from the peak point signal data of each sample.

6. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to claim 5, characterized in that: The method of using a genetic algorithm based on partial least squares to screen out peak point signals with a high correlation with protein content from the peak point signal data of each sample specifically includes: Step 1) Echo data encoding: each peak signal is regarded as a gene and binary encoded. If the encoded gene is 1, the peak signal is selected, and if it is 0, it is not selected; Step 2) Initializing the number of echo points: randomly generating n groups of encoded peak point signals, each group having different genes, representing different combinations of peak point signal data; Step 3) using the partial least squares regression analysis model to calculate the fitness function of each gene population and sort them from small to large; The fitness function F(x)=RMSE+1 / R 2 , where RMSE represents the root mean square error of the model prediction after modeling all peak point signals in each population, R 2 represents the prediction correlation coefficient of the model; Step 4) according to the sorting in step 3), exclude the population with a larger total in proportion, and then perform replication, crossover and mutation operations to form a new population. After each operation, record the peak signal to confirm the selected frequency and record the frequency probability; Step 5) Iterate according to step 3) and step 4) until the loop ends, and obtain the optimal peak signal for the final regression analysis operation, and finally screen out the peak signal with the highest frequency.

7. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to claim 4, characterized in that: In the step b), weights are assigned to the selected peak signals according to the frequencies, and the frequency weight parameter is set to 0.02; Preferably, the weight of the sampling points screened in step 6) is greater than 0.

2.

8. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to claim 4, characterized in that: The modeling algorithm is partial least squares analysis with non-negative NN restrictions.

9. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to claim 8, characterized in that: The protein content calculation model cross-validation correlation coefficient R 2 >0.

9.

10. The method for detecting protein content in soybeans based on low-field nuclear magnetic resonance technology according to any one of claims 1 to 9, characterized in that: The step 4) also includes: using the protein content calculation model to obtain the predicted value of the protein content of each sample in the soybean model sample series; then using the predicted value of the protein content of each sample in the soybean model sample series and the known protein content to establish a working curve; then using low-field nuclear magnetic resonance to detect the soybean sample to be tested; and using the protein content calculation model and the working curve to calculate the protein content of the soybean sample to be tested.