Establishment and evaluation method of prediction model suitable for processing dissolved leguminous plant varieties
By establishing a prediction model for solubility of legume plants based on machine learning, the problem that the existing technology cannot effectively evaluate and classify pea germplasm resources is solved, and the effect of high accuracy and rapid screening of suitable dissolving varieties is achieved.
Patent Information
- Application Number
- CN202411917121.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art cannot effectively evaluate and classify large-scale pea germplasm resources, especially in the quality of processing solubilized beans.
By obtaining the values of sensory quality, physical and chemical nutritional quality and solubility-related indicators of legume plant varieties, multiple machine learning algorithms are used to compare them to establish the best legume plant protein solubility quality evaluation prediction model. The model obtains the comprehensive value Y through a combination of principal component analysis and regression model, and is further simplified and classified by ridge regression model and K-means clustering analysis.
A high-accuracy and good prediction effect evaluation and prediction model for solubility of legume plant proteins is established, which can quickly screen out the varieties of legume plant suitable for processing and dissolving, simplify the evaluation process, and meet the market's demand for high-quality dissolved legume plant protein products.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of agricultural product processing, and in particular to an establishment and evaluation method of a prediction model suitable for processing dissolving legume varieties. Background Art
[0002] Variety quality is a key factor in market competitiveness, and quality evaluation is an important basis for the selection of improved varieties and the selection of varieties. Therefore, it is crucial for food processing companies to establish a variety quality evaluation prediction model for the production of corresponding functional products. Although traditional food quality evaluation methods have certain value, they still have limitations in scalability, ability to solve complex problems, and data-driven decision-making. The core advantage of machine learning technology is that it can discover complex patterns that traditional methods cannot detect, reducing the need for human intervention when dealing with tedious and repetitive tasks. Therefore, machine learning has opened up new horizons for variety quality evaluation and provided a new method for screening varieties with high quality. At present, domestic and foreign scholars have conducted relevant research on this. Some scholars have studied the prediction model of the association between wheat flour quality and bread quality, and some scholars have established a comprehensive quality evaluation model for walnuts and a comprehensive quality evaluation model for soy protein on low-salt emulsified sausages. At the same time, some scholars have established artificial neural network models to quickly screen high-quality beer during processing.
[0003] Pea (Pisum sativum L.) is a legume crop with high nutritional value and wide application value. It is also the second largest edible legume crop in the world. Peas are rich in a large number of nutrients, including protein (20% to 25%), starch (36.9% to 49%), dietary fiber (14% to 26%), non-starch polysaccharides (12% to 24%) and lipids (1.2% to 2.4%). Research results show that pea protein has good solubility. At present, the solubility characteristics of protein are widely used in the food industry, such as adding a large amount of protein to beverages and ham to increase their nutritional value. Therefore, proteins with good solubility are very popular among food processing companies.
[0004] At present, China has collected more than 5,000 pea germplasm resources and safely preserved more than 4,000 pea germplasm resources at home and abroad. There are large differences in the quality of different varieties of peas. Therefore, it is impossible to scientifically evaluate and classify the existing large-scale pea germplasm resources. In addition, there is no established solubility evaluation prediction model for bean protein and a quality evaluation and classification method for beans suitable for processing solubility. Summary of the invention
[0005] The present invention aims to solve one of the technical problems in the related art to at least a certain extent. To this end, one object of the present invention is to provide a method for establishing a prediction model for legume varieties suitable for processing and dissolving, and a method for evaluating legume varieties suitable for processing and dissolving. The model establishment method obtains the comprehensive value Y of legume varieties suitable for processing and dissolving, and uses multiple machine learning algorithms for comparison based on sensory quality-related indicators, physical and chemical nutritional quality-related indicators, etc., so as to establish an optimal legume protein solubility quality evaluation prediction model, and the model has high accuracy and good prediction effect. The evaluation method can quickly screen legume varieties suitable for processing and dissolving, greatly simplifies the evaluation process, enables scientific researchers and enterprises to quickly determine whether legume varieties are suitable for the development of soluble products, and helps to meet the market demand for high-quality soluble legume protein products.
[0006] To this end, the first aspect of the present invention provides a method for establishing a prediction model for processing dissolving legume varieties. In some embodiments of the present invention, the method comprises:
[0007] S1: Obtaining values of sensory quality related indicators, physical and chemical nutritional quality related indicators and solubility related indicators of fruit samples of multiple legume varieties;
[0008] S2: Standardizing the values of the solubility-related indexes of the fruit samples of the plurality of legume varieties, and adding them together to obtain a comprehensive value Y suitable for processing the solubility-type legume varieties;
[0009] S3: performing principal component analysis based on the values of the sensory quality-related indicators, the values of the physical and chemical nutrition quality-related indicators and the comprehensive value Y to determine the principal component scores;
[0010] S4: Establish multiple regression models based on the values of the sensory quality-related indicators, the values of the physical and chemical nutritional quality-related indicators, and the comprehensive value Y, and establish multiple principal component regression models based on the principal component scores to obtain the determination coefficient R of each model 2 value;
[0011] S5: R 2 The regression model with the value closest to 1 is taken as the prediction model suitable for processing dissolving legume varieties.
[0012] Among them, the solubility-related indicators refer to the protein subunit content, the content of proteins with different molecular weights, the content of allergenic proteins and the content of water-soluble proteins.
[0013] Existing methods only use a single method to establish an evaluation model, which has a low accuracy rate. The present invention obtains the comprehensive value Y of legume varieties suitable for processing dissolution, and uses multiple machine learning algorithms (such as principal component regression, stepwise regression, ridge regression algorithm, etc.) for comparison based on sensory quality related indicators, physical and chemical nutrition quality related indicators, etc., so as to establish an optimal legume protein solubility quality evaluation prediction model, which has high accuracy and good prediction effect.
[0014] In some embodiments of the present invention, the legume varieties include pea varieties, soybean varieties, red bean varieties, mung bean varieties, and black bean varieties.
[0015] In some embodiments of the present invention, the sensory quality related indicators include at least: 100-grain weight, length, width and thickness.
[0016] In some embodiments of the present invention, the physical and chemical nutritional quality related indicators include at least: crude protein, crude fat, crude fiber, starch, water, and amino acid content.
[0017] In some embodiments of the present invention, when the legume species is a pea species, the different molecular weight proteins include proteins of 14 kDa, 20 kDa, 34 kDa, 43 kDa, 44 kDa, 47 kDa, and 63 kDa.
[0018] In some embodiments of the invention, the protein subunits include globulin and albumin.
[0019] In some embodiments of the present invention, the allergenic proteins include a 44 kDa protein and a 63 kDa protein.
[0020] In some embodiments of the present invention, the standardization process is to subtract the mean value from each indicator data and then divide it by the standard deviation.
[0021] In some embodiments of the invention, the plurality of regression models include a principal component regression model, a stepwise regression model, a principal component-stepwise regression model, a ridge regression model, and a principal component-ridge regression model.
[0022] In some embodiments of the present invention, the establishment method further comprises:
[0023] The prediction model of the legume variety suitable for processing and dissolving determined in step S5 is further simplified based on the correlation, wherein the simplification is determined by the index value and / or protein content that are positively correlated with the dependent variable solubility.
[0024] In some embodiments of the invention, the plurality of legume species is at least 25 species.
[0025] In some embodiments of the present invention, the establishment method further comprises:
[0026] Before performing step S1, outlier analysis of quality and solubility is performed on fruit samples of multiple selected legume varieties to remove outlier legume varieties.
[0027] A second aspect of the present invention provides a method for evaluating leguminous plant varieties suitable for processing dissolving types. In some embodiments of the present invention, the evaluation method comprises:
[0028] Based on the prediction model established by the method for establishing a prediction model for legume varieties suitable for processing and dissolving as described in the first aspect, an evaluation factor for whether the legume variety to be tested is a legume variety suitable for processing and dissolving is obtained, so as to determine the suitability of the legume variety to be tested as a legume variety suitable for processing and dissolving.
[0029] In some embodiments of the present invention, the evaluation method comprises:
[0030] (I) Based on the prediction model established by the method for establishing a prediction model for processing dissolving legume varieties described in the first aspect, the sensory quality-related indexes and / or the physicochemical nutritional quality-related indexes and / or the protein content of the fruit samples of the legume varieties to be tested in the prediction model are measured;
[0031] (II) calculating, based on the values measured in step (I), an evaluation factor Z for determining whether the legume variety to be tested is a legume variety suitable for processing dissolving type;
[0032] (III) evaluating the suitability of the tested legume variety as a processing-dissolving legume based on the evaluation factor Z.
[0033] In some embodiments of the present invention, the evaluation method comprises:
[0034] A: Determine the width, serine content, tyrosine content, glycine content, 44kDa protein content and allergenic protein content of the fruit samples of the tested legume varieties;
[0035] B: Based on the indicators measured in step A, calculating the evaluation factor Z of the legume variety to be tested as a legume variety suitable for processing dissolving type;
[0036] C: evaluating the suitability of the tested legume variety as a processing-soluble legume based on the evaluation factor Z,
[0037] Wherein, the allergenic proteins include 44kDa protein and 63kDa protein.
[0038] In some embodiments of the present invention, the evaluation factor Z of the legume variety suitable for processing dissolving type is obtained by the following formula:
[0039] Z=-0.041a-0.058b+0.075c+0.026×d+0.172e+0.172f
[0040] Where a is the width value in millimeters;
[0041] b is the serine content in 100 g of the fruit of the tested legume species, in g;
[0042] c is the tyrosine content in 100 g of the fruit of the tested legume species, in g;
[0043] d is the glycine content in 100 g of the fruit of the tested legume species, in g;
[0044] e is the relative content of 44 kDa protein;
[0045] f is the relative content of allergenic protein.
[0046] In some embodiments of the present invention, step C further comprises:
[0047] Based on the evaluation factor Z, the ridge regression model combined with K-means cluster analysis was used to divide the leguminous plant varieties into three categories;
[0048] (1) If the Z value of the tested legume variety is ≥2.28, it is suitable for processing and dissolution;
[0049] (2) If the Z value of the leguminous plant variety to be tested is 0.77-2.28, it is basically suitable for processing and dissolution;
[0050] (3) If the Z value of the leguminous plant variety to be tested is ≤0.77, it is not suitable for processing and dissolution.
[0051] In some embodiments of the present invention, when the number of leguminous plant varieties to be tested is at least one, the evaluation method further comprises:
[0052] D: AHP hierarchical analysis combined with K-means cluster analysis was used to grade and assign scores to the sensory quality-related indicators, physical and chemical nutritional quality-related indicators, and comprehensive value Y of the fruit samples of legume varieties that were determined to be suitable for processing and dissolution and / or basically suitable for processing and dissolution, so as to obtain a comprehensive score for legume varieties that are suitable for processing and dissolution;
[0053] E: Compare the comprehensive scores and take the variety with the highest comprehensive score as the best legume variety suitable for processing and dissolving.
[0054] A third aspect of the present invention provides an evaluation device suitable for processing dissolving legume varieties. In some embodiments of the present invention, the evaluation device comprises:
[0055] An acquisition module, used to acquire sensory quality related indicators and / or physical and chemical nutritional quality related indicators and / or protein content of a fruit sample of a leguminous plant variety to be tested;
[0056] A calculation module is used to calculate the evaluation factor Z of the legume variety to be tested as a legume variety suitable for processing dissolving type according to the various indicators measured by the acquisition module;
[0057] Evaluation module: used for evaluating the suitability of the legume variety to be tested as a processing dissolving legume according to the evaluation factor Z,
[0058] Among them, the sensory quality-related indicators and / or physical and chemical nutritional quality-related indicators and / or protein content of the fruit sample are determined by the correlation of the prediction model established according to the method for establishing a prediction model for processing solubility-type legume varieties described in the first aspect.
[0059] A fourth aspect of the present invention provides an evaluation device suitable for processing dissolving legume varieties. In some embodiments of the present invention, the evaluation device comprises: an acquisition module for acquiring the width, serine content, tyrosine content, glycine content, 44 kDa protein content and allergenic protein content of a fruit sample of the legume variety to be tested;
[0060] A calculation module is used to calculate the evaluation factor Z of the legume variety to be tested as a legume variety suitable for processing dissolving type according to the various indicators measured by the acquisition module;
[0061] Evaluation module: used for evaluating the suitability of the legume variety to be tested as a processing dissolving legume according to the evaluation factor Z,
[0062] Among them, the allergenic proteins include 44kDa protein and 63kDa protein
[0063] In some embodiments of the present invention, the calculation module is further used to obtain the evaluation factor Z value according to the formula Z=-0.041a-0.058b+0.075c+0.026×d+0.172e+0.172f,
[0064] Where, in the formula, a is the width value in millimeters;
[0065] b is the serine content in 100 g of the fruit of the tested legume species, in g;
[0066] c is the tyrosine content in 100 g of the fruit of the tested legume species, in g;
[0067] d is the glycine content in 100 g of the fruit of the tested legume species, in g;
[0068] e is the relative content of 44 kDa protein;
[0069] f is the relative content of allergenic protein.
[0070] In some embodiments of the present invention, the evaluation module is further used to classify leguminous plant varieties into three categories according to the evaluation factor Z using a ridge regression model combined with a K-means cluster analysis;
[0071] (1) If the Z value of the tested legume variety is ≥2.28, it is suitable for processing and dissolution;
[0072] (2) If the Z value of the leguminous plant variety to be tested is 0.77-2.28, it is basically suitable for processing and dissolution;
[0073] (3) If the Z value of the leguminous plant variety to be tested is ≤0.77, it is not suitable for processing and dissolution.
[0074] In some embodiments of the present invention, the evaluation module is also used to make the following judgments:
[0075] When the number of the leguminous plant species to be tested is at least one, the method further comprises:
[0076] D: AHP hierarchical analysis combined with K-means cluster analysis was used to grade and assign scores to the sensory quality-related indicators, physical and chemical nutritional quality-related indicators, and comprehensive value Y of the fruit samples of legume varieties that were determined to be suitable for processing and dissolution and / or basically suitable for processing and dissolution, so as to obtain a comprehensive score for legume varieties that are suitable for processing and dissolution;
[0077] E: Compare the comprehensive scores and take the variety with the highest comprehensive score as the best legume variety suitable for processing and dissolving.
[0078] A fifth aspect of the present invention provides an electronic device. In some embodiments of the present invention, the electronic device includes:
[0079] at least one processor and at least one memory in communication with the processor,
[0080] The memory stores program instructions that can be executed by the processor.
[0081] The processor is capable of executing the evaluation method described in the second aspect.
[0082] A sixth aspect of the present invention provides a computer-readable storage medium. In some embodiments of the present invention, the computer-readable storage medium stores computer instructions for executing the evaluation method described in the second aspect.
[0083] Taking peas as an example, the existing means of evaluating the solubility of peas require the extraction of pea protein and the determination of the water-soluble protein content by the Coomassie Brilliant Blue method. The operation time is long and there are many steps. At present, a method for evaluating the quality of peas suitable for processing solubility has not been established. The present invention establishes a method for evaluating the suitability of pea solubility processing. By measuring the width, serine content, tyrosine content, glycine content, 44kDa content and allergenic protein content of pea samples, it is possible to predict whether an unknown pea variety is suitable for the processing of soluble products.
[0084] The present invention has the following beneficial effects:
[0085] (1) The present invention obtains the comprehensive value Y of the legume varieties suitable for processing solubility, and uses multiple machine learning algorithms (such as principal component regression, stepwise regression, ridge regression algorithm, etc.) for comparison based on sensory quality related indicators, physical and chemical nutrition quality related indicators, etc., so as to establish the best legume protein solubility quality evaluation prediction model, which has high accuracy and good prediction effect. The establishment process of the solubility evaluation prediction model can provide new ideas for other plant varieties.
[0086] (2) By measuring the width, serine content, tyrosine content, glycine content, 44kDa content and allergenic protein content of the legume samples to be tested, legume varieties suitable for processing soluble products can be quickly screened, which greatly simplifies the evaluation process and enables researchers and companies to quickly determine whether legume varieties are suitable for the development of soluble products, helping to meet the market demand for high-quality soluble legume protein products.
[0087] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0089] Figure 1 Box plots showing properties of pea raw materials (center black line is mean, n=30);
[0090] Figure 2The significant box plots of the quality characteristics of peas suitable for processing and dissolving are shown, including (a) protein subunit components and allergenic protein box plots, (b) 16 amino acids box plots), the central black line is the average value, n = 30;
[0091] Figure 3 The heat map of the correlation analysis between pea quality and the quality of pea varieties suitable for processing dissolving type is shown (n=25);
[0092] Figure 4 The ridge trace plot of ridge regression (a) and the ridge trace plot of principal component-ridge regression (b) are shown (n=20). DETAILED DESCRIPTION
[0093] The embodiments of the present invention are described in detail below. The embodiments described below are exemplary and are only used to explain the present invention, and should not be understood as limiting the present invention.
[0094] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. Further, in the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.
[0095] The endpoints and any values of the ranges disclosed in this article are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoint values of each range, the endpoint values of each range and the individual point values, and the individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed in this article.
[0096] In order to make the present invention more easily understood, certain technical and scientific terms are specifically defined below. Unless otherwise clearly defined elsewhere in this document, all other technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which the present invention belongs.
[0097] In this document, the terms “include” or “comprising” are open expressions, that is, including the contents specified in the present invention but not excluding other contents.
[0098] As used herein, the terms "optionally", "optional" or "optionally" generally mean that the subsequently described event or circumstance may but need not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0099] According to a specific embodiment of the present invention, the present invention provides a method for establishing a prediction model suitable for processing dissolving legume varieties, comprising:
[0100] S1: Obtaining values of sensory quality related indicators, physical and chemical nutritional quality related indicators and solubility related indicators of fruit samples of multiple legume varieties;
[0101] S2: Standardizing the values of the solubility-related indexes of the fruit samples of the plurality of legume varieties, and adding them together to obtain a comprehensive value Y suitable for processing the solubility-type legume varieties;
[0102] S3: performing principal component analysis based on the values of the sensory quality-related indicators, the values of the physical and chemical nutrition quality-related indicators and the comprehensive value Y to determine the principal component scores;
[0103] S4: Establish multiple regression models based on the values of the sensory quality-related indicators, the values of the physical and chemical nutritional quality-related indicators, and the comprehensive value Y, and establish multiple principal component regression models based on the principal component scores to obtain the determination coefficient R of each model 2 value;
[0104] S5: R 2 The regression model with the value closest to 1 is taken as the prediction model suitable for processing dissolving legume varieties.
[0105] Among them, the solubility-related indicators refer to the protein subunit content, the content of proteins with different molecular weights, the content of allergenic proteins and the content of water-soluble proteins.
[0106] According to a specific embodiment of the present invention, the sensory quality related indicators include at least: 100-grain weight, length, width and thickness; the physical and chemical nutritional quality related indicators include at least: crude protein, crude fat, crude fiber, starch, water, and amino acid content.
[0107] According to a specific embodiment of the present invention, when the leguminous plant variety is a pea variety, the different molecular weight proteins include proteins of 14kDa, 20kDa, 34kDa, 43kDa, 44kDa, 47kDa, and 63kDa. The protein subunits include globulin and albumin.
[0108] According to a specific embodiment of the present invention, the allergenic protein includes a 44 kDa protein and a 63 kDa protein.
[0109] According to a specific embodiment of the present invention, the standardization process includes but is not limited to subtracting the mean value of each indicator data and then dividing it by the standard deviation.
[0110] According to a specific embodiment of the present invention, the multiple regression models include but are not limited to a principal component regression model, a stepwise regression model, a principal component-stepwise regression model, a ridge regression model and a principal component-ridge regression model.
[0111] According to a specific embodiment of the present invention, the establishment method further comprises:
[0112] The prediction model of the legume variety suitable for processing and dissolving determined in step S5 is further simplified based on the correlation, wherein the simplification is determined by the index value and / or protein content that are positively correlated with the dependent variable solubility.
[0113] According to a specific embodiment of the present invention, the plurality of leguminous plant varieties is at least 25, for example, 25, 3, 40 or more.
[0114] According to a specific embodiment of the present invention, the establishment method may further include:
[0115] Before performing step S1, outlier analysis of quality and solubility is performed on fruit samples of multiple selected legume varieties to remove outlier legume varieties.
[0116] According to a specific embodiment of the present invention, the present invention provides an evaluation method for leguminous plant varieties suitable for processing dissolution, comprising:
[0117] (I) Based on the prediction model established by the method for establishing a prediction model suitable for processing dissolving legume varieties as described above, the sensory quality related indexes and / or the physicochemical nutritional quality related indexes and / or the protein content of the fruit samples of the legume varieties to be tested in the prediction model are measured;
[0118] (II) calculating, based on the values measured in step (I), an evaluation factor Z for determining whether the legume variety to be tested is a legume variety suitable for processing dissolving type;
[0119] (III) evaluating the suitability of the tested legume variety as a processing-dissolving legume based on the evaluation factor Z.
[0120] According to a specific embodiment of the present invention, the present invention provides an evaluation method for leguminous plant varieties suitable for processing dissolution, comprising:
[0121] A: Determine the width, serine content, tyrosine content, glycine content, 44kDa protein content and allergenic protein content of the fruit samples of the tested legume varieties;
[0122] B: Based on the indicators measured in step A, calculating the evaluation factor Z of the legume variety to be tested as a legume variety suitable for processing dissolving type;
[0123] C: evaluating the suitability of the tested legume variety as a processing-soluble legume based on the evaluation factor Z,
[0124] Wherein, the allergenic proteins include 44kDa protein and 63kDa protein.
[0125] According to a specific embodiment of the present invention, the evaluation factor Z suitable for processing dissolving legume varieties is obtained by the following formula:
[0126] Z=-0.041a-0.058b+0.075c+0.026×d+0.172e+0.172f
[0127] Where a is the width value in millimeters;
[0128] b is the serine content in 100 g of the fruit of the tested legume species, in g;
[0129] c is the tyrosine content in 100 g of the fruit of the tested legume species, in g;
[0130] d is the glycine content in 100 g of the fruit of the tested legume species, in g;
[0131] e is the relative content of 44 kDa protein;
[0132] f is the relative content of allergenic protein.
[0133] According to a specific embodiment of the present invention, step C further comprises:
[0134] Based on the evaluation factor Z, the ridge regression model combined with K-means cluster analysis was used to divide the leguminous plant varieties into three categories;
[0135] (1) If the Z value of the tested legume variety is ≥2.28, it is suitable for processing and dissolution;
[0136] (2) If the Z value of the leguminous plant variety to be tested is 0.77-2.28, it is basically suitable for processing and dissolution;
[0137] (3) If the Z value of the leguminous plant variety to be tested is ≤0.77, it is not suitable for processing and dissolution.
[0138] According to a specific embodiment of the present invention, when the number of leguminous plant varieties to be tested is at least one, the evaluation method further comprises:
[0139] D: AHP hierarchical analysis combined with K-means cluster analysis was used to grade and assign scores to the sensory quality-related indicators, physical and chemical nutritional quality-related indicators, and comprehensive value Y of the fruit samples of legume varieties that were determined to be suitable for processing and dissolution and / or basically suitable for processing and dissolution, so as to obtain a comprehensive score for legume varieties that are suitable for processing and dissolution;
[0140] E: Compare the comprehensive scores and take the variety with the highest comprehensive score as the best legume variety suitable for processing and dissolving.
[0141] According to a specific embodiment of the present invention, the present invention provides a method for evaluating pea varieties suitable for processing dissolving type, comprising:
[0142] The width, serine content, tyrosine content, glycine content, 44kDa content and allergenic protein content of the pea sample to be tested are measured and substituted into formula (I) to obtain the quality index of pea varieties suitable for processing and dissolving. The optimal ridge regression equation for the quality index of pea varieties suitable for processing and dissolving is:
[0143] Z = -0.041×width (mm) -0.058×serine (g / 100g) +0.075×tyrosine (g / 100g) +0.026×glycine (g / 100g) +0.172×44kDa (%) +0.172×allergenic protein (%) (I)
[0144] The score after entering the formula was used as the final score of a pea variety. The ridge regression model combined with K-means cluster analysis was used to divide the final scores of the pea varieties into three categories.
[0145] (1) If the calculated value of the pea variety is ≥2.28, the pea sample to be tested is suitable for processing and dissolution;
[0146] (2) If the calculated value of the pea variety is 0.77-2.28, the pea sample to be tested is basically suitable for processing and dissolution;
[0147] (3) If the calculated value of the pea variety is ≤0.77, the pea sample to be tested is not suitable for processing and dissolution.
[0148] This study selected 30 naturally air-dried high-quality pea seeds from different provinces in China as the research objects, and measured the sensory quality, physicochemical nutritional quality and solubility of pea protein of each pea variety. The relationship between various indicators was explored through correlation analysis. Subsequently, multiple machine learning algorithms (including principal component regression, stepwise regression, and ridge regression algorithms) were used for comparison to determine the best pea protein solubility quality evaluation prediction model. At the same time, a classification method for pea quality evaluation suitable for processing solubility was established through K-means cluster analysis, and the AHP hierarchical analysis method was used to screen out the pea varieties most suitable for solubility product processing, aiming to more effectively optimize the pea planting management structure and provide a reference for the future breeding of high-quality pea varieties.
[0149] The scheme of the present disclosure will be explained below in conjunction with the examples. Those skilled in the art will appreciate that the following examples are only used to illustrate the present disclosure and should not be considered to limit the scope of the present disclosure. Where specific techniques or conditions are not indicated in the examples, the techniques or conditions described in the literature in this area or the product instructions are used. Where the manufacturers of reagents or instruments are not indicated, they are all conventional products that can be obtained commercially.
[0150] Example 1 Establishment of a Pea Quality Determination Model Suitable for Dissolving Processing
[0151] 1. Pea varieties and outlier analysis
[0152] Thirty naturally air-dried high-quality pea seeds from Hebei, Shandong, Yunnan, Sichuan and Ningxia provinces (autonomous regions) in China were selected, including different fruit types and seed coat colors. The names and serial numbers of pea varieties are shown in Table 1.
[0153] Table 1 Names of 30 pea varieties
[0154] Serial number Variety name Origin Serial number Variety name Origin Serial number Variety name Origin 1 Gancui No.2 Shandong 11 Yunwan21 Yunnan 21 Ba Wan No.1 Hebei 2 Yunwan129 Yunnan 12 Zhonghua No.3 Hebei 22 Yunwan 125 Yunnan 3 YW90 Sichuan 13 Emerald Green No. 99 Ningxia 23 Ji Zhang Wan No. 5 Hebei 4 Taiwan longevity peas Ningxia 14 Yunwan123 Yunnan 24 Zhongqin No.1 Hebei 5 Zhong Wan No. 18 Shandong 15 Zhong Wan No. 4 Shandong 25 Zhong Wan No. 8 Shandong 6 Zhong Wan No.11 Shandong 16 Zhong Wan No.9 Shandong 26 Longevity Bean No.1 Shandong 7 YP5102 Sichuan 17 Tang Sweet Pea 895 Hebei 27 Tang Wan No. 3 Hebei 8 Zhong Wan No.6 Hebei 18 Yunwan128 Yunnan 28 Sugar snap peas Hebei 9 S2055 Sichuan 19 Ji Zhang Wan No.3 Hebei 29 Qizhen No. 76 Shandong 10 Yunwan121 Yunnan 20 Cloud Pea 122 Yunnan 30 Advance No. 1 Hebei
[0155] According to the box-type significance analysis diagram ( Figure 1 and Figure 2 ) The analysis of solubility grouping (dependent variable) and pea quality (independent variable) found that there were five pea varieties (Gancui 2, Zhongwan 6, Tangtianwan 895, Changshoudou 1 and Tangwan 3) that were comprehensive outliers in terms of the quality of pea varieties suitable for processing solubility and the quality of pea raw materials. Therefore, the remaining 25 pea varieties were used for subsequent analysis.
[0156] 2. Determination of pea raw material quality
[0157] The sensory quality and physicochemical nutritional quality of the remaining 25 pea varieties after removing the outliers were measured, with a total of 25 indicators.
[0158] Analysis of sensory quality of peas: 100-grain weight: refer to SN / T 0798-1999; length, width, thickness: refer to NY / T136-1989.
[0159] Analysis of pea physical and chemical nutritional quality: starch: refer to GB 5009.9-2016 method 2; crude protein: refer to GB5009.5-2016 method 1; moisture: refer to GB / T 21305-2007; crude fiber: refer to GB / T 5009.10-2003; crude fat: refer to GB 5009.6-2016 method 1; 16 kinds of amino acids content: refer to GB 5009.124-2016.
[0160] The variation range, mean, coefficient of variation, median and standard deviation of the basic data of the remaining 25 pea varieties after removing the outliers were analyzed, and the results are shown in Table 2. Among them, the coefficients of variation of 100-grain weight (0.04), length (0.09), width (0.07), crude protein (0.08), crude fat (0.09), starch (0.09), moisture (0.08), leucine (0.08) and glycine (0.09) were all less than 0.1. The coefficient of variation reflects the degree of data dispersion. The larger the coefficient of variation, the greater the difference between the indicators of different varieties.
[0161] Table 2 Descriptive analysis of quality characteristics of pea raw materials (n=25)
[0162] Quality indicators Distribution Mean Coefficient of variation Upper four scores median Lower Four Score Standard Deviation 100-grain weight (g) 16.07-30.53 23.00±0.71 0.04 20.30 22.57 24.83 0.97 Thickness(mm) 5.10-6.88 5.96±0.09 0.60 5.54 5.97 6.32 3.57 Length(mm) 6.83-9.70 7.79±0.13 0.09 7.24 7.91 8.22 0.67 Width(mm) 5.76-7.24 6.58±0.09 0.07 6.32 6.46 7.07 0.46 Crude protein (g / 100g) 20.85-28.45 23.83±0.39 0.08 22.25 23.45 25.05 1.97 Crude fat (g / 100g) 3.00-4.25 3.52±0.06 0.09 3.30 3.40 3.78 0.32 Crude fiber (g / 100g) 5.40-8.75 6.61±0.17 0.13 5.95 6.50 6.93 0.86 Starch (g / 100g) 43.00-57.55 50.90±0.96 0.09 45.30 52.15 54.65 4.80 Water content (g / 100g) 10.14-13.31 11.94±0.20 0.08 11.09 12.16 12.74 0.98 Glutamic acid (g / 100g) 2.24-4.10 3.40±0.08 0.12 3.15 3.43 3.67 0.40 Aspartic acid (g / 100g) 1.65-2.89 2.35±0.05 0.11 2.23 2.32 2.54 0.26 Arginine (g / 100g) 1.40-2.13 1.78±0.04 0.11 1.68 1.74 1.96 0.20 Lysine(g / 100g) 1.40-2.33 1.82±0.04 0.13 1.65 1.74 2.00 0.23 Leucine (g / 100g) 1.18-1.76 1.43±0.02 0.08 1.37 1.42 1.51 0.12 Serine(g / 100g) 0.85-2.30 1.02±0.05 0.26 0.92 0.97 1.03 0.27 Phenylalanine(g / 100g) 0.91-1.21 0.98±0.02 0.10 0.90 1.00 1.03 0.10 Alanine(g / 100g) 0.72-1.20 0.88±0.02 0.11 0.83 0.86 0.92 0.10 Glycine (g / 100g) 0.76-1.14 0.91±0.02 0.09 0.86 0.88 0.94 0.08 Valine (g / 100g) 0.72-1.20 0.88±0.02 0.11 0.83 0.86 0.92 0.10 Threonine (g / 100g) 0.66-1.02 0.82±0.02 0.10 0.78 0.84 0.89 0.08 Histidine(g / 100g) 0.44-0.67 0.52±0.01 0.12 0.48 0.53 0.57 0.06 Tyrosine (g / 100g) 0.38-0.99 0.48±0.02 0.27 0.43 0.45 0.50 0.13 Proline (g / 100g) 0.68-1.23 0.95±0.03 0.16 0.83 0.88 1.09 0.15 Isoleucine (g / 100g) 0.62-1.03 0.76±0.02 0.11 0.72 0.75 0.78 0.08 Methionine (g / 100g) 0.07-0.13 0.09±0.00 0.22 0.08 0.09 0.10 0.02
[0163] 3. Determination of the quality of pea varieties suitable for processing dissolving
[0164] (1) Pea protein extraction
[0165] 100 g of peas of different varieties were soaked, peeled manually, dried, and pressed into powder. The pea powder was dissolved in a beaker with 30 °C deionized water at a ratio of 1:3, and then stirred evenly on a magnetic stirrer. -1 Centrifuge for 20 min, collect the supernatant through a 100-mesh sieve, and then slowly add 0.1 mol·L -1 HCL, adjust the pH to 5.8-6.2, and then add 1 mL of lactic acid streptococcus suspension (1.58×10 8 CFU / mL), mixed evenly on a magnetic stirrer, placed in a 4°C refrigerator for 2 h, and then stirred at 4000 r·min -1 Centrifuge for 20 minutes, collect the precipitate, bag it for freeze-drying, and then place the pea protein powder in a -80°C refrigerator for use.
[0166] (2) Determination of pea protein subunits and allergenic protein content
[0167] Pea protein was subjected to SDS-PAGE electrophoresis. After the gel was stained in a staining solution for 1 hour, it was placed on a shaker for decolorization and image collection. The gel staining solution used Coomassie Brilliant Blue solution, and the decolorization used 45% methanol in acetic acid solution. The relative contents of 14kDa, 20kDa, 34kDa, 43kDa, 44kDa, 47kDa, 63kDa and protein subunits (globulin and albumin) were calculated using the optical density analysis software Image Lab. The allergenic protein content is the sum of the 44kDa and 63kDa contents.
[0168] (3) Solubility analysis
[0169] Prepare Coomassie Brilliant Blue solution: weigh 100 mg of Coomassie Brilliant Blue G-250, dissolve it in 50 mL of 95% ethanol, then add 100 mL of 85% phosphoric acid, and finally dilute to 1000 mL with distilled water. -1 Bovine serum albumin (BSA) standard solution: weigh 0.1 g of BSA and dissolve it in 100 mL of distilled water. Store at 4°C after aliquoting. Draw a protein standard curve: take 6 stoppered test tubes and add different volumes of 1000 μg mL -1 of standard BSA solution and mix well. Take 0.1mL of the above series of concentrations of BSA solution, add 4mL of Coomassie Brilliant Blue solution, shake well and let it stand for about 2 minutes, then use a cuvette to measure the absorbance at 595nm. Use the standard protein content as the horizontal axis and the absorbance at 595nm as the vertical axis to make a standard curve. Determine the soluble protein content: Take 0.1g of pea protein powder, place it in a 250mL conical flask, add 40mL of distilled water, and rotate at 4000r·min -1 Centrifuge for 10 min, take 100 μL of supernatant, add 4 mL of Coomassie Brilliant Blue solution, shake well and let stand for about 2 min, then use a cuvette to measure the absorbance at 595 nm to determine the water-soluble protein content (μg mL -1 ).
[0170] After removing the outliers, the remaining 25 pea varieties were subjected to a descriptive analysis based on the basic data of the quality associated with whether they were suitable for processing into soluble products, and the results are shown in Table 3. As shown in Table 3, the coefficients of variation of globulin (0.06) and 20 kDa (0.06) were both less than 0.1, indicating that the quality of different pea varieties was suitable for processing soluble products. In addition to the above four indicators, the other indicators of pea variety quality varied greatly among different pea varieties.
[0171] Table 3 Descriptive analysis of quality characteristics of peas suitable for processing dissolution type (n=25)
[0172]
[0173]
[0174] 4. Analysis of comprehensive quality values of pea varieties suitable for processing dissolving
[0175] As shown in Table 3, there are 11 quality indicators for evaluating pea varieties suitable for processing solubility. It is difficult to simply use one indicator to evaluate whether a pea variety is suitable for processing into a solubility product. Therefore, this study fits the 11 indicators for evaluating the quality of pea varieties suitable for processing solubility into one indicator for subsequent model establishment.
[0176] (1) Changes in quality indices of pea varieties suitable for processing solubility. Among the 11 indices, some indices are more suitable when they are larger, while others are more suitable when they are smaller. Therefore, for the convenience of subsequent calculations, the 11 quality indices for evaluating pea varieties suitable for processing solubility are changed to the larger the better.
[0177] (2) Data standardization. After the quality indicators of the 11 pea varieties suitable for processing solubility were changed to the larger the better, they were standardized, that is, the mean value of each indicator was subtracted and divided by the standard deviation.
[0178] (3) Calculation of comprehensive value. The processed standardized data are added up and recorded as Y, which is the comprehensive value of the quality of pea varieties suitable for processing solubility, as shown in Table 4.
[0179] Table 4 Comprehensive values of pea varieties (n=25)
[0180] Serial number Variety name Comprehensive value Y Serial number Variety name Comprehensive value Y 1 Yunwan129 -0.19 16 Ji Zhang Wan No.3 -2.12 2 YW90 -0.39 17 Cloud Pea 122 -3.15 3 Taiwan longevity peas 3.21 18 Ba Wan No.1 4.00 4 Zhong Wan No. 18 -1.17 19 Yunwan 125 2.99 5 Zhong Wan No.11 2.31 20 Ji Zhang Wan No. 5 6.20 6 YP5102 1.53 21 Zhongqin No.1 -5.45 7 S2055 -2.72 22 Zhong Wan No. 8 -3.11 8 Yunwan121 -2.06 23 Sugar snap peas 4.62 9 Yunwan21 2.82 24 Qizhen No. 76 3.59 10 Zhonghua No.3 -8.77 25 Advance No. 1 4.60 11 Emerald Green No. 99 -1.87 12 Yunwan123 -0.14 13 Zhong Wan No. 4 -1.20 14 Zhong Wan No.9 0.10 15 Yunwan128 -0.76
[0181] 5. Establishment of a quality evaluation prediction model for pea varieties suitable for processing dissolving
[0182] (1) Correlation analysis between pea quality and quality of pea varieties suitable for processing dissolving
[0183] Correlation analysis of 36 indicators of 25 pea varieties revealed that ( Figure 3), pea length was significantly positively correlated with glycine (0.50), leucine (0.47), and histidine (0.53); pea thickness was significantly negatively correlated with 34kDa (-0.50), alanine (-0.47), methionine (-0.47), and phenylalanine (-0.57), and was significantly positively correlated with 47kDa (0.42); pea starch was significantly negatively correlated with glycine (-0.58), alanine (-0.66), valine (-0.63), isoleucine (-0.59), leucine (-0.56), and lysine (-0.57), and was significantly positively correlated with 47kDa (0.48); pea crude protein was significantly negatively correlated with aspartic acid (0.78), threonine (0.63), glutamic acid (0.56), glycine (0.81), alanine (0.8 1), valine (0.76), methionine (0.57), isoleucine (0.72), leucine (0.81), phenylalanine (0.58), lysine (0.72), histidine (0.65), and arginine (0.86) were extremely significantly positively correlated; pea crude fiber was extremely significantly positively correlated with 14kDa (0.48), 34kDa (0.68), threonine (0.47), glycine (0.57), alanine (0.65), valine (0.55), methionine (0.50), isoleucine (0.51), leucine (0.53), lysine (0.49), and histidine (0.53), and was extremely significantly negatively correlated with globulin (-0.53) and 47kDa (-0.58); pea crude fat was significantly negatively correlated with arginine (-0.43).
[0184] (2) Multicollinearity test
[0185] Considering that there may be multicollinearity between variables due to high correlation, the stability and explanatory power of the model will be negatively affected, resulting in unstable and inaccurate model parameter estimation. Therefore, it is necessary to test the variables for multicollinearity before model construction to improve the reliability of regression analysis results. Variance inflation factor (VIF) and tolerance factor (1 / VIF) are commonly used indicators for judging multicollinearity. Generally speaking, a VIF value greater than 10 and a 1 / VIF value less than 0.2 indicate that there is multicollinearity between variables. To solve the problem of multicollinearity, the commonly used methods are to increase the sample size, delete redundant variables, and adjust the model construction method. The commonly used models to solve multicollinearity are ridge regression model, principal component regression model, stepwise regression model, and lasso regression model. As shown in Table 5, the results of the multicollinearity test of modeling indicators are shown. The VIF values of most modeling indicators are greater than 10 and 1 / VIF is less than 0.2, which means that there is a serious multicollinearity problem between the modeling indicators, so this problem must be corrected. In this study, principal component regression model, stepwise regression model, and ridge regression model will be established to solve the problem of multicollinearity.
[0186] Table 5 Multicollinearity test
[0187]
[0188]
[0189] (3) Construction of principal component analysis and principal component regression model
[0190] In order to reduce the impact of multicollinearity on the model, a principal component regression model was first established to solve the collinearity problem. However, the principal component analysis of the modeling index data was required before the principal component regression model was established. All pea samples were randomly divided into a training set (n = 20) and a validation set (n = 5) for subsequent model construction. The results of the principal component analysis of the 20 samples in the training set are shown in Table 6. With F1-F6 as the scores of principal components 1-6, the functional expressions of the six principal components are obtained as follows:
[0191] F1=0.288X 1 +0.286X 2 +...-0.096X 25 -0.002X 26
[0192] F2=0.004X 1 +0.061X 2 +...-0.107X 25 +0.036X 26
[0193] F3=0.030X 1 +0.053X 2 -...+0.078X 25 +0.069X 26
[0194] F4=0.115X 1 +0.087X 2 +...+0.271X 25 +0.007X 26
[0195] F5=-0.029X 1 -0.020X 2 +...+0.534X 25 +0.484X 26
[0196] F6=0.081X 1 +0.104X 2 +...-0.101X 25 +0.674X26
[0197] As shown in Table 6, the eigenvalues of the first six principal components of the quality index are 11.38, 3.47, 3.08, 2.32, 1.75 and 1.11, all greater than 1; the variance contribution rates of the first six principal components are 43.77%, 13.36%, 11.83%, 8.94%, 6.71% and 4.26%, respectively, and the cumulative contribution rate is 88.86%. Therefore, these six components can reflect most of the information of the modeling indicators, so the first six components can be selected for analysis.
[0198] Table 6 Principal component analysis results
[0199]
[0200]
[0201] According to the results of the principal component, the weight of each index is calculated. Among them, glutamic acid and histidine have the highest weight. It can be seen that the content of glutamic acid and histidine plays an important role in evaluating whether the unknown pea variety is suitable for the processing of soluble products. The principal component scores of each pea sample are calculated and ranked. The results are shown in Table 7. From Table 7, it can be seen that the highest ranking pea variety is Zhongwan No. 9, which shows that Zhongwan No. 9 has certain feasibility as a raw material for the processing of soluble products. The principal component regression method is used to construct a multivariate linear regression model with the 6 principal component scores of 20 randomly selected pea samples as independent variables and the solubility grouping as the dependent variable. The significance and multicollinearity of the regression model are tested, and the regression equation is finally obtained as Z = -0.02Z1 + 0.11Z2 + 0.11Z3 + 0.16Z5 + 0.46Z6, and its correlation coefficient is 0.36, indicating that the explanation rate of the model for whether the unknown pea variety is suitable for the processing of soluble products is 36%.
[0202] Table 7 Principal component scores
[0203]
[0204] (4) Model comparison
[0205] In addition to the principal component regression model, which can solve the problem of multicollinearity, the stepwise regression model and the ridge regression model can also solve the problem of multicollinearity well.
[0206] Multivariate stepwise regression analysis and principal component-stepwise regression analysis were used to establish a solubility evaluation prediction model for 20 randomly selected pea varieties. A stepwise regression model was constructed with 26 modeling indicators as independent variables and solubility grouping as the dependent variable. At the same time, combined with Table 7, a principal component-stepwise regression model was constructed with 6 principal component scores as independent variables and solubility grouping as the dependent variable. The significance and multicollinearity of the two regression models were tested. The final results can be seen in Table 8. The regression equations obtained were Z = -1.96E-6 + 0.46 × comprehensive value Y (R 2 =0.21) and Z = -1.50E-6 + 0.46Z6 (R 2 =0.24).
[0207] Table 8 Pea quality and suitable processing solubility Pea quality overall regression analysis after stepwise regression and principal component-stepwise regression coefficient parameter estimates (n = 20)
[0208]
[0209]
[0210] Since the model explanation rates of the established principal component regression model, stepwise regression model and principal component-stepwise regression model are relatively low, in order to establish a more stable and effective pea protein solubility evaluation prediction model, a ridge regression model was constructed with 26 modeling indicators as independent variables and solubility grouping as the dependent variable. At the same time, combined with Table 7, a principal component-ridge regression model was constructed with 6 principal component scores as independent variables and solubility grouping as the dependent variable. The ridge traces of the two models are shown in Figure 2. Figure 4When k = 0.42, the ridge regression coefficient begins to stabilize, and the obtained ridge regression equation is Z = -0.16×length (mm) -0.04×width (mm) -0.10×thickness (mm) +0.13×100-grain weight (g) -0.05×starch (g / 100g) +0.08×protein (g / 100g) +0.16×water (g / 100g) -0.06×crude fiber (g / 100g) + 0.16 × crude fat (g / 100g) - 0.09 × aspartic acid (g / 100g) + 0.11 × threonine (g / 100g) - 0.06 × serine (g / 100g) + 0.05 × glutamic acid (g / 100g) - 0.17 × proline (g / 100g) + 0.08 × glycine (g / 100g) + 0.13 × alanine ( g / 100g) + 0.07 × valine (g / 100g) + 0.07 × methionine (g / 100g) + 0.07 × isoleucine (g / 100g) + 0.13 × leucine (g / 100g) + 0.03 × tyrosine (g / 100g) + 0.04 × phenylalanine (g / 100g) - 0.11 × lysine (g / 100g) + 0.11 × histidine ( g / 100g)-0.10×arginine (g / 100g)+0.17×comprehensive value Y, and its correlation coefficient is 0.97. When k=0.62, the principal component-ridge regression coefficient begins to stabilize, and the obtained ridge regression equation is Z=-0.14Z1+0.15Z2+0.13Z3+0.15Z4+0.15Z5+0.35Z6, and its correlation coefficient is 0.57.
[0211] The principal component regression model, stepwise regression model, principal component-stepwise regression model, ridge regression model and principal component-ridge regression model established above were comprehensively compared. Among them, the ridge regression model has the highest model explanation rate (R 2 =0.97), indicating that this model can well predict whether unknown pea varieties are suitable for the processing of soluble products. Therefore, the ridge regression model is the best pea protein solubility evaluation prediction model.
[0212] Since the ridge regression equation needs to detect all indicators, it is further simplified. According to the correlation analysis results ( Figure 3), the solubility of pea protein was significantly positively correlated with width (r = 0.39), 44kDa (r = 0.48), and allergenic protein (r = 0.43), and was positively correlated with serine (r = 0.18), tyrosine (r = 0.16), and glycine (r = 0.11), so the simplified ridge regression equation was Z = -0.041 × width (mm) -0.058 × serine (g / 100g) + 0.075 × tyrosine (g / 100g) + 0.026 × glycine (g / 100g) + 0.172 × 44kDa (%) + 0.172 × allergenic protein (%)
[0213] Example 2 Determination of Pea Solubility Quality
[0214] The various indicators of the 20 pea varieties in the training set and the 5 pea varieties in the validation set were substituted back into the ridge regression equation for internal validation and external validation, respectively. The results are shown in Table 9. The internal validation accuracy of the ridge regression model was 50%, and the external validation accuracy was 60%.
[0215] Table 9 Internal and external validation of ridge regression model
[0216]
[0217] Example 3 Establishment of a quality evaluation method for peas suitable for processing and dissolving
[0218] The K-means cluster analysis method was used to preliminarily classify the standardized values of the original grouping of 25 pea varieties, and they were initially divided into three categories (suitable, basically suitable and unsuitable), and the cluster center of each category was determined. At the same time, the ridge regression model combined with K-means cluster analysis was used to classify the 25 pea varieties, and the classification results are shown in Table 10. The matching degree between the two was obtained as follows: 50.00% of the pea varieties suitable for soluble product processing, 50.00% of the pea varieties basically suitable for soluble product processing, and 66.67% of the pea varieties unsuitable for soluble product processing, indicating that the evaluation result is good and suitable as an evaluation standard for screening pea varieties suitable for soluble product processing.
[0219] Table 10 Classification of pea varieties (n=25)
[0220]
[0221]
[0222] AHP hierarchical analysis combined with K-means cluster analysis was used to grade and assign scores to the modeling indicators, and the results are shown in Table 11. Finally, the comprehensive score of pea varieties suitable for processing soluble products was obtained (Table 12), and the results showed that sweet snap peas were the best pea variety suitable for processing soluble products.
[0223] Table 11 The scores of each index and each level obtained by AHP hierarchical analysis combined with K-means cluster analysis
[0224]
[0225]
[0226]
[0227] Note: '+' means the larger the index, the better; '-' means the smaller the index, the better
[0228] Table 12 Comprehensive scores of pea varieties suitable for soluble product processing
[0229] serial number Variety name Comprehensive score 1 Sugar snap peas 61 2 Zhong Wan No.9 60 3 Qizhen No. 76 53 4 Zhong Wan No.4 53 5 Yunwan 125 51 6 Zhong Wan No. 18 49 7 Zhong Wan No.11 49 8 Ji Zhang Wan No.3 45 9 Ba Wan No.1 43 10 Yunwan129 43
[0230] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "some implementation schemes" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0231] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A method for establishing a prediction model for processing dissolving legume varieties, characterized in that: include: S1: Obtaining values of sensory quality related indicators, physical and chemical nutritional quality related indicators and solubility related indicators of fruit samples of multiple legume varieties; S2: Standardizing the values of the solubility-related indexes of the fruit samples of the plurality of legume varieties, and adding them together to obtain a comprehensive value Y suitable for processing the solubility-type legume varieties; S3: performing principal component analysis based on the values of the sensory quality-related indicators, the values of the physical and chemical nutrition quality-related indicators and the comprehensive value Y to determine the principal component scores; S4: Establish multiple regression models based on the values of the sensory quality-related indicators, the values of the physical and chemical nutritional quality-related indicators, and the comprehensive value Y, and establish multiple principal component regression models based on the principal component scores to obtain the determination coefficient R of each model 2 value; S5: R 2 The regression model with the value closest to 1 is taken as the prediction model suitable for processing dissolving legume varieties. Among them, the solubility-related indicators refer to the protein subunit content, the content of proteins with different molecular weights, the content of allergenic proteins and the content of water-soluble proteins.
2. The establishment method according to claim 1, characterized in that: The legume plant varieties include pea varieties, soybean varieties, red bean varieties, mung bean varieties, and black bean varieties.
3. The establishment method according to claim 1, characterized in that: The sensory quality related indicators include at least: 100-grain weight, length, width and thickness; Optionally, the physical and chemical nutritional quality related indicators include at least: crude protein, crude fat, crude fiber, starch, water, and amino acid content.
4. The establishment method according to claim 1, characterized in that: When the leguminous plant variety is a pea variety, the different molecular weight proteins include proteins of 14 kDa, 20 kDa, 34 kDa, 43 kDa, 44 kDa, 47 kDa, and 63 kDa; Optionally, the protein subunits include globulin and albumin; Optionally, the allergenic proteins include a 44 kDa protein and a 63 kDa protein.
5. The establishment method according to claim 1, characterized in that: The standardization process is to subtract the mean value from each indicator data and then divide it by the standard deviation.
6. The establishment method according to claim 1, characterized in that: The multiple regression models include a principal component regression model, a stepwise regression model, a principal component-stepwise regression model, a ridge regression model and a principal component-ridge regression model.
7. The establishment method according to claim 1, characterized in that: The establishment method further comprises: The prediction model of the legume variety suitable for processing and dissolving determined in step S5 is further simplified based on the correlation, wherein the simplification is determined by the index value and / or protein content that are positively correlated with the dependent variable solubility.
8. The establishment method according to claim 1, characterized in that: The plurality of leguminous plant varieties are at least 25; Optionally, the establishment method further comprises: Before performing step S1, outlier analysis of quality and solubility is performed on fruit samples of multiple selected legume varieties to remove outlier legume varieties.
9. A method for evaluating leguminous plant varieties suitable for processing dissolving, characterized in that: include: Based on the prediction model established by the method for establishing a prediction model for a legume variety suitable for processing and dissolving according to any one of claims 1 to 8, an evaluation factor for whether the legume variety to be tested is a legume variety suitable for processing and dissolving is obtained, so as to determine the suitability of the legume variety to be tested as a legume variety suitable for processing and dissolving.
10. The evaluation method according to claim 9, characterized in that: include: (I) Based on the prediction model established by the method for establishing a prediction model for leguminous plant varieties suitable for processing and dissolving according to any one of claims 1 to 8, the sensory quality-related indexes and / or the physicochemical nutritional quality-related indexes and / or the protein content of the fruit samples of the leguminous plant varieties to be tested in the prediction model are measured; (II) calculating, based on the values measured in step (I), an evaluation factor Z for determining whether the legume variety to be tested is a legume variety suitable for processing dissolving type; (III) evaluating the suitability of the tested legume variety as a processing-dissolving legume based on the evaluation factor Z.
11. The evaluation method according to claim 9, characterized in that: include: A: Determine the width, serine content, tyrosine content, glycine content, 44kDa protein content and allergenic protein content of the fruit samples of the tested legume varieties; B: Based on the indicators measured in step A, calculating the evaluation factor Z of the legume variety to be tested as a legume variety suitable for processing dissolving type; C: evaluating the suitability of the tested legume variety as a processing-soluble legume based on the evaluation factor Z, Wherein, the allergenic proteins include 44kDa protein and 63kDa protein.
12. The evaluation method according to claim 11, characterized in that: The evaluation factor Z of the legume variety suitable for processing dissolving type is obtained by the following formula: Z=-0.041a-0.058b+0.075c+0.026×d+0.172e+0.172f Where a is the width value in millimeters; b is the serine content in 100 g of the fruit of the tested legume species, in g; c is the tyrosine content in 100 g of the fruit of the tested legume species, in g; d is the glycine content in 100 g of the fruit of the tested legume species, in g; e is the relative content of 44 kDa protein; f is the relative content of allergenic protein.
13. The evaluation method according to claim 11, characterized in that: Step C further comprises: Based on the evaluation factor Z, the ridge regression model combined with K-means cluster analysis was used to divide the leguminous plant varieties into three categories; (1) If the Z value of the tested legume variety is ≥2.28, it is suitable for processing and dissolution; (2) If the Z value of the leguminous plant variety to be tested is 0.77-2.28, it is basically suitable for processing and dissolution; (3) If the Z value of the leguminous plant variety to be tested is ≤0.77, it is not suitable for processing and dissolution.
14. The evaluation method according to claim 11, characterized in that: When the number of leguminous plant varieties to be tested is at least one, the evaluation method further comprises: D: AHP hierarchical analysis combined with K-means cluster analysis was used to grade and assign scores to the sensory quality-related indicators, physical and chemical nutritional quality-related indicators, and comprehensive value Y of the fruit samples of legume varieties that were determined to be suitable for processing and dissolution and / or basically suitable for processing and dissolution, so as to obtain a comprehensive score for legume varieties that are suitable for processing and dissolution; E: Compare the comprehensive scores and take the variety with the highest comprehensive score as the best legume variety suitable for processing and dissolving.
15. An evaluation device suitable for processing dissolving leguminous plant varieties, characterized in that: include: An acquisition module, used to acquire sensory quality related indicators and / or physical and chemical nutritional quality related indicators and / or protein content of a fruit sample of a leguminous plant variety to be tested; A calculation module is used to calculate the evaluation factor Z of whether the legume variety to be tested is suitable for processing dissolving legume variety according to the various indicators measured by the acquisition module; Evaluation module: used for evaluating the suitability of the legume variety to be tested as a processing dissolving legume according to the evaluation factor Z, The sensory quality-related indexes and / or physicochemical nutritional quality-related indexes and / or protein content of the fruit sample are determined by the correlation of the prediction model established according to the method for establishing a prediction model for processing-soluble legume varieties according to any one of claims 1 to 8.
16. An evaluation device suitable for processing dissolving leguminous plant varieties, characterized in that: include: An acquisition module, used for acquiring the width, serine content, tyrosine content, glycine content, 44kDa protein content and allergenic protein content of the fruit sample of the legume plant variety to be tested; A calculation module is used to calculate the evaluation factor Z of whether the legume variety to be tested is suitable for processing dissolving legume variety according to the various indicators measured by the acquisition module; Evaluation module: used for evaluating the suitability of the legume variety to be tested as a processing dissolving legume according to the evaluation factor Z, Wherein, the allergenic proteins include 44kDa protein and 63kDa protein.
17. The evaluation device according to claim 16, characterized in that The calculation module is also used to obtain the evaluation factor Z value according to the formula Z=-0.041a-0.058b+0.075c+0.026×d+0.172e+0.172f. Where, in the formula, a is the width value in millimeters; b is the serine content in 100 g of the fruit of the tested legume species, in g; c is the tyrosine content in 100 g of the fruit of the tested legume species, in g; d is the glycine content in 100 g of the fruit of the tested legume species, in g; e is the relative content of 44 kDa protein; f is the relative content of allergenic protein.
18. The evaluation device according to claim 16, characterized in that The evaluation module is also used to classify leguminous plant varieties into three categories according to the evaluation factor Z by using a ridge regression model combined with K-means cluster analysis; (1) If the Z value of the tested legume variety is ≥2.28, it is suitable for processing and dissolution; (2) If the Z value of the leguminous plant variety to be tested is 0.77-2.28, it is basically suitable for processing and dissolution; (3) If the Z value of the tested legume variety is ≤0.77, it is not suitable for processing and dissolution; Optionally, the evaluation module is further used to make the following judgments: When the number of the leguminous plant species to be tested is at least one, the method further comprises: D: AHP hierarchical analysis combined with K-means cluster analysis was used to grade and assign scores to the sensory quality-related indicators, physical and chemical nutritional quality-related indicators, and comprehensive value Y of the fruit samples of legume varieties that were determined to be suitable for processing and dissolution and / or basically suitable for processing and dissolution, so as to obtain a comprehensive score for legume varieties that are suitable for processing and dissolution; E: Compare the comprehensive scores and take the variety with the highest comprehensive score as the best legume variety suitable for processing and dissolving.
19. An electronic device, characterized in that: include: at least one processor and at least one memory in communication with the processor, The memory stores program instructions that can be executed by the processor. The processor is capable of executing the evaluation method according to any one of claims 9 to 14.
20. A computer-readable storage medium, characterized in that: The computer program product stores computer instructions for executing the evaluation method according to any one of claims 9 to 14.