A method for evaluating conformational changes of plant fibrin molecules

Through Gaussian fitting and feature peak analysis, a method for evaluating the conformational change of plant fibrin molecules was constructed, which solved the evaluation problems caused by large molecular weight and complex structure, and achieved an accurate assessment of the mimicking ability of plant fibrin molecules and animal proteins.

CN119104514BActive Publication Date: 2025-08-08NINGBO SULIAN FOOD CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411109073.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2025-08-08
Estimated Expiration
2044-08-13

AI Technical Summary

Technical Problem

Due to the large molecular weight and complex structure of the existing plant fibrin molecules, it is difficult to extract features with traditional secondary structure infrared spectroscopy processing methods, and thus it is difficult to accurately conduct conformation evaluation.

Method used

The Gaussian fitting algorithm was used to obtain the infrared spectral characteristic peaks of plant fibrin and animal proteins. Through the analysis of feature peak ambiguity weight and confidence, the conformation change evaluation decision function was constructed, the content differences between the two secondary structures were compared, and the support vector machine algorithm was trained for evaluation.

Benefits of technology

The accurate evaluation of plant fibrin molecules and animal protein mimicking ability is achieved, avoiding subjective errors and noise interference in traditional methods, and improving the accuracy of conformational evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119104514B_ABST
    Figure CN119104514B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of protein molecule conformational change evaluation, and specifically to a method for evaluating conformational changes in plant fibrin molecules. The method comprises: obtaining plant fibrin atlas data and animal protein atlas data; using a Gaussian fitting algorithm to obtain several characteristic peaks and corresponding data pairs of the plant fibrin atlas data and the animal protein atlas data, respectively; training a conformational change evaluation decision function by performing feature analysis on the characteristic peaks in the two atlas data within different secondary structure ranges; and using a conformational change evaluation decision function to obtain an evaluation value for the conformational change of the plant fibrin molecule to be tested based on the content difference between the plant fibrin molecule to be tested and the animal protein sample it is intended to imitate under all secondary structures. The present application aims to solve the problem that the conformational evaluation of plant fibrin molecules is difficult to accurately perform due to the excessively large molecular weight and overly complex spectrum of artificial plant fibrin.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of protein molecule conformational change evaluation, and specifically to a method for evaluating plant fiber protein molecule conformational change. Background Art

[0002] Plant fiber protein molecular processing is a technology that processes plant protein into a fiber structure to simulate animal protein, thereby giving plant protein food the taste of animal protein. The animal protein content in existing meat products is mainly in lean meat, but the price of lean meat is facing cyclical price increases and instability, and lean meat loses water and shrinks during the application and maturation process, reducing the output rate. Accordingly, in order to meet the needs of corporate development and different consumers, this application processes plant fiber protein molecules to give them the taste of meat protein, aiming to provide a wider variety of protein sources and optimize the nutritional content and taste characteristics of food through technical means.

[0003] The properties of chemical substances come from their molecular conformation. As a biological macromolecule, the structural feature of protein that determines its physical properties (such as taste and oil absorption) at the molecular conformation level is its secondary structure. The secondary structure is formed by the polypeptide chains of the protein coiled together through hydrogen bonds. Proteins with similar secondary structures often have similar physical properties. Currently, the most sensitive method for detecting hydrogen bonds is Fourier infrared spectroscopy. When this method is applied to the judgment of protein secondary structure, it can determine the proportion of different types of secondary structures of proteins. The proportion of protein secondary structures can be used to compare the similarity of the secondary structures of two proteins. However, in the process of evaluating the molecular conformation of plant fiber proteins, due to the large molecular weight and complex structure of plant fiber proteins, the infrared spectrum corresponding to its secondary structure is too complex, making it difficult to extract features by traditional secondary structure infrared spectroscopy processing methods, and thus the problem of difficult to accurately evaluate the conformation of plant protein fiber molecules. Summary of the Invention

[0004] In order to solve the above technical problems, the present application provides a method for evaluating conformational changes of plant fiber protein molecules to solve the existing problems.

[0005] The present invention provides a method for evaluating conformational changes in plant fiber protein molecules using the following technical solutions:

[0006] One embodiment of the present application provides a method for evaluating conformational changes of plant fiber protein molecules, the method comprising the following steps:

[0007] S1, adding potassium bromide to the processed protein mixture, grinding, mixing and tableting, and scanning with an infrared spectrometer to obtain an FTIR spectrum of the plant fiber protein, from which the wavelength band of the amide I band is intercepted to obtain the plant fiber protein spectrum data; accordingly, the animal protein spectrum data is obtained;

[0008] S2, using Gaussian fitting algorithm to obtain several characteristic peaks of plant fiber protein map data and animal protein map data and their corresponding data pairs; by performing feature analysis on the characteristic peaks in the two map data within different secondary structure ranges, training the conformational change evaluation decision function, specifically:

[0009] A1, based on the distribution area and central wavenumber difference of any characteristic peak under any secondary structure, determine the characteristic peak ambiguity weight of any characteristic peak for any secondary structure;

[0010] A2, based on the difference between the first-order derivatives of the Gaussian function of any characteristic peak and the spectrum data, and the fluctuation of the characteristic peak ambiguity weight of any characteristic peak in all secondary structures, determines the characteristic peak confidence of any characteristic peak;

[0011] A3, based on the peak confidence, overlapping area ratio, and peak ambiguity weight of any two characteristic peaks between the two chromatographic data, determine the content difference of any secondary structure;

[0012] A4: construct an input data set based on the content differences of all secondary structures in any two atlas data sets under all combinations of species, as well as the manually evaluated evaluation values. This input data set is then trained to obtain a conformational change evaluation decision function.

[0013] S3, based on the content difference of all secondary structures between the plant fiber protein molecule to be tested and the animal protein sample it imitates, a conformational change evaluation decision function is used to obtain an evaluation value of the conformational change of the plant fiber protein molecule to be tested.

[0014] Preferably, the data pair is related information of the corresponding characteristic peak, and the related information of the corresponding characteristic peak contained in the data pair includes: the central wave number, half-height width, peak height and peak area of the characteristic peak.

[0015] Preferably, the method for determining the characteristic peak ambiguity weight includes:

[0016] Obtain the corresponding sequence of each secondary structure corresponding to the wavenumber range expanded under the preset step size;

[0017] Accumulate and sum the values corresponding to the wavelength of any characteristic peak in the corresponding sequence of any secondary structure to obtain the area of any characteristic peak in the wavelength band of any secondary structure;

[0018] Determining the conformity of the central wavenumber of any characteristic peak within any secondary structure band based on the central wavenumber difference between any characteristic peak and any secondary structure;

[0019] The normalized value of the product of the area and the central wave number conformity is used as the characteristic peak ambiguity weight.

[0020] Preferably, the method for determining the central wavenumber conformity includes: calculating the difference between the central wavenumber of the data pair of any characteristic peak and the central wavenumber of any secondary structure, and taking the normalized value of the inverse of the difference as the central wavenumber conformity of any characteristic peak in any secondary structure band.

[0021] Preferably, the method for determining the characteristic peak confidence includes:

[0022] The fitting confidence of any characteristic peak is determined based on the numerical difference between the first-order derivative of the Gaussian function of any characteristic peak and the first-order derivative of the plant fiber protein atlas data within a preset range;

[0023] Determine the inter-class confidence of any characteristic peak based on the fluctuation of the characteristic peak fuzziness weight of all secondary structures;

[0024] The ratio of the inter-class confidence to the fitting confidence is used as the characteristic peak confidence.

[0025] Preferably, the method for determining the fitting confidence includes:

[0026] Obtaining the wave number value of any characteristic peak within a preset range to form a range sequence of the any characteristic peak;

[0027] Calculating the first-order derivative of the Gaussian function of any characteristic peak, and normalizing the elements of the first-order derivative of the Gaussian function obtained by taking the range sequence as an independent variable to obtain a variation regularity sequence of any characteristic peak;

[0028] The first-order derivative of the plant fiber protein map data is obtained, and the elements of the first-order derivative of the plant fiber protein map data obtained by taking the range sequence as an independent variable are normalized to obtain a change rule sequence of the plant fiber protein map data;

[0029] The reciprocal of the difference between the two variation regularity sequences is taken as the fitting confidence of any characteristic peak.

[0030] Preferably, the method for determining the content difference includes:

[0031] Calculate the sum of the characteristic peak confidences of the two arbitrary characteristic peaks, and record it as the first characteristic peak coefficient;

[0032] Calculate the overlapping area ratio of any two characteristic peaks, and record it as the second characteristic peak coefficient;

[0033] Calculate the sum of the characteristic peak ambiguity weights of any two characteristic peaks under any secondary structure, and record it as the third characteristic peak coefficient;

[0034] The first, second and third characteristic peak coefficients are merged to obtain the characteristic peak difference between any two characteristic peaks under any secondary structure;

[0035] The cumulative sum of the characteristic peak differences between any two characteristic peaks under any secondary structure is recorded as the content difference.

[0036] Preferably, the calculation method of the overlapping area ratio is: taking the overlapping area of any two characteristic peaks as the numerator, taking the sum of the non-overlapping area and the overlapping area of any two characteristic peaks as the denominator, and taking the ratio of the numerator to the denominator as the overlapping area ratio.

[0037] Preferably, the method for calculating the characteristic peak difference includes: taking the product of the first, second and third characteristic peak coefficients as the characteristic peak difference of any two characteristic peaks under any secondary structure.

[0038] Preferably, the method for training the input data set to obtain the conformational change evaluation decision function is a support vector machine algorithm.

[0039] This application has at least the following beneficial effects:

[0040] This application aims to evaluate the ability of plant fiber proteins to imitate animal proteins, and proposes a method for evaluating the conformational changes of plant fiber protein molecules. The method solves the problem that the molecular weight of artificial plant fiber protein molecules is too large, which makes it difficult to extract features by traditional secondary structure infrared spectroscopy processing, and thus it is difficult to accurately evaluate the conformation of plant protein fiber molecules. The method includes: obtaining multiple characteristic peaks from the protein spectrum of the amide I band through Gaussian fitting; constructing characteristic peak ambiguity weights based on the position characteristics of the characteristic peaks and the bands corresponding to the four protein secondary structures, and corresponding the characteristic peaks to the four protein secondary structures, avoiding the traditional method based on technical analysis. The subjective errors caused by direct classification and division of characteristic peaks based on the experience of technical personnel are eliminated; by comparing the changing patterns of characteristic peaks and amide I band protein spectra, combined with the fluctuation of characteristic peak ambiguity weight of characteristic peaks, characteristic peak confidence is obtained, thus avoiding the adverse effects of pseudo peaks caused by noise on the final comparison data; finally, by comparing the characteristic peak areas of amide I band infrared spectra of the two proteins, combined with characteristic peak ambiguity weight and characteristic peak confidence, the difference in secondary structure content is obtained, and data training is performed on the difference in secondary structure content to construct a conformational change evaluation decision function for evaluating the imitation ability of plant fiber protein relative to animal protein. This application completes the comparison of the secondary structure similarity of the two proteins by comparing the infrared spectra of amide I band of the two proteins, solving the problem that the conformational evaluation of plant fiber protein molecules is difficult to accurately carry out due to the excessive molecular weight and overly complex spectrum of artificial plant fiber protein molecules. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] Figure 1 A flow chart of a method for evaluating conformational changes of plant fiber protein molecules provided in one embodiment of the present application;

[0043] Figure 2 A schematic diagram of the bands corresponding to the four secondary structures of the amide I band in a protein provided in one embodiment of the present application;

[0044] Figure 3 A flowchart of the steps for training a conformational change evaluation decision function provided in one embodiment of the present application;

[0045] Figure 4 A flowchart of a method for determining the confidence level of a characteristic peak provided in one embodiment of the present application;

[0046] Figure 5 A flowchart of the steps of a method for determining content differences provided in one embodiment of the present application. DETAILED DESCRIPTION

[0047] To further illustrate the technical means and effectiveness of this application to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the method for evaluating conformational changes in plant fiber protein molecules proposed in this application, including its specific implementation, structure, features, and effectiveness. In the following description, references to different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0048] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0049] The specific scheme of the method for evaluating conformational changes of plant fiber protein molecules provided by the present application is described in detail below with reference to the accompanying drawings.

[0050] An embodiment of the present application provides a method for evaluating conformational changes of plant fiber protein molecules.

[0051] Specifically, a method for evaluating conformational changes of plant fibrous protein molecules is provided below. Figure 1 , the method comprises the following steps:

[0052] The first step is to obtain plant fiber protein map data and animal protein map data.

[0053] In practical applications, plant fibrous proteins are typically derived from plant materials such as soybeans, peas, wheat, soy beans, and tofu. The processing involves protein extraction, hydrolysis, polymerization, or fiberization to achieve the desired food functions and nutritional value.

[0054] In this application, a fibrous texture of meaty texture is achieved by processing plant protein into a stringy structure. The stringy structure is produced by using a protein mixture with a certain ratio, using protein isolate and gluten as the main raw materials, utilizing the hydrophilicity of soy protein and wheat protein, and adding corn starch to fully expand the protein. After high-temperature expansion and extrusion, the fiber texture is elastic and non-sticky.

[0055] In the present application, the ratio of plant raw materials in the processing of the above-mentioned plant protein drawing structure is: 40-50 parts of soy protein isolate, 20-30 parts of gluten, 15-25 parts of corn starch, and 10-20 parts of low-temperature soybean meal.

[0056] It should be noted that the purpose of using puffing extrusion equipment in this application to perform high-temperature expansion and extrusion of the protein mixture is to utilize the characteristics of soy protein and wheat protein being hydrophilic proteins, so that the proteins can be easily fully mixed during the puffing and extrusion process, and the maximum degree of protein modification can be achieved during the heating process. At the same time, by using a screw to extrude it, the protein structure of soybeans and wheat is modified from a spherical shape to a fiber drawing structure, so that it achieves a fine and tough protein fiber structure, thereby changing the taste of the meat.

[0057] In addition, in addition to the elastic meat fiber texture, the overall chewy feel and meaty texture of plant protein lean meat need to be achieved by mixing with the emulsified slurry of soy protein isolate.

[0058] Based on this, the present application makes full use of the respective adhesiveness and protein emulsification characteristics of soy protein isolate, gluten and distarch phosphate, and mixes the processed protein mixture with the emulsified slurry of soy protein isolate, thereby improving the overall chewiness and meatiness of the processed protein mixture.

[0059] Among them, the main proportions of the emulsified slurry of soy protein isolate are: 50-60 parts of water, 15-20 parts of soy protein isolate, 5-10 parts of vegetable oil, 5-10 parts of gluten, 5-10 parts of distarch phosphate, and 3-5 parts of corn starch.

[0060] It is noteworthy that the plant protein lean meat of the present invention is prepared by mixing and maturing a plant protein filamentous structure and an emulsified material mainly composed of soy protein isolate. The key lies in the processing of the plant protein filamentous structure and the ratio of the emulsified material mainly composed of soy protein isolate. Moreover, the plant protein lean meat can be used as a substitute for lean meat in industrial production channels such as meat products, sauces, fillings, and toppings, and has a very wide range of application channels. The main features of the product are:

[0061] a. After high-temperature steaming, it will not mix with other materials in the product and can maintain a clear sense of meat fiber; after frying or boiling, it can also maintain the granular form containing meat fiber;

[0062] b. It has excellent oil and water retention properties, which can lock the oil and water in the mixture of the product, making the product feel moisturizing after chewing;

[0063] c. Soy protein isolate can make plant protein lean meat maintain a certain elasticity and taste, and increase the chewiness of lean meat;

[0064] d. Every 100 grams of plant protein lean meat contains 27.8g of protein, 1.5g of fat, and 3.5g of carbohydrates. These nutrients are much healthier than animal protein lean meat, especially the protein content is significantly higher than animal protein lean meat;

[0065] e. The plant protein fiber structure in the product is not mushy, greasy, sticky or loose after boiling or frying, and can still maintain the elasticity of the fiber;

[0066] f. When the product is added to various hams, luncheon meats, and meatballs, it will not shrink like lean meat, but will expand after absorbing water and oil, thereby increasing the output rate of the product and improving the cost performance of the product.

[0067] Next, the present application completes the evaluation of the ability of plant fiber protein molecules to imitate animal protein molecules by comparing the secondary structure spectrum parts in the infrared spectra of plant fiber protein molecules and animal protein molecules.

[0068] The specific method for obtaining plant fiber protein atlas data and animal protein atlas data is as follows:

[0069] Potassium bromide is added to the processed protein mixture for grinding, mixing and tableting, and an infrared spectrometer is used for scanning to obtain an FT IR spectrum of the plant fiber protein, from which the wavelength band of the amide I band is intercepted to obtain the plant fiber protein spectrum data.

[0070] Preferably, in one embodiment of the present application, 5 mg of the processed protein mixture is weighed, 0.2 g of potassium bromide is added, and the mixture is ground and mixed evenly. The mixture is then compressed using an infrared tablet press and the temperature is balanced in a constant temperature box for 5 minutes. The compressed tablet is then scanned using an infrared spectrometer with a scanning band of 400 to 4000 cm -1 , the scanning band interval is 1cm -1 , thereby obtaining the FT IR spectrum of plant fibrous protein, the vertical axis of the FT IR spectrum of plant fibrous protein is absorbance, and the horizontal axis is wavelength.

[0071] In this embodiment, the weight of the processed protein mixture and potassium bromide, as well as the time for the thermostat to equilibrate the temperature, can be set by the implementer according to actual conditions.

[0072] According to experience, the FT IR spectrum of protein secondary structure is located between 1615 and 1700 cm -1 The wavelength band is called amide I band; therefore, the FT IR spectrum of plant fiber protein is intercepted at 1615~1700cm -1 The spectrum of the band is used to obtain the plant fiber protein spectrum data. The storage data of the plant fiber spectrum data is a vector with a length of N=85, where the nth element represents the wavelength of (1614+n)cm -1 The absorbance of the sample is .

[0073] It should be noted that the position and shape of the amide I band can provide information about the protein's molecular structure, such as the composition and changes in its secondary structure. Infrared spectroscopy allows for preliminary assessment and comparison of protein structures, which is particularly useful when studying protein folding states, conformational changes, or interactions with other molecules.

[0074] Potassium bromide is added to the freeze-dried animal protein powder with a meaty texture similar to the processed protein mixture, and the powder is ground, mixed and pressed into tablets. The powder is scanned with an infrared spectrometer to obtain the FT IR spectrum of the animal protein, from which the band of the amide I band is intercepted to obtain the animal protein spectrum data.

[0075] In one embodiment of the present application, the schematic diagram of the bands corresponding to the four secondary structures of the amide I band in the protein is shown in the attached figure. Figure 2 shown.

[0076] In the attached Figure 2 In the figure, A represents α-helix structure, B represents β-sheet structure, C represents random coil structure, and D represents β-turn structure. The horizontal axis represents the wave number in cm. -1 , among which, 1615~1700cm -1 The band is amide I band.

[0077] In the second step, the conformational change evaluation decision function is trained by performing feature analysis on the characteristic peaks in the two atlas data within different secondary structure ranges.

[0078] Traditional infrared spectrum processing of the characteristic peaks of the amide I band uses Gaussian fitting and other techniques to decompose this portion of the infrared spectrum into multiple characteristic peaks, thereby extracting the features of the aliased infrared spectrum. Although this method produces too many characteristic peaks for large and complex protein molecules, making it difficult to map them to the four protein secondary structures, it can still eliminate interference caused by the overlapping of characteristic peaks. Therefore, this example applies the same processing to the plant fibrous protein spectrum data to obtain a dataset of secondary structure characteristic peaks for the plant fibrous protein spectrum data.

[0079] In one embodiment of the present application, the process of acquiring the secondary structure characteristic peak dataset of plant fibrous protein atlas data is as follows:

[0080] The plant fibrous protein atlas data was used as input, and the Gaussian fitting algorithm was used. The fitting stop condition was the fitting correlation coefficient R 2≥0.98, the output is K characteristic peaks, each characteristic peak corresponds to a data pair, where the kth data pair is (m, w, h, s), each data pair represents the relevant information of the corresponding characteristic peak, m, w, h, s represent the central wave number m, half-height width w, peak height h, and peak area s of the kth characteristic peak, respectively; the number of data pairs K is determined by the calculation results of the algorithm, and the K data pairs are combined to form a data set of secondary structure characteristic peaks of plant fibrous protein. The Gaussian fitting algorithm is a well-known processing technique commonly used for protein data and will not be described in detail in this application.

[0081] The same processing is performed on the animal protein atlas data to obtain the secondary structure characteristic peak data set of the animal protein atlas data.

[0082] In one embodiment of the present application, the characteristic peaks in the two atlas data are analyzed in different secondary structure ranges to train the conformational change evaluation decision function. Figure 3 As shown, specifically:

[0083] A1, based on the distribution area and central wavenumber difference of any characteristic peak under any secondary structure, determine the characteristic peak ambiguity weight of any characteristic peak for any secondary structure.

[0084] Since the molecular weight of plant protein fiber is too large, the number of characteristic peaks in its secondary structure characteristic peak data set is usually much greater than 4, resulting in a relatively uniform distribution of these characteristic peaks on the amide I band, and sometimes spanning the band range of two or more secondary structures. At this time, it is necessary to perform subordinate division on these characteristic peaks to determine the secondary structure content of the plant protein fiber.

[0085] Based on this, this embodiment uses fuzzy mathematics to perform subordinate division on the characteristic peaks, that is, the characteristic peaks are divided into different parts according to the secondary structure band range, where the rth part corresponds to the rth secondary structure. The larger the area of the rth part occupies in the entire characteristic peak, the more likely the secondary structure represented by the characteristic peak is the rth secondary structure, that is, the characteristic peak fuzziness weight for the rth secondary structure is also greater.

[0086] Specifically, the method for determining the characteristic peak ambiguity weight includes:

[0087] Obtain a corresponding sequence of the wavenumber range corresponding to each secondary structure expanded at a preset step size; accumulate and sum the values corresponding to the wavelength of any characteristic peak in the corresponding sequence of any secondary structure to obtain the area of any characteristic peak in any secondary structure band; based on the central wavenumber difference between any characteristic peak and any secondary structure, determine the central wavenumber conformity of any characteristic peak in any secondary structure band; and use the normalized value of the product of the area and the central wavenumber conformity as the characteristic peak ambiguity weight.

[0088] In this embodiment, a specific analysis is performed by taking the kth characteristic peak and the rth secondary structure as an example:

[0089] For the kth characteristic peak in the secondary structure characteristic peak data set of plant fibrous protein, since the present embodiment adopts the Gaussian fitting algorithm to obtain the characteristic peak, the kth characteristic peak corresponds to a Gaussian function f(x) k , that is, when the wave number is λ, the function value of the kth characteristic peak is f(λ) k .

[0090] At the same time, the wave number range corresponding to the rth secondary structure can be expanded to form a corresponding sequence with a step size of 1. In other embodiments of the present application, the step size can be set according to actual conditions. Among them, the wave number range of the α-helix structure can be represented by the sequence (1646, 1647...1664), and the corresponding sequence of the rth secondary structure is Num r .

[0091] Furthermore, in this embodiment, the corresponding sequence Num of the rth secondary structure r As the independent variable, the Gaussian function f(x) of the kth characteristic peak is k As a function, obtain the corresponding sequence of the rth secondary structure under the Gaussian function f(x) k The element values in the sequence are then added up to obtain the area DS of the kth characteristic peak in the rth secondary structure band. k,r .

[0092] In addition, the closer the central wavenumber m in the data pair of the kth characteristic peak is to the central wavenumber of the rth secondary structure, the more likely the kth characteristic peak represents the rth secondary structure, and therefore the characteristic peak ambiguity weight of the kth characteristic peak to the rth secondary structure should also be greater.

[0093] Therefore, this embodiment calculates the absolute value of the difference between the central wavenumber of the data pair of the kth characteristic peak and the central wavenumber of the rth secondary structure, then calculates the inverse of the absolute value of the difference, and finally performs normalization to obtain the central wavenumber conformity Ci of the kth characteristic peak in the rth secondary structure band. k,r It should be understood that the larger the consistency value, the more likely the kth characteristic peak is to represent the rth secondary structure.

[0094] It should be noted that when calculating the reciprocal of the absolute value of the difference, an extremely small number is added to the absolute value of the difference. In this embodiment, the value is 0.1 to prevent the denominator from being 0.

[0095] Finally, calculate the fuzziness weight of the kth characteristic peak to the rth secondary structure characteristic peak, and convert the area DS of the kth characteristic peak in the rth secondary structure band intok,r The degree of conformity Ci with the central wave number of the kth characteristic peak in the rth secondary structure band k,r The normalized value of the product of is used as the characteristic peak fuzziness weight of the kth characteristic peak to the rth secondary structure.

[0096] It should be understood that the greater the fuzzy weight of the kth characteristic peak for the rth secondary structure, the more likely the kth characteristic peak is the characteristic peak of the rth secondary structure. Using fuzzy characteristic weights as a method for determining the expression of secondary structure by characteristic peaks avoids the subjective errors caused by traditional methods that directly classify characteristic peaks based on the technician's experience.

[0097] At this point, the characteristic peak fuzziness weight of each secondary structure has been calculated for each characteristic peak in the secondary structure characteristic peak data set of plant fiber protein atlas data. The same calculation can be performed on each characteristic peak in the secondary structure characteristic peak data set of animal protein atlas data to obtain the characteristic peak fuzziness weight of each characteristic peak.

[0098] A2, based on the difference in the first-order derivative between the Gaussian function of any characteristic peak and the spectrum data, and the fluctuation of the characteristic peak ambiguity weight of any characteristic peak in all secondary structures, determines the characteristic peak confidence of any characteristic peak.

[0099] Since this embodiment characterizes the difference in protein taste by comparing the difference in secondary structure content of two large molecular proteins, when comparing the characteristic peaks of the secondary structures of the two proteins, the more the characteristic peak conforms to the changing trend of the original data, the better the fitting effect of the characteristic peak and the higher the confidence; the greater the fluctuation of the characteristic peak fuzziness weight of the characteristic peak, the more the characteristic peak represents a single and clear secondary structure, the more clearly it can represent a certain secondary structure, and the higher the confidence; the higher the confidence of the characteristic peak, the clearer its expression of the secondary structure content characteristics and the better the expression effect, and when comparing the characteristic peaks, the greater the calculation weight of the characteristic peak should be.

[0100] Due to the characteristic peak separation method of infrared spectra, such as Gaussian fitting, when decomposing the infrared spectrum of the amide I band, noise interference may occur and produce false peaks. Such peaks are called pseudo-peaks. When performing infrared spectroscopy analysis of the secondary structure of large molecular weight proteins, due to the excessive number of characteristic peaks, the pseudo-peaks that appear are difficult to distinguish. The characteristic peak confidence can eliminate the adverse effects of pseudo-peaks on the final comparative data.

[0101] Specifically, the method for determining the characteristic peak confidence includes:

[0102] The fitting confidence of any characteristic peak is determined based on the numerical difference between the first-order derivative of the Gaussian function of any characteristic peak and the first-order derivative of the plant fiber protein atlas data within a preset range; the inter-class confidence of any characteristic peak is determined based on the fluctuation of the characteristic peak fuzziness weight of any characteristic peak in all secondary structures; the ratio of the inter-class confidence to the fitting confidence is used as the characteristic peak confidence.

[0103] In one embodiment of the present application, the method for determining the confidence level of a characteristic peak is shown in the flowchart of the attached figure. Figure 4 shown.

[0104] The method for determining the fitting confidence level includes:

[0105] Obtain the wavenumber value of any characteristic peak within a preset range to form a range sequence of the any characteristic peak; calculate the first-order derivative of the Gaussian function of the any characteristic peak, and normalize the elements in the first-order derivative of the Gaussian function obtained by taking the range sequence as the independent variable to obtain a change law sequence of the any characteristic peak; calculate the first-order derivative of the plant fiber protein atlas data, and normalize the elements in the first-order derivative of the plant fiber protein atlas data obtained by taking the range sequence as the independent variable to obtain a change law sequence of the plant fiber protein atlas data; and use the inverse of the difference between the two change law sequences as the fitting confidence of the any characteristic peak.

[0106] In one embodiment of the present application, taking the kth characteristic peak as an example, the fitting confidence of the kth characteristic peak is analyzed, specifically:

[0107] For the kth characteristic peak in the secondary structure characteristic peak data set of plant fibrous protein, obtain its Gaussian function f(x) k The central wave number m k According to the 3sigma principle, the main range of the kth characteristic peak is mainly in the interval [m k -3×σ k ,m k +3×σ k ], obtain the wave number value in this interval to form the range sequence Int of the kth characteristic peak k .

[0108] At the same time, for the Gaussian function f(x) k Find the first derivative and use the range sequence Int k Substitute the first-order derivative for the independent variable to obtain the first-order derivative sequence of the k-th characteristic peak. Normalize the first-order derivative sequence to obtain the change law sequence of the k-th characteristic peak. This sequence represents the numerical change law of the k-th characteristic peak, where the positive or negative value represents the rise or fall of the k-th characteristic peak, and the size of the value represents the severity of the change.

[0109] The plant fiber protein map data is a vector. The first-order derivative vector calculation algorithm is used to calculate the plant fiber protein map data using the vector as input, and the output is the first-order derivative vector of the plant fiber protein map data. The first-order derivative vector calculation method is a commonly used technology in the field of data processing.

[0110] Furthermore, according to the range sequence Int of the kth characteristic peak k The first-order derivative vector is intercepted within the wavenumber range to obtain the first-order derivative sequence of the plant fibrin map data. The first-order derivative sequence is normalized to obtain the change law sequence of the plant fibrin map data. The sequence represents the numerical change law of the plant fibrin map data within the k-th characteristic peak range, where the positive or negative value represents the rise or fall of the plant fibrin map, and the size of the value represents the severity of the change.

[0111] Finally, the reciprocal of the cumulative sum of the absolute values of the differences between the k-th characteristic peak and the corresponding elements of the change pattern sequence of the plant fiber protein spectrum is used as the fitting confidence of the k-th characteristic peak.

[0112] In another embodiment of the present application, the reciprocal of the DTW distance between the kth characteristic peak and the variation regularity sequence of the plant fiber protein pattern is used as the fitting confidence of the kth characteristic peak. The DTW distance is a well-known technology and will not be described in detail.

[0113] In other embodiments of the present application, the L1 norm of the sequence composed of the differences between the kth characteristic peak and the variation regularity sequence of the plant fiber protein spectrum is calculated, and the inverse of the L1 norm is calculated and recorded as the fitting confidence of the kth characteristic peak.

[0114] It should be understood that, the greater the fitting confidence of the k-th characteristic peak, the better the fitting effect of the k-th characteristic peak to the original data, and the higher the confidence.

[0115] It should be noted that before calculating the reciprocal, a very small number is added to the L1 norm, which is 0.1 in this embodiment to prevent the denominator from being 0.

[0116] The method for determining the inter-class confidence includes:

[0117] The fluctuation of the characteristic peak fuzziness weight of any characteristic peak in all secondary structures is used as the inter-class confidence of any characteristic peak.

[0118] In one embodiment of the present application, taking the kth characteristic peak as an example, the inter-class confidence of the kth characteristic peak is analyzed, and the standard deviation of the characteristic peak fuzziness weights of the kth characteristic peak in all secondary structures is used as the inter-class confidence of the kth characteristic peak. In other embodiments of the present application, variance, information entropy, etc. can also be used as fluctuation conditions for analysis.

[0119] It should be understood that the greater the inter-class confidence of the k-th characteristic peak, the more the k-th characteristic peak represents a single and clear secondary structure, the more clearly it can represent a certain secondary structure, and the higher the confidence.

[0120] Finally, the inter-class confidence is divided by the fitting confidence to obtain the peak confidence β of the kth peak. k The larger the value, the better the kth characteristic peak expresses the secondary structure and the higher the confidence. The confidence level can be used to screen characteristic peaks that express the secondary structure content characteristics well, thereby avoiding the adverse effects of pseudo peaks caused by noise on the final comparison data.

[0121] The above characteristic peak confidence is calculated for each characteristic peak in the secondary structure characteristic peak data set of plant fiber protein atlas data. The same calculation can be performed on each characteristic peak in the secondary structure characteristic peak data set of animal protein atlas data to obtain the characteristic peak confidence of each characteristic peak.

[0122] A3, based on the characteristic peak confidence, overlapping area ratio, and characteristic peak ambiguity weight of any two characteristic peaks between the two chromatographic data, determines the content difference of any secondary structure.

[0123] When comparing the conformations of two proteins, the greater the difference in the characteristic peaks of the amide I band in the infrared spectra of the two proteins, the greater the difference in their secondary structures and the greater the difference in their taste. Therefore, this example compares the differences in the characteristic peaks of the amide I band of plant fiber protein and animal protein.

[0124] This example uses area comparison as one method for comparing characteristic peak differences. A numerical relationship is established between the overlapping and non-overlapping areas. The greater the percentage of overlapping area to the total area, the smaller the difference; the greater the percentage of non-overlapping area to the total area, the greater the difference. The peak confidence levels between characteristic peaks, as well as the peak ambiguity weights for different secondary structures, are combined to analyze the differences in the content of any secondary structure.

[0125] Wherein, the method for determining the content difference includes:

[0126] Calculate the sum of the characteristic peak confidences of any two characteristic peaks, which is recorded as the first characteristic peak coefficient; calculate the overlapping area ratio of the any two characteristic peaks, which is recorded as the second characteristic peak coefficient; calculate the sum of the characteristic peak fuzziness weights of the any two characteristic peaks under any secondary structure, which is recorded as the third characteristic peak coefficient; fuse the first, second and third characteristic peak coefficients to obtain the characteristic peak difference of the any two characteristic peaks under any secondary structure; and record the cumulative sum of the characteristic peak differences of all any two characteristic peaks under any secondary structure as the content difference.

[0127] It can be understood that fusion can be divided into forward fusion and reverse fusion. This embodiment adopts the forward fusion method. Forward fusion is a fusion method such as addition and multiplication between data. The specific forward fusion method is determined by the implementer according to the actual situation. The application does not impose any special restrictions.

[0128] In one embodiment of the present application, the flow chart of the method for determining the content difference is shown in the attached figure. Figure 5 shown.

[0129] In this embodiment, the k1th characteristic peak of the secondary structure characteristic peak data set of plant fibrous protein and the k2th characteristic peak of the secondary structure characteristic peak data set of animal protein are analyzed as examples, specifically:

[0130] First, the characteristic peak confidences of the two characteristic peaks are summed and recorded as the first characteristic peak coefficients of the two characteristic peaks. The larger the value, the better the expression effect of the comparison result between the two on the secondary structure and the higher the confidence.

[0131] Secondly, the overlapping area between the two characteristic peaks is recorded as S1, and the non-overlapping area between the two characteristic peaks is recorded as S2; the overlapping area ratio between the k1th characteristic peak and the k2th characteristic peak is: S1 is used as the numerator, the sum of S1 and S2 is used as the denominator, and the ratio of the numerator to the denominator is used as the overlapping area ratio between the k1th characteristic peak and the k2th characteristic peak, which is recorded as the second characteristic peak coefficient of the two characteristic peaks.

[0132] Then, the rth characteristic peak fuzziness weights of the two characteristic peaks are summed and recorded as the third characteristic peak coefficient of the two characteristic peaks. The larger the value, the more the comparison result of the k1 and k2 characteristic peaks represents the difference of the rth secondary structure.

[0133] Finally, the product of the first, second, and third characteristic peak coefficients was calculated as the characteristic peak difference between the two characteristic peaks under the rth secondary structure. The larger the value, the more significant the difference between the two proteins in the rth secondary structure based on the comparison results of the k1 and k2 characteristic peaks.

[0134] The characteristic peak differences between all characteristic peaks under the r-th secondary structure are summed to obtain the content difference ΔB of the r-th secondary structure of the two proteins. r .

[0135] In another embodiment of the present application, the sum of the first, second and third characteristic peak coefficients is calculated as the characteristic peak difference between the two characteristic peaks under the rth secondary structure.

[0136] In other embodiments of the present application, the sum of the first and third characteristic peak coefficients is calculated and then multiplied by the second characteristic peak coefficient to obtain the characteristic peak difference between the two characteristic peaks under the rth secondary structure.

[0137] To facilitate data processing, the R secondary structure content differences were arranged into a sequence, recorded as the content difference feature sequence. The larger the r-th value in the content difference feature sequence, the more obvious the content difference of the r-th secondary structure between the two proteins was.

[0138] A4, construct an input data set based on the content differences of all secondary structures in any two atlas data under all combinations of species and the manually evaluated evaluation values, and train the input data set to obtain a conformational change evaluation decision function.

[0139] Based on the characteristic sequences of the differences in the secondary structure content between plant fiber proteins and animal proteins, the differences in the secondary structure content of the two proteins can be determined, and then the ability of plant fiber proteins to imitate the taste of animal proteins can be evaluated; during the evaluation, the differences in different secondary structures have different effects on the taste imitation ability. This embodiment uses the support vector machine method to process the characteristic sequences of the differences in secondary structure content to complete the evaluation of the changes in the molecular conformation of plant fiber proteins; optionally, neural networks and PCA algorithms can also be used as alternative evaluation methods.

[0140] Among them, the specific method of processing the secondary structure content difference feature sequence by support vector machine and training the conformational change evaluation decision function is as follows:

[0141] First, different plant fiber protein map data and different animal protein map data are combined in pairs to obtain multiple combinations; for one of the combinations, a characteristic sequence of secondary structure content differences is obtained according to the above-mentioned calculation method of this embodiment; the artificial meat products of the plant fiber protein of this combination and the meat products corresponding to the animal protein are evaluated by tasters, and the evaluation results of the tasters are collected. The evaluation results are sorted and classified into 4 types of evaluation, 1 represents that the two proteins have similar tastes, 2 represents that the two proteins have relatively similar tastes, 3 represents that the two proteins have different tastes, and 4 represents that the two proteins have very different tastes.

[0142] Finally, each combination can obtain a secondary structure content difference characteristic sequence as the input part of a set of data; each combination can also obtain an evaluation value, with evaluation values of 1, 2, 3, and 4 as the data evaluation of a set of data; the input parts composed of all types and the corresponding evaluation values constitute the input data set.

[0143] Taking the input data set as input, the support vector machine algorithm is used to output a decision function, which is recorded as the conformational change evaluation decision function. The support vector machine algorithm is a commonly used technology in the field of data processing.

[0144] The third step is to use the conformational change evaluation decision function to obtain the evaluation value of the conformational change of the plant fiber protein molecule to be tested based on the content difference in all secondary structures between the plant fiber protein molecule to be tested and the animal protein sample it imitates.

[0145] During the detection process, for the artificial meat plant fiber protein molecule to be detected, its protein samples and the target imitation animal meat are obtained, and then a secondary structure content difference characteristic sequence is obtained according to the above method of this embodiment. The secondary structure content difference characteristic sequence is used as the input of the conformational change evaluation decision function, and an evaluation value is output. According to the evaluation values 1, 2, 3, and 4, the taste evaluation of the artificial meat plant fiber protein molecule is obtained accordingly: the two proteins have similar tastes, the two proteins have relatively similar tastes, the two proteins have different tastes, and the two proteins have very different tastes.

[0146] Since this embodiment uses fuzzy mathematics to extract the difference in characteristic peaks of the amide I bands of the infrared spectra of two proteins, it solves the problem of difficult to accurately evaluate the conformation of plant fiber protein molecules due to the complex spectra of large molecular weight proteins.

[0147] The various embodiments in this application are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0148] It should be noted that, unless otherwise specified and limited, terms such as "include", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such article or device. In the absence of further restrictions, the phrase "including a ..." defines an element, does not exclude the presence of other identical elements in the article or device including the element. In addition, the term "and\or" used herein includes any and all combinations of one or more related listed items.

[0149] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not invented herein.

[0150] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. A method for evaluating conformational changes of plant fiber protein molecules, characterized in that: The method comprises the following steps: S1, adding potassium bromide to the processed protein mixture for grinding, mixing and tableting, and scanning with an infrared spectrometer to obtain an FTIR spectrum of the plant fiber protein, from which the wavelength band of the amide I band is intercepted to obtain the plant fiber protein spectrum data; accordingly, the animal protein spectrum data is obtained; S2, using Gaussian fitting algorithm to obtain several characteristic peaks of plant fiber protein map data and animal protein map data and their corresponding data pairs; by performing feature analysis on the characteristic peaks in the two map data within different secondary structure ranges, training the conformational change evaluation decision function, specifically: A1, based on the distribution area and central wavenumber difference of any characteristic peak under any secondary structure, determine the characteristic peak ambiguity weight of any characteristic peak for any secondary structure; A2, determining the characteristic peak confidence of any characteristic peak based on the difference between the first-order derivatives of the Gaussian function of any characteristic peak and the spectrum data, and the fluctuation of the characteristic peak ambiguity weight of any characteristic peak in all secondary structures; the method for determining the characteristic peak confidence includes: The fitting confidence of any characteristic peak is determined based on the numerical difference between the first-order derivative of the Gaussian function of any characteristic peak and the first-order derivative of the plant fiber protein atlas data within a preset range; Determine the inter-class confidence of any characteristic peak based on the fluctuation of the characteristic peak fuzziness weight of all secondary structures; The ratio of the inter-class confidence to the fitting confidence is used as the characteristic peak confidence; A3, based on the characteristic peak confidence, overlapping area ratio, and characteristic peak ambiguity weight of any two characteristic peaks between the two atlas data, determine the content difference of any secondary structure; the method for determining the content difference includes: Calculate the sum of the characteristic peak confidences of the two arbitrary characteristic peaks, and record it as the first characteristic peak coefficient; Calculate the overlapping area ratio of any two characteristic peaks, and record it as the second characteristic peak coefficient; Calculate the sum of the characteristic peak ambiguity weights of any two characteristic peaks under any secondary structure, and record it as the third characteristic peak coefficient; The first, second and third characteristic peak coefficients are merged to obtain the characteristic peak difference between any two characteristic peaks under any secondary structure; The cumulative sum of the characteristic peak differences between any two characteristic peaks under any secondary structure is recorded as the content difference; A4: construct an input data set based on the content differences of all secondary structures in any two atlas data sets under all combinations of species, as well as the manually evaluated evaluation values. This input data set is then trained to obtain a conformational change evaluation decision function. S3, based on the content difference of all secondary structures between the plant fiber protein molecule to be tested and the animal protein sample it imitates, a conformational change evaluation decision function is used to obtain an evaluation value of the conformational change of the plant fiber protein molecule to be tested.

2. The method for evaluating conformational changes of plant fiber protein molecules according to claim 1, wherein: The data pair is related information of the corresponding characteristic peak, and the related information of the corresponding characteristic peak contained in the data pair includes: the central wave number, half-height width, peak height and peak area of the characteristic peak.

3. The method for evaluating conformational changes of plant fiber protein molecules according to claim 2, wherein: The method for determining the characteristic peak ambiguity weight includes: Obtain the corresponding sequence of each secondary structure corresponding to the wavenumber range expanded under the preset step size; Accumulate and sum the values corresponding to the wavelength of any characteristic peak in the corresponding sequence of any secondary structure to obtain the area of any characteristic peak in the wavelength band of any secondary structure; Determining the conformity of the central wavenumber of any characteristic peak within any secondary structure band based on the central wavenumber difference between any characteristic peak and any secondary structure; The normalized value of the product of the area and the central wave number conformity is used as the characteristic peak ambiguity weight.

4. The method for evaluating conformational changes of plant fiber protein molecules according to claim 3, wherein: The method for determining the central wavenumber conformity includes: calculating the difference between the central wavenumber of the data pair of any characteristic peak and the central wavenumber of any secondary structure, and taking the normalized value of the reciprocal of the difference as the central wavenumber conformity of the any characteristic peak in any secondary structure band.

5. The method for evaluating conformational changes of plant fiber protein molecules according to claim 1, wherein: The method for determining the fitting confidence includes: Obtaining the wave number value of any characteristic peak within a preset range to form a range sequence of the any characteristic peak; Calculating the first-order derivative of the Gaussian function of any characteristic peak, and normalizing the elements of the first-order derivative of the Gaussian function obtained by taking the range sequence as an independent variable to obtain a variation regularity sequence of any characteristic peak; The first-order derivative of the plant fiber protein map data is obtained, and the elements of the first-order derivative of the plant fiber protein map data obtained by taking the range sequence as an independent variable are normalized to obtain a change rule sequence of the plant fiber protein map data; The reciprocal of the difference between the two variation regularity sequences is taken as the fitting confidence of any characteristic peak.

6. The method for evaluating conformational changes of plant fiber protein molecules according to claim 1, wherein: The calculation method of the overlapping area ratio is: taking the overlapping area of the arbitrary two characteristic peaks as the numerator, taking the sum of the non-overlapping area and the overlapping area of the arbitrary two characteristic peaks as the denominator, and taking the ratio of the numerator to the denominator as the overlapping area ratio.

7. The method for evaluating conformational changes of plant fiber protein molecules according to claim 1, wherein: The method for calculating the characteristic peak difference includes: taking the product of the first, second and third characteristic peak coefficients as the characteristic peak difference of any two characteristic peaks under any secondary structure.

8. The method for evaluating conformational changes of plant fiber protein molecules according to claim 1, wherein: The method for training the input data set to obtain the conformational change evaluation decision function is the support vector machine algorithm.

Citation Information

Patent Citations

  • Raman nondestructive inspection method for quality of frozen pacific white shrimps

    CN103424394A