Methods for predicting the values of material properties of interest

Through principal component analysis and singular value decomposition, spectral outliers are identified and removed to generate prediction functions, which solves the problem that spectral outliers affect calibration function in the prior art, and improves the prediction accuracy of near-infrared spectroscopy under non-laboratory conditions.

CN114599957BActive Publication Date: 2025-08-12EVONIK OPERATIONS GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080072592.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-17
Filing Date
2020-10-05
Publication Date
2025-08-12
Estimated Expiration
2040-10-05

AI Technical Summary

Technical Problem

When the prior art predicts material attribute values through near-infrared spectroscopy, spectral outliers are not effectively identified and processed, resulting in inaccurate calibration functions.

Method used

The principal component analysis and singular value decomposition method are used to calculate the score index and distance measurement threshold of the spectrum, and spectral outliers are automatically identified and removed, and a prediction function is generated to predict the value of the attribute value of the material.

Benefits of technology

Reliable identification and removal of spectral outliers when creating calibration functions is achieved, improving the prediction accuracy and robustness of near-infrared spectroscopy under non-laboratory conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114599957B_ABST
    Figure CN114599957B_ABST
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method for predicting the value of a property of interest in a sample studied by infrared spectroscopy. The method aims to generate a calibration function. To this end, a set of calibration samples is selected, from which outliers are identified and removed from the set of calibration samples. The outliers are determined using principal component analysis and singular value decomposition. A threshold value for separating the outliers from the remaining samples is calculated based on a predetermined formula. The threshold value can also be set dynamically by increasing it stepwise, which is preferred for spectroscopic devices that are not operated under laboratory conditions. #imgabs0##imgabs1#
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a computer-implemented method for predicting the value of a property of interest of a material and to an apparatus for such a method, said apparatus comprising a processing unit adapted to perform said method. Background Art

[0002] Near-infrared (NIR) spectroscopy is a useful tool for predicting the values of interesting properties of materials, particularly when located far from analytical laboratories where qualitative and / or quantitative analyses are typically performed. Specifically, agricultural products, such as feed ingredients and / or feeds, can be analyzed for the presence and concentration of specific ingredients before and after processing stages such as roasting and pressure treatment, during or after storage, and after preparation of feed containing these ingredients. However, sufficient calibration is required to provide accurate and reliable predictions using NIR spectroscopy.

[0003] The calibration function is typically generated by multivariate analysis. This allows for proper consideration of correlations to estimate the composition of complex mixtures and to compensate for interference from background signals. Predicting the value of a property of interest using NIR spectroscopy and a calibration function is a two-step process. In the first step, a calibration model is constructed using data sets obtained by indirect measurements such as light signal intensity and direct measurements such as analyte concentration in several scenarios across a variety of analytes and tool conditions. The general formula for the relationship between direct (e.g., analyte concentration) and indirect (e.g., light signal intensity) measurements is y = f(x1, x2, ..., x n ), where y is the expected attribute value of interest to be predicted (e.g., analyte concentration), f is a certain function (model), and x1, x2, …, x n are the independent variables of the model, and specifically, the results of any indirect measurements, such as optical measurements converted at a (specific) number of wavelengths. The goal of this first step is to develop a useful function f that reflects the relationship between the indirect measurement and the expected property value of interest to be predicted. In the second step (prediction), the measured values x1, x2, ..., x n This functionality can be evaluated to obtain an estimate of a direct measurement (e.g., the concentration of an analyte) when an optical measurement is performed at some future time, without requiring a corresponding direct measurement.

[0004] There is a vast literature on creating calibration functions to predict the values of material properties of interest using NIR spectroscopy. Most of this prior art addresses general and specific concepts for creating calibration functions. Some of this literature specifically discusses compensating for interferences, such as environmental disturbances and instrument-specific errors, such as measurement errors and aging emitters, when creating calibration functions. However, this prior art does not consider identifying spectral outliers, let alone how to handle them when creating calibration functions, even though such outliers are likely to occur during the accumulation of indirect (optical) measurements. Therefore, they are also likely to impair the creation of calibration functions. Outliers can be due to measurement variability or instrument error; the latter are sometimes excluded from the dataset. Generally speaking, and particularly in statistics, outliers are considered data points that differ significantly from other observations, but this can be subject to subjective interpretation and misinterpretation. On the other hand, including data points at the edges of a dataset is essential for a meaningful and robust calibration and should not be simply skipped simply because they appear odd. This suggests that one, if not the main, issue with outliers is their detection or identification, as there is no strict definition of an outlier. Ultimately, therefore, determining whether an observation is an outlier remains a subjective exercise. Due to the lack of a universally accepted definition of outliers, a variety of methods exist for detecting them. Some are graphical, such as the normal probability plot, others are model-based, and hybrid methods, such as so-called boxplots, exist. The choice of method and how to handle outliers often depends on the specific circumstances. Even if a normal distribution model is appropriate for the analyzed data, outliers in large sample sizes must be anticipated and should not be automatically discarded when present. Applications should use classification algorithms that are robust to outliers to model data with naturally occurring outliers. Removing outliers is a controversial practice opposed by many scientists and science lecturers. It is more acceptable in practice because the underlying model of the process and the general distribution of measurement errors are reliable. Outliers caused by instrument reading errors can be excluded, but at least verification of the readings is desirable. Specifically, in the case of infrared spectroscopy, identifying outliers as accurately and efficiently as possible, and preferably simultaneously, is a significant challenge. This is even more relevant when the corresponding infrared spectrum (indirect measurement) and corresponding reference data (direct measurement) are used as the basis for creating a calibration function suitable for predicting the value of a material property of interest using infrared spectroscopy.

[0005] Therefore, there is a need for a method of predicting the value of a property of interest of a material from infrared spectroscopy that enables reliable and automatic identification and removal of spectral outliers during creation of a calibration function. Summary of the Invention

[0006] Therefore, the object of the present invention is a computer-implemented method for predicting the value of a property of interest of a material, said method comprising the steps of:

[0007] a) providing a population of infrared spectra of samples, wherein the spectra form an m x n input data matrix X, where m is the number of samples in a row and n is the data points in a column,

[0008] b) removing spectral outliers from the spectral group of step a), comprising the following steps:

[0009] b1) Obtain the principal components by performing principal component analysis on the matrix X,

[0010] b2) Generate a diagonal matrix ∑ from the input data matrix X, which contains the singular values σ of the matrix X m and the loading matrix V,

[0011] b3) Calculate the score x for each spectrum by multiplying each data point of the input data matrix X with the loading of each component from step b2) m , forming the mean of each column of the X matrix to provide B 0,m The score index si is calculated by the following formula

[0012]

[0013] b4) Determine the component N whose eigenvalue causes the regression of X to converge for at least 99% of the scores C The number of, and the distance metric threshold D of each spectrum in step a) is calculated by the following formula i

[0014]

[0015] b5) calculating the average of all scores for each principal component of each spectrum of step a) and calculating a distance measure between said average and each score of each principal component,

[0016] b6) when the distance metric value of the score of the principal component obtained in step b5) is greater than the distance metric threshold of step b4), the sample spectrum is regarded as a spectral outlier,

[0017] b7) removing the spectral outliers of step b6) from the spectral group of step a) to obtain a clean spectral group,

[0018] c) generating a prediction function on the cleaned spectral population of step b7),

[0019] d) providing an infrared spectrum of a sample of unknown origin and / or composition or a sample of the same origin and / or composition as the sample in step a), and

[0020] e) predicting the value of the attribute of interest from the spectrum of step d) using the prediction function of step c). DETAILED DESCRIPTION

[0021] The sample infrared spectra provided in step a) define an n-dimensional data space, where n is the number of data points in the spectrum and m is the number of samples. m×n The input data matrix is represented by m rows for the samples and n columns for the data table, often also written as an m×n matrix X. The large amount of data makes the resulting matrix quite complex. Therefore, the input data matrix undergoes data reduction without losing relevant information. This is usually done in principal component analysis (PCA). This is a statistical procedure in which an orthogonal transformation is used to convert a set of observations of possibly correlated variables (each entity has a different value) into a set of linearly uncorrelated variable values called principal components. This transformation is defined in a specific way so that the first principal component has the largest possible variance, that is, it causes the greatest possible variability in the data. Each further component is orthogonal to the previous component and is second only to the previous component in terms of the highest possible variance. The resulting vectors (each of which is a linear combination of the variables and contains n observations) are uncorrelated orthogonal basis sets. In principal component analysis, the input data matrix X m×n is decomposed into two mutually orthogonal matrices, namely U m×n Matrix and Matrix. Mathematically, this step is described by the following formula:

[0022]

[0023] is the so-called loading matrix, which contains the new axes generated by the transformation in and , and U m×n Also known as the so-called score matrix, which contains the new coordinates. The loadings can be understood as the weights that are multiplied by the original variables to calculate the principal components.

[0024] According to the present invention, the input data matrix X is decomposed into the product of several matrices by singular value decomposition, which is mathematically described as:

[0025]

[0026] where X m×n is a matrix m×n of rank k, U m×m is a simple m×m matrix, ∑ m×n is the singular value σ of the matrix X m A real m×n matrix, and is the adjoint matrix of the unitary n×n matrix V containing the load; the real number of the adjoint matrix is equivalent to the transposed matrix

[0027] In the context of the present invention, the infrared spectrum group is not subject to any restrictions regarding the number of spectra, as long as the number of spectra is sufficient to produce a meaningful prediction function. Thus, the number of spectra in the spectrum group can range from 50 to 10,000, from 50 to 5,000, from 50 to 2,500, from 50 to 2,000, from 50 to 1,500, from 50 to 1,000, from 100 to 1,000, from 50 to 500, from 100 to 500, from 50 to 250, from 100 to 250, or from 50 to 100.

[0028] One or more infrared spectra can be recorded at wavelengths between 400 and 2,800 nm, using any suitable spectrometer operating according to the monochromator principle or the Fourier transform principle. Preferably, infrared spectra are recorded between 1,100 and 2,500 nm. Wavelengths are easily converted into corresponding wavenumbers, and therefore infrared spectra can also be recorded at corresponding wavenumbers. Typically, infrared spectra require the presence of functional groups or moieties that allow the radiation to excite molecular vibrations in the material, and the frequencies of the light passing through are then recorded by the spectrometer. However, infrared spectra can also be recorded for materials without any functional groups or moieties, such as minerals that, on their own, do not allow any (molecular) vibrations. However, this requires the co-existence of an infrared-active material, which interacts with an inactive material, for example, in the formation of a composite compound, thereby causing vibrational changes in the excitation of molecular vibrations. Comparing the infrared spectrum thus obtained with the spectrum of the (pure) infrared-active material without the additional infrared-inactive material allows conclusions to be drawn regarding the properties and concentration of the infrared-inactive material. Therefore, infrared spectra enable prediction of the values of properties of interest for a wide range of different materials. However, biological samples, such as feed, contain a large number of diverse organic and inorganic compounds and thus represent a rather complex matrix. Despite this, each biological substance provides a unique near-infrared spectrum, comparable to a single fingerprint. Therefore, two biological substances that give identical spectra can be assumed to have the same physical and chemical composition and, therefore, to be identical. On the other hand, if two biological substances give different spectra, they can be assumed to be different, either in terms of physical or chemical properties, or both. Due to the distinct and highly specific absorption bands of organic and inorganic compounds, the signals and intensities of organic and inorganic compounds in infrared spectra can be readily correlated with the specific organic compound and its concentration in a sample of known weight. Therefore, infrared spectroscopy enables the prediction of properties of interest in quite complex materials, such as the identity and concentration of different amino acids and proteins in a biological sample. However, infrared spectroscopy also enables the identification of abstract parameters in complex samples, while general trends can be discerned across samples of the same type due to variations in these abstract parameters. For example, the absorption and intensity in the infrared spectra of complex samples of feed or feed ingredients can be attributed and correlated with abstract parameters such as trypsin inhibitor activity, urease activity, protein solubility in alkali, and protein dispersibility and concentration. In the next step, the infrared spectrometer in question must be calibrated. Once the absorption intensities at individual wavelengths or wavenumbers have been successfully matched—that is, attributed and correlated—with the parameters of interest and their values, near-infrared spectroscopy can reliably predict the values of the material's properties of interest. To this end, a sufficient number of infrared spectra are recorded for the relevant material—if necessary, 100, 200, 300, 400, 500, or more—and the absorption intensities at individual wavelengths or wavenumbers are matched to the corresponding parameters and their values. Infrared spectra can be measured in transmission or reflection mode.In academic applications, the most common measurement mode is transmission mode, but for measurements of insoluble or granular materials, reflection mode is preferred. In the latter case, the reflectance of the light returning from the sample is measured, and the difference between the emitted and reflected light is given as absorbance.

[0029] Typically, the method according to the present invention is not subject to any restrictions on the material whose attribute value is being predicted. Therefore, it can be a single substance or a mixture of different substances. As explained in more detail above, the material must contain at least one infrared active substance to allow the excitation of molecular vibrations. In addition, the material can be any type of component and / or source. Suitable materials in the context of the present invention are organic compounds, such as amino acids, peptides, proteins, carboxylic acids, fatty acids, saturated and / or unsaturated fatty acids, polyunsaturated fatty acids or mixtures thereof, or substances of human, animal or plant origin, such as digesta, body parts, animal parts or plant parts, or technical production products based on plant and / or animal-derived products, such as feed, animal feed and / or feed ingredients, such as wet distiller's grains (DDGS) containing solubles, hydrolyzed fish meal or hydrolyzed feather meal, and inorganic compounds, such as carbonates, phosphates, nitrates and sulfates, or mixtures thereof. Preferably, the material whose property value of interest is to be predicted is an amino acid, protein, peptide, carboxylic acid, saturated, unsaturated and / or partially saturated fatty acid, polyunsaturated fatty acid, feed, feedstuff, feed raw material, substance of human, animal and / or plant origin or a mixture thereof.

[0030] The material that is subjected to the method according to the invention strongly influences the attribute value of being paid attention to that will be predicted. For example, when feed, feedstuff or feed raw material are materials that have been subjected to the method according to the invention, the attribute value of being paid attention to that will be predicted is particularly amino acid, polyunsaturated fatty acid and / or mineral nutrient. For example, when digesta, body parts or animal parts are materials that have been subjected to the method according to the invention, the attribute value of being paid attention to that will be predicted is specifically protein, peptide and / or amino acid. Therefore, the method according to the invention is not subject to any restriction about the attribute value of being paid attention to that will be predicted in principle. Therefore, the attribute value of being paid attention to can be any type of material parameter that can be determined directly or indirectly by infrared spectroscopy.

[0031] According to the present invention, the sample whose infrared spectrum is provided in step d) is a sample of unknown origin and / or composition, or a sample with the same origin and / or composition as the sample in step a). Preferably, the sample in step d) has an unknown origin and / or composition. Alternatively, it is preferred that the sample in step d) has the same origin as the sample in step a), but an unknown composition. Preferably, the sample whose infrared spectrum is provided in step d) is a sample of unknown origin and / or composition, or a sample with the same origin and / or composition as the sample group in step a). Preferably, the sample in step d) has an unknown origin and / or composition. Alternatively, it is preferred that the sample in step d) has the same origin as the sample group in step a), but an unknown composition.

[0032] Since infrared spectra can be measured in transmission or reflection mode, the samples can be unground and / or ground samples, where the ground samples can have different particle sizes.

[0033] In principle, the method according to the invention is not subject to any restrictions with respect to the distance metric used. Thus, in principle, any suitable distance metric can be used in the method, such as a Euclidian distance measure, a Pearson distance measure, or a Mahalanobis distance measure. In general, a distance metric can be defined from the similarity measure s(i, j) of two observations i and j as the distance coefficient d(i, j) of the two observations i and j by the following formula:

[0034]

[0035] Therefore, the distance metric can also be obtained indirectly, that is, first determining the similarity metric and then obtaining the distance metric from the similarity metric.

[0036] In an embodiment of the method according to the invention, the distance metric is a Euclidean distance metric, a Pearson distance metric, a Mahalanobis distance metric or a distance metric obtained from a similarity metric.

[0037] Typically, any type of similarity analysis uses thresholds that define certain limits, i.e., which observations are still acceptable and which are no longer. Using strict thresholds for distance or similarity metrics often works well in laboratory conditions, and therefore under very reproducible conditions, with well-trained personnel and spectrometers that are perfectly calibrated and part of a network. However, using strict thresholds for distance or similarity metrics does not work well when using so-called stand-alone spectrometers, which are not typically operated under laboratory conditions. Furthermore, choosing a strict threshold for a distance metric can be quite subjective, especially if the threshold is simply randomly selected or adopted from other processes. Therefore, using strict, predefined thresholds for distance or similarity metrics is likely to lead to similar problems as in the prior art. However, the method according to the present invention is also applicable to predicting the value of the attribute of interest using so-called stand-alone spectrometers, which are not typically operated under laboratory conditions.

[0038] In the context of the present invention, it was discovered that these problems are addressed by using a so-called dynamic threshold for the distance metric. This dynamic thresholding approach involves incrementally increasing the distance metric threshold obtained in step b4) by +1. The resulting "new" distance metric threshold serves as the basis for evaluating whether a sample spectrum is a spectral outlier, replacing the distance metric threshold obtained in step b4). This process is repeated until the difference between the two maximum distance metric values is at least 1, and the maximum distance value is at most 8.

[0039] In another embodiment of the method according to the present invention, step b) further comprises the following steps:

[0040] b5.1) Increase the distance metric threshold obtained in step b4) by +1,

[0041] b5.2) using the distance metric threshold of step b5.1) to determine the two distance metrics obtained in step b5) with the highest values,

[0042] b5.3) determining the difference between the distance metric values determined in step b5.2), and

[0043] b5.4) Repeat steps b5.1) to b5.3) with the distance metric threshold value of step b5.1) until the difference value determined in step b5.3) is at least 1 and the maximum value of the distance metric is at most 8.

[0044] Each time steps b5.1) to b5.3) are repeated, the distance metric threshold of the last run replaces the distance metric threshold of the previous run (i.e., the run before the last run), and is then incremented by +1 in the repeated step b5.1). Thus, when steps b5.1) to b5.3) are repeated for the first time, the distance metric threshold obtained in step b4) is replaced by the distance metric threshold obtained in step b5.1), and this distance metric threshold is then incremented by +1 in the repeated step b5.1). When steps b5.1) to b5.3) are repeated again, the distance metric threshold of the last run replaces the distance metric threshold used in the second to last run of repeating step b5.1), and the distance metric threshold in the last run is then incremented by +1 in the repeated step b5.1).

[0045] Preferably, the generation of the prediction function for the clean spectral population of step b7) involves the use of a reference data set obtained from determining the value of the property of interest in each sample of step a) in the quantitative analysis.

[0046] In another embodiment, the method according to the present invention further comprises the following steps:

[0047] a1) determining the value of the attribute of interest in each sample of step a) in a quantitative analysis to provide a reference data set.

[0048] In the context of the present invention, the term quantitative analysis is used in its broadest sense known to those skilled in the art, preferably in the field of analytical chemistry, and refers to any chemical and / or physical method used to determine the absolute or absolute presence of one, several, or all specific parameters, such as the presence of a specific substance in a sample. For example, when the sample in question is derived from feed, feed, or feed ingredients, the specific parameter may be the absolute or absolute presence of one, several, or all of the following amino acids: methionine, cysteine, cystine, threonine, leucine, arginine, isoleucine, valine, histidine, phenylalanine, tyrosine, tryptophan, glycine, serine, proline, alanine, aspartic acid, glutamic acid, the total amount of lysine, the amount of lysine reaction, and the ratio of the amount of lysine reaction to the total amount of lysine. The term lysine reaction is used to denote the amount of lysine that is actually available to the animal, in particular for digestion. In contrast, the term "total lysine" is used to denote the sum of the amount of lysine that is actually available to the animal, in particular for digestion, and the amount of lysine that is not available to the animal, in particular for digestion. The latter lysine amount is usually due to lysine degradation reactions, such as the already mentioned Maillard reaction.

[0049] Many feeds undergo processing, which can damage amino acids. This can render some amino acids nutritionally unavailable. This is particularly true for lysine, which has an ε-amino group that can react with carbonyl groups of other compounds present in the diet, such as reducing sugars, to produce compounds that are partially absorbed from the intestine but have no nutritional value to the animal. The reaction of the ε-amino group of free and / or protein-bound lysine with reducing sugars during heat processing is called the Maillard reaction. This reaction produces early and late Maillard products. Early Maillard products are structurally altered lysine derivatives known as Amadori compounds, while late Maillard products are known as melanoidins. Melanoidins do not interfere with normal lysine analysis and have no effect on calculated digestibility values. They simply result in lower lysine concentrations being absorbed. Therefore, melanoidins are not typically identified in routine amino acid analyses. In contrast, Amadori compounds do interfere with amino acid analysis and provide inaccurate lysine concentrations in the analyzed sample. Lysine bound to these compounds is called "blocked lysine" and is biologically unavailable because it blocks any enzymatic degradation in the gastrointestinal tract.

[0050] The amount of active lysine in a sample can be determined using the Sanger reagent, 1-fluoro-2,4-dinitrobenzene (FNDB). Therefore, the lysine measured using this method is also known as FDNB-lysine. The Sanger reagent converts lysine into yellow dinitrophenyl (DNP)-lysine, which can be extracted and measured spectrophotometrically or by high-performance liquid chromatography at a wavelength of 435 nm.

[0051] Alternatively, the reactive lysine content in a sample can be determined by guanidation using the mild reagent n-methylisourea. In this method, n-methylisourea reacts only with the ε-amino group of lysine, but not with the α-amino group. Therefore, guanidation can be used to determine both free and peptide-bound lysine. Therefore, guanidation is preferred for determining reactive lysine. Guanidination of lysine produces homoarginine, which is further derivatized with ninhydrin, and the resulting absorbance change is measured at 570 nm. The derivatized sample is then hydrolyzed to yield homoarginine again. Reactive lysine can also be determined by guanidation of intact, protein-bound lysine in alkaline medium to yield homoarginine. In this type of reaction, guanidation is typically achieved via the action of n-methylisourea (OMIU).

[0052] Because this method is easier to use, it is preferred to use the guanidation reaction to determine reactive lysine. The guanidation reaction involves incubating the feed raw material and / or feed sample in n-methylisourea. Preferably, the ratio of n-methylisourea to lysine is greater than 1000. The sample thus treated, obtained from step i), is dried and analyzed for homoarginine, preferably by using ion exchange high performance liquid chromatography. Subsequently, the sample is derivatized with ninhydrin, and the absorbance of the derivatized sample is measured at a wavelength of 570 nm. Thereafter, the sample is hydrolyzed. The weight and molar amount of homoarginine in the sample are determined. Finally, the amount of reactive lysine is calculated from the molar amount of homoarginine.

[0053] However, during the processing of feed raw materials and / or feeds, not only lysine but also other amino acids are thermally destroyed. According to the method of the present invention, the amino acids methionine, cysteine, cystine, threonine, leucine, arginine, isoleucine, valine, histidine, phenylalanine, tyrosine, tryptophan, glycine, serine, proline, alanine, aspartic acid and glutamic acid are quantitatively analyzed in feed raw materials and / or feed samples. To a certain extent, amino acids exist not only in the form of single compounds, but also in the form of oligopeptides formed by two, three or even more amino acids in equilibrium reactions, such as dipeptides, tripeptides or higher peptides. The amino group of an amino acid is usually too weak as a nucleophile to react directly with the carboxyl group of another amino acid, or in the protonated form (–NH3 +) exists. Therefore, under standard conditions, the equilibrium of this reaction is generally on the left side. However, depending on the conditions of the individual amino acids and the sample solution, some amino acids to be determined may not exist as single compounds, but rather as oligopeptides composed of two, three, or even more amino acids, such as dipeptides, tripeptides, or higher peptides. Therefore, samples of feed materials and / or feed should be hydrolyzed using, for example, hydrochloric acid or barium hydroxide, preferably acidic or alkaline hydrolysis. To facilitate the separation of free amino acids and / or the identification and determination of amino acids, the free amino acids are derivatized using a chromogenic reagent, if necessary. Suitable chromogenic reagents are known to those skilled in the art. Subsequently, the free amino acids or derivatized free amino acids are subjected to chromatographic separation, where different amino acids are separated from each other due to different retention times caused by the different functional groups of the individual amino acids. Suitable chromatographic columns, such as ion exchange columns or reversed-phase columns, as well as suitable elution solvents for chromatographic separation of amino acids are known to those skilled in the art. The separated amino acids are ultimately identified in the eluate of the chromatographic step by comparison with calibrated standards prepared for analysis. Typically, amino acids eluting from the column are detected using an appropriate detector, such as conductivity, mass-specific, fluorescence, or UV / VIS, depending on the time of derivatization with a chromogenic reagent. This produces a chromatogram of peak areas and heights for each amino acid. Individual amino acids are determined by comparing the peak areas and heights with calibrated standards or a calibration curve for each amino acid. Since both cystine (HO₂C(–H₂N)CH–CH₂–S–S–CH₂–CH(NH₂)–CO₂H) and cysteine (HS–CH₂–CH(NH₂)–CO₂H) are measured as cysteine (HO₃S–CH₂–CH(NH₂)–CO₂H) after acidic hydrolysis, no distinction is made between the two amino acids in the quantitative analysis due to their high oxidizability.

[0054] The above procedure is generally used for the quantitative analysis of the total amount of lysine, which is necessary for determining the ratio between the reacted amount of lysine and the total amount of lysine, and for the quantitative analysis of at least one amino acid selected from the group consisting of methionine, cysteine, cystine, threonine, leucine, arginine, isoleucine, valine, histidine, phenylalanine, tyrosine, tryptophan, glycine, serine, proline, alanine, aspartic acid, and glutamic acid.

[0055] The most critical aspect of amino acid quantification is sample preparation, which varies depending on the type of component and the primary amino acid of interest. Most amino acids can be hydrolyzed in hydrochloric acid (6 mol / l) for up to 24 hours. For the sulfur-containing amino acids methionine, cysteine, and cystine, oxidation with performic acid is preferred prior to hydrolysis. For the quantitative analysis of tryptophan, barium hydroxide (1.5 mol / l) is used for hydrolysis for 20 hours.

[0056] When the sample in question originates from a feed, a fodder or a feed ingredient, the specific parameters may also be one, several or all of the following from the list consisting of trypsin inhibitor activity, urease activity, protein solubility in alkali and protein dispersibility index of the sample.

[0057] Quantitative analysis of trypsin inhibitor activity is based on the inhibitor's ability to form a complex with the enzyme trypsin, thereby reducing its activity. Trypsin catalyzes the synthesis of the substrates N-α-benzoyl-D,L-arginine-p-nitroanilide (DL-BAPNA, IUPAC name N-[5-(diaminomethyleneamino)-1-(4-nitroanilino)-1-oxopentan-2-yl]benzylamide) and N-α-benzoyl-L-arginine-p-nitroanilide (L-BAPNA, IUPAC name N-[5-(diaminomethyleneamino)-1-(4-nitroanilino)-1-oxopentan-2-yl]benzylamide). This catalytic hydrolysis releases the yellow product, free nitroaniline, which results in a change in absorbance. Trypsin activity is proportional to the yellow color. The concentration of p-nitroaniline can be determined spectroscopically at a wavelength of 410 nm. L-BAPNA is commonly used in the ISO 14902 (2001) method, and DL-BAPNA is commonly used in the AACC 22.40-01 method (a modification of the method originally invented by Hamerstrand in 1981).

[0058] In the ISO 14902 method, the sample is first finely ground using a 0.50 mm sieve. Any exotherm should be avoided during the grinding process. The ground sample is mixed with an alkaline aqueous solution, for example, 1 gram of sample dissolved in 50 ml of 0.01 N sodium hydroxide solution. The resulting solution, suspension, dispersion, or emulsion is then stored at 4°C for up to approximately 24 hours. The pH of the mixture thus obtained is between 9 and 10, particularly between 9.4 and 9.6. The resulting solution is diluted with water and allowed to stand. A sample of this solution, for example, 1 ml, is taken and diluted, depending on its estimated or previously estimated trypsin inhibitor activity, such that 1 ml of the diluted solution results in 40% to 60% inhibition of the enzymatic reaction. For example, 1 ml of the trypsin working solution is added to a mixture of L-BAPNA, water, and a diluted sample extraction solution, for example, 5 ml of L-BAPNA, 2 ml of (distilled) water, and 1 ml of the appropriately diluted sample extraction solution. The sample is then incubated at 37°C for precisely 10 minutes. The reaction is terminated by the addition of 1 ml of 30% acetic acid. Blank samples were prepared as described above, but trypsin was added after the acetic acid. After centrifugation at 2.5 g, the absorbance at a wavelength of 410 nm was measured.

[0059] In the AACC 22-40.01 method, the sample is first finely ground using a 0.15 mm sieve. Any exotherm should be avoided during the grinding process. The ground sample is then mixed with an alkaline aqueous solution, for example, 1 gram of sample dissolved in 50 ml of 0.01 N sodium hydroxide solution, and stirred slowly at 20°C for 3 hours. The pH of the resulting solution, suspension, dispersion, or emulsion should be between 8 and 11, preferably between 8.4 and 10. The resulting solution, suspension, dispersion, or emulsion is diluted with water, shaken, and allowed to stand. A sample of this solution, for example, 1 ml, is taken and diluted based on its estimated or previously estimated trypsin inhibitor activity, such that 1 ml of the diluted solution results in 40% to 60% inhibition of the enzymatic reaction. For example, 2 ml of the trypsin working solution is added to a mixture of D,L-BAPNA, water, and a diluted sample extract solution, for example, 5 ml of D,L-BAPNA, 1 ml of (distilled) water, and 1 ml of the appropriately diluted sample extract solution. The sample is then incubated at 37°C for exactly 10 minutes. The reaction was stopped by adding 1 ml of acetic acid (30%). The blank samples were prepared as described above, but trypsin was added after the acetic acid. After centrifugation at 2.5 g, the absorbance at a wavelength of 410 nm was measured.

[0060] Independent of the method used, trypsin inhibitor activity was calculated as mg trypsin inhibitor / g trypsin using the following formula:

[0061]

[0062] i = inhibition percentage (%);

[0063] Ar = absorbance of the solution relative to the standard;

[0064] Abr = absorbance of blank relative to standard;

[0065] As = absorbance of the solution relative to the sample;

[0066] Abs = absorbance of the blank relative to the sample;

[0067]

[0068] TIA = trypsin inhibitor activity (mg / g);

[0069] i = inhibition percentage (%);

[0070] m0 = mass of the test sample (g);

[0071] m1 = mass of trypsin (g);

[0072] f1 = dilution factor of the sample extract; and

[0073] f2 = conversion factor based on trypsin purity.

[0074] One trypsin unit is defined as the amount of enzyme that increases the absorbance at 410 nm by 0.01 units per 1 ml reaction volume after 10 minutes of reaction. Trypsin inhibitor activity is defined as the number of trypsin inhibited units (TUI). TUI per ml is calculated using the following formula:

[0075]

[0076] in

[0077] A blank = absorbance of blank

[0078] A sample = Sample absorbance

[0079] V dl.smp. = volume of diluted sample solution (ml).

[0080] The TUI thus obtained was plotted against the volume of the diluted sample solution, with the inhibitor volume extrapolated to 0 ml to give the final TUI [ml]. Finally, the TUI per gram of sample was calculated using the formula

[0081] TUI[g]=TUI[ml -1 ]×d×50

[0082] Where d = dilution factor (final volume divided by amount taken).

[0083] The results of this analytical method should not exceed 10% of the mean of the replicate samples.

[0084] Urease catalyzes the degradation of urea into ammonia and carbon dioxide. Since urease is naturally present in soybeans, the quantitative analysis of this enzyme is the most common test for evaluating the quality of processed soybeans. Preferably, the quantitative analysis of urease is carried out according to the method of ISO 5506 (1988) or AOCS Ba 9-58. The method of AOCS Ba 9-58 determines the residual activity of urease as an indirect index to assess whether protease inhibitors are destroyed during the processing of feed materials and / or feeds. The residual activity of the urease is measured as the increase in pH value caused by the release of the alkaline compound ammonia into the medium in the test. The recommended level for the increase in pH value is 0.01 to 0.35 units (NOPA, 1997). A typical quantitative analysis of urease activity in feed ingredients and / or feedstuffs is performed as follows: First, a solution of urea dissolved in a buffer containing NaH2PO4 and KH2PO4 is prepared, for example, by adding 30 g of urea to 1 L of a buffer solution consisting of 4.45 g of Na2HPO4 and 3.4 g of KH2PO4, and the pH value thus obtained is measured. Subsequently, a sample of the feed ingredient and / or feedstuff, for example, a 0.2 g sample of soybeans, is added to this solution. The test tube or beaker containing the solution, suspension, dispersion, or emulsion thus obtained is placed in a water bath, for example, at a temperature of 30 ± 5°C, preferably 30°C, for 20 to 40 minutes, preferably 30 minutes. Finally, the pH value of the solution, suspension, dispersion, or emulsion is measured, compared with the pH value of the original urea solution, and the difference is reported as the increase in pH value.

[0085] The solubility of proteins in alkali, also referred to below as the solubility of proteins in alkaline solutions or the alkaline solubility of proteins, for example according to DIN EN ISO 14244, is an effective method for distinguishing between overprocessed and correctly processed products.

[0086] The solubility of protein in alkali, or the alkaline solubility of protein, involves determining the percentage of protein that dissolves in an alkaline solution. Before dissolving a sample of feed raw material and / or feed of known weight, the nitrogen content of the sample of a specific weight is determined using a standard method for determining nitrogen, such as the Kjeldahl method or the Dumas method. The nitrogen content thus determined is referred to as the total nitrogen content. Subsequently, the same weight of the sample from the same source is suspended in an alkaline solution of a defined concentration, preferably an alkali metal hydroxide solution, specifically a potassium hydroxide solution. An aliquot of the suspension thus obtained is removed and centrifuged. Another aliquot of the suspension thus obtained is removed. The nitrogen content of this liquid portion is determined using a standard method for determining nitrogen, such as the Kjeldahl method or the Dumas method. The nitrogen content thus determined is compared with the total nitrogen content and expressed as a percentage of the original nitrogen content of the sample.

[0087] For example, the pH value of a typical alkaline solution for measuring protein alkaline solubility is 12.5, and the concentration of potassium hydroxide solution is 0.036 mol / l or 0.2 weight %. In step ii), for example, 1.5 grams of soybean samples are placed in 75 ml of potassium hydroxide solution and subsequently stirred at 20° C. for 20 minutes at 8500 rpm (revolutions per minute). Subsequently, an aliquot of, for example, about 50 ml of the suspension, solution, dispersion or emulsion thus obtained is taken out and immediately centrifuged at 2500 g for 15 minutes. Then, an aliquot of, for example, 10 ml of the supernatant of the suspension, solution, dispersion or emulsion thus obtained is taken out and the nitrogen content in the aliquot is measured by a standard method for measuring nitrogen, for example, the Kjeldahl method or the Dumas method. Finally, the result is expressed as the percentage of nitrogen content in the sample.

[0088] The determination of the Protein Dispersibility Index (PDI) measures the solubility of a protein in water after mixing the sample with water. This method also involves determining the nitrogen content of a sample of known weight, which is usually performed according to the same procedures as for wet chemical analysis of proteins. The nitrogen content thus obtained is also referred to as the total nitrogen content. In addition, the method also includes preparing a suspension of the same weight of sample as that used to determine the nitrogen content when suspended in water, which is usually achieved using a high-speed mixer. The suspension thus obtained is filtered and the filtrate is centrifuged. The nitrogen content of the supernatant thus obtained is determined by again using standard methods for determination, such as the Kjeldahl method or the Dumas method as described above. The nitrogen content thus obtained is also referred to as the nitrogen content in solution. The Protein Dispersibility Index is ultimately calculated as the ratio between the nitrogen content in solution and the total nitrogen content and expressed as a percentage of the original nitrogen content of the sample.

[0089] Since the value of the protein dispersibility index increases with decreasing particle size, the result obtained in determining the protein dispersibility index depends on the particle size of the sample. Therefore, the sample to be determined for the protein dispersibility index is preferably ground, in particular to a 1 mm mesh size.

[0090] The above procedure is carried out according to the official method Ba 10-65 of the American Oil Chemists Society (AOCS), and the protein dispersibility index is preferably determined according to this method. For example, the nitrogen content of the soybean sample is determined by a standard method for determining nitrogen, such as the Kjeldahl method or the Dumas method. For example, a 20-gram aliquot of the soybean sample is placed in a blender, 300 ml of (deionized) water is added at 25° C., and then stirred for 10 minutes at 8500 rpm. The suspension, solution, dispersion or emulsion thus obtained is filtered and the solution, dispersion or emulsion thus obtained is centrifuged at 1000 g for 10 minutes. Finally, the nitrogen content in the supernatant is determined by a standard method for determining nitrogen, such as the Kjeldahl method or the Dumas method.

[0091] The reference dataset obtained from step a1) is used to find data correlations with the clean spectral population of step b1) to generate a prediction function in step c). In the case of the absolute or relative presence of one, several, or all relevant substances in the sample, the data correlation refers to the concentration of the relevant substances in the sample.

[0092] In a preferred embodiment of the method according to the invention, the generation of the prediction function in step c) comprises analyzing the data correlation of the reference data set of step a1) and the cleaned spectral population of step b7) to give a prediction function.

[0093] By verifying the prediction function generated in step c), the quality of the prediction achieved by the method according to the present invention can be improved.

[0094] In another embodiment of the method according to the invention, the generation of the prediction function in step c) comprises the validation of said prediction function.

[0095] In principle, method according to the present invention is not subject to any restriction about verification type. Therefore, in method according to the present invention, any suitable verification program can be used. Nevertheless, in the context of the present invention, preferably use hold-out cross validation, k-fold validation and / or the verification for ring test data. A kind of selection of the back is to verify for external spectrum and corresponding external reference data. The advantage of this selection is that the calibration group is not only evaluated according to its performance on the spectrograph of main spectrometer, i.e. for creating prediction function. On the contrary, this selection also makes it possible to evaluate the performance of the calibration group with respect to a large amount of, for example, hundreds of other spectrographs that also use calibration function. Therefore, preferential use is made of the verification for ring test data.

[0096] In a preferred embodiment of the method according to the invention, the validation is a leave-one-out cross validation, a k-fold validation and / or a validation on circular test data.

[0097] In the context of the present invention, the terms leave-one-out cross-validation, k-fold validation and validation on circular test data are used as known to a person skilled in the art.

[0098] For example, when the validation is a leave-one-out cross validation, step c) of the method according to the present invention further comprises the following steps:

[0099] cV1) dividing the reference data of step a1) and the clean spectrum group of step b7) into a training set S0 and a test set S1, wherein the size of the test set S1 is smaller than the size of the training set S0,

[0100] cV2) generating a preliminary prediction function on the training set S0 of step cV1),

[0101] cV3) predicting the value of the property of interest by applying the preliminary prediction function of step cV2) to the clean spectral population,

[0102] cV4) calculating the average prediction error of the preliminary function of step cV2) using the predicted attribute values of step cV3) and the corresponding reference data of step cV1),

[0103] cV5) when the average prediction error of the preliminary prediction function is outside the limit, repeating steps cV1) to cV5), or when the average prediction error of the preliminary prediction function is within the limit, continuing with step cV6), and

[0104] cV6) Approving the preliminary prediction function of step cV2) as the prediction function.

[0105] For example, when the validation is a k-fold cross validation, step c) of the method according to the present invention further comprises the following steps:

[0106] cV1) dividing the set of clean spectra from step b7) and the corresponding reference data from step a1) into n subsets of equal size, where n is an integer with a minimum value of at least 2 and a maximum value less than the number of spectra, one of the subsets serving as a test set and the remaining subsets serving as training sets,

[0107] cV2) generates a preliminary prediction function on the training set of step cV1),

[0108] cV3) predicting the value of the property of interest by applying the preliminary prediction function of step cV2) to the clean spectral population,

[0109] cV4) calculating the average prediction error of the preliminary prediction function using the predicted attribute value of step cV3) and the corresponding reference data of step cV1),

[0110] cV5) when the average prediction error of the preliminary prediction function is outside the limit, repeating steps cV1) to cV5), or when the average prediction error of the preliminary prediction function is within the limit, continuing with step cV6),

[0111] cV6) approving the preliminary prediction function of step cV2) as the prediction function, and performing steps cV2) to cV6) k times, using each of the n subsets as a test set once to give k prediction functions, and

[0112] cV7) Averaging the parameters of each approved prediction function in step cV6) to give a calibration function.

[0113] For example, when the verification is for ring test data, step c) of the method according to the present invention further comprises the following steps:

[0114] cV1) provides a set of external reference spectra and corresponding reference data,

[0115] cV2) generating a preliminary prediction function for the cleaned spectral population of step b7) and the reference data of step a1),

[0116] cV3) predicting the value of the property of interest by applying the preliminary prediction function of step cV2) to the external reference spectrum of step cV1),

[0117] cV4) calculating the average prediction error of the preliminary prediction function of step cV2) using the predicted attribute values of step cV3) and the corresponding external reference data of step cV1),

[0118] cV5) when the average prediction error of the preliminary prediction function is outside the limit, repeating steps cV1) to cV5), or when the average prediction error of the preliminary prediction function is within the limit, continuing with step cV6), and

[0119] cV6) Approving the preliminary prediction function of step cV2) as the prediction function.

[0120] Depending on the type of validation, the average prediction error is also called the root mean square error of cross validation (RMSECV), when the validation is internal validation, specifically cross validation, and is calculated as follows:

[0121]

[0122] Where M = the number of predicted attribute values,

[0123]

[0124]

[0125] Generally, verification of the prediction function obtained by the method according to the present invention or any embodiment thereof provides a highly reliable and useful prediction function. Nevertheless, in rare cases, further evaluation of the prediction function generated by the method according to the present invention or any embodiment thereof may be appropriate. One useful approach is to plot the predicted property values from the cleaned spectrum population of step b7) and the prediction function from step c) as regression lines into a measurement-prediction plot to produce a scatter plot. Optionally, the property values of the reference data from step a1) are also plotted on the plot. An angle bisector is then plotted on the resulting scatter plot to divide the plane of the plot into two triangular planes of equal size. In theory, if the prediction function is perfect, the plotted prediction function and the angle bisector will be congruent. However, in practice, this rarely occurs. Next, the distance of each predicted property value from the angle bisector in the scatter plot is determined, and an overall ranking of the distances is formed, with the highest distances ranked first. Spectra corresponding to at least three of the top-ranked distances, i.e., predicted property values, are removed from the cleaned spectrum population of step b7). The prediction function from step c) is then corrected to compensate for the removal of these spectra. This method allows further fine-tuning of the prediction function. The results of this fine-tuning can be easily followed as the method is repeated step by step: each repetition brings the prediction function closer to the angle bisector.

[0126] In another preferred embodiment of the method according to the present invention, step c) further comprises the following steps:

[0127] c1) plotting the property value predicted by the clean spectrum group of step b7) and the prediction function of step c) as a regression line into a measurement-prediction graph to give a scatter plot,

[0128] c2) drawing an angle bisector between the axes of the graph of step c1),

[0129] c3) determining the distance from each predicted attribute value in step c1) to the angle bisector in the scatter plot in step c2),

[0130] c4) forming an overall ranking of the distances obtained in step c3) with the highest distances ranked first, and removing spectra corresponding to at least three of the top-ranked distances from the group of clean spectra of step b7), and

[0131] c5) Modifying the prediction function of step c) to compensate for the removal of the spectrum in step c4).

[0132] In the context of the present invention, the term modified prediction function is used to indicate that a new prediction function is generated on the cleaned spectral population due to the additional removal of spectra in step c4) above.

[0133] There may be cases where two or more spectra corresponding to the top-ranked distances to be removed from the clean spectrum group in step b7) are in (close) proximity in the scatter plot. However, removing two or more spectra from the (close) proximity can excessively affect the slope of the modified prediction function. In the worst case, the slope of the prediction function can be significantly affected, for example, by modifying the prediction function to cause the slope to be higher or lower than its previous value.

[0134] This effect of removing two or more spectra corresponding to the top-ranked distances from the clean spectrum group of step b7) when they are in (close) proximity in the scatter plot can be avoided or at least significantly reduced by weighting them. For this purpose, first, a perpendicular line is drawn through the angle bisector at half the length of the angle bisector, thereby dividing the plane above and below the angle bisector into two planes, respectively, where the planes mirrored on the angle bisector are equally sized. Typically, the plane between the axis of the scatter plot and the angle bisector and the perpendicular line drawn through the angle bisector at half its length is divided into four planes: two triangles of equal size and two squares of equal size, one triangle and one square above and below the angle bisector. Next, when one or more spectra corresponding to the top-ranked distances to be removed from the clean spectrum group in step b7) are present in the same plane given by the angle bisector and the perpendicular line, they are weighted by a factor of 0.5.

[0135] In another preferred embodiment, the method according to the invention therefore further comprises:

[0136] - in step c2), drawing a perpendicular line through the angle bisector at half the length of the angle bisector to divide the plane between the axes of the scatter plot into four planes, and

[0137] - In step c4), when the spectra corresponding to at least two top-ranked distances are present in the same plane, they are weighted with a factor of 0.5.

[0138] In principle, the above method can be performed until the prediction function and the angle bisector are congruent. In practice, steps c1) to c5) are repeated until there is no further improvement in the convergence of the prediction function to the angle bisector, or in other words, until the curve of the modified prediction function obtained as a regression line is as close as possible to the angle bisector.

[0139] In a further preferred embodiment, the method according to the present invention further comprises the following steps:

[0140] c6) Repeating steps c1) to c5) using the modified prediction function and the clean spectral population obtained from the previous correction, until the plot of the modified prediction function thus obtained as a regression line is as close as possible to the angle bisector.

[0141] As an alternative to or in addition to the process of steps c1) to c5), with or without step c6), the quality of the prediction function of step c) can be further improved by using confidence intervals. This involves first plotting the predicted property values of the cleaned spectrum population of step b7) and the prediction function of step c) as a regression line on a measurement-prediction plot to produce a scatter plot. Next, as in the above process, an angle bisector is drawn between the two axes of the scatter plot, and a confidence interval of a predetermined width is drawn around the angle bisector. Spectra corresponding to points outside the confidence interval are then removed from the cleaned spectrum population of step b7), and the prediction function of step c) is corrected to compensate for the removal of these spectra. In the context of the present invention, the width of the confidence interval is +5x RMSECV to –5x RMSECV, preferably +4x RMSECV to –4x RMSECV, +3x RMSECV to –3xRMSECV, +2x RMSECV to –2x RMSECV, +1x RMSECV to –1x RMSECV, or its width is any positive real number to the corresponding negative real number between +5x RMSECV and –5x RMSECV, around the corresponding value of the angle bisector, where RMSECV is the root mean square error of the cross validation defined above.

[0142] In another embodiment, the method according to the present invention further comprises the following steps:

[0143] c7) plotting the property value predicted by the clean spectrum group of step b7) and the prediction function of step c) as a regression line into a measurement-prediction graph to give a scatter plot,

[0144] c8) drawing an angle bisector between the two axes of the graph of step c7),

[0145] c9) plotting a confidence interval having a predetermined width around the angle bisector as the curve graph in step c8),

[0146] c10) removing spectra corresponding to points outside the confidence interval in the graph of step c9) from the cleaned spectrum group of step b7), and

[0147] c11) Modifying the prediction function of step c) to compensate for the removal of the spectrum in step c10).

[0148] In the context of the present invention, the term modified prediction function is used to indicate that a new prediction function is generated on the clean spectral population due to the additional removal of spectra in step c10) above.

[0149] Furthermore, there may be cases where two or more spectra corresponding to the top-ranked distances to be removed from the clean spectrum group in step b7) are located in (close) neighborhoods in the scatter plot. However, removing two or more spectra from the (close) neighborhood can significantly affect the slope of the modified prediction function. In the worst case, the slope of the prediction function can be significantly affected, for example, by modifying the prediction function to cause the slope to be higher or lower than its previous value.

[0150] This effect of removing two or more spectra corresponding to the top-ranked distances from the clean spectrum group of step b7) when they are in (close) proximity in the scatter plot can be avoided or at least significantly reduced by weighting them. For this purpose, first, a perpendicular line is drawn through the angle bisector at half the length of the angle bisector, thereby dividing the plane above and below the angle bisector into two planes, respectively, where the planes mirrored on the angle bisector are equally sized. Typically, the plane between the axis of the scatter plot and the angle bisector and the perpendicular line drawn through the angle bisector at half its length is divided into four planes: two triangles of equal size and two squares of equal size, one triangle and one square above and below the angle bisector. Next, when one or more spectra corresponding to the top-ranked distances to be removed from the clean spectrum group in step b7) are present in the same plane given by the angle bisector and the perpendicular line, they are weighted by a factor of 0.5.

[0151] In another preferred embodiment, the method according to the invention therefore further comprises:

[0152] - in step c8), drawing a perpendicular line through the angle bisector at half the length of the angle bisector to divide the plane between the axes of the scatter plot into four planes, and

[0153] - In step c10), when the spectra corresponding to at least two top-ranked distances are present in the same plane, they are weighted with a factor of 0.5.

[0154] In principle, the above method can be performed until the prediction function and the angle bisector are congruent. In practice, steps c7) to c11) are repeated until there is no further improvement in the convergence of the prediction function to the angle bisector, or in other words, until the curve of the modified prediction function obtained as a regression line is as close as possible to the angle bisector.

[0155] In a further preferred embodiment, the method according to the present invention further comprises the following steps:

[0156] c12) Repeating steps c7) to c11) using the modified prediction function and the clean spectral population obtained from the previous correction, until the plot of the modified prediction function thus obtained as a regression line is as close as possible to the angle bisector.

[0157] It may be beneficial to perform data pre-processing on the spectra of step a) to facilitate subsequent analysis of the spectra.

[0158] In another embodiment, the method according to the present invention further comprises the following steps before step b):

[0159] a2) performing data preprocessing on the infrared spectrum of step a), wherein the data preprocessing is one or more selected from the group consisting of smoothing, multiplicative scatter correction, standard normal variate transformation, detrending, spectral differentiation, and continuous piecewise direct standardization.

[0160] Inherent in the collection of spectra over time is some form of random variation or noise. Smoothing is a suitable method for removing noise and better exposing the signal of the underlying causal processes. In the context of the present invention, smoothing is preferably moving average smoothing, polynomial smoothing with Savitzky-Golar filtering, and / or extended signal-to-noise ratio polynomial smoothing.

[0161] Calculating a moving average involves creating a new series whose values consist of the averages of the original observations in the original time series. A moving average requires specifying a window size, called the window width. This defines the number of original observations used to calculate the moving average. The "moving" part of moving average refers to sliding the window defined by the window width along the time series to calculate the average value in the new series.

[0162] The Savitzky-Golay filter is a digital filter that can be applied to a collection of digital data points to smooth the data—that is, to improve the accuracy of the data without distorting the signal's trend. This is achieved by fitting a low-degree polynomial to consecutive subsets of adjacent data points using a linear least-squares method, in a process called convolution. When the data points are equally spaced, an analytical solution to the least-squares equation can be found in the form of a single set of "convolution coefficients" that can be applied to all data subsets to give an estimate of the smoothed signal (or its derivative) at the center point of each subset.

[0163] Detrending is a type of baseline correction; it removes sloping baseline variations, typically present in NIR reflectance spectra of powder samples, by subtracting a linear or polynomial fit of the baseline from the raw spectrum.

[0164] The standard normal variate (SNV) is another commonly used preprocessing method because of its algorithmic simplicity and effectiveness in scatter correction. It is often used for spectra where baseline and pathlength variations cause differences between otherwise identical spectra.

[0165] Multiplicative scatter correction (MSC) is performed by regressing the measured spectrum against a reference spectrum and then correcting the measured spectrum using the slope and intercept of this linear fit. This preprocessing method has been shown to be effective in minimizing baseline shifts and multiplicative effects. In many cases, the results of MSC are very similar to SNV. Despite this, many spectroscopists prefer SNV to MSC because SNV corrects each spectrum individually and does not require the entire data set. For example, the extended multiplicative signal correction preprocessing method enables the separation of physical light scattering effects from chemical absorption effects in spectra from powders or turbid solutions. Model-based methods are particularly useful in minimizing wavelength-dependent light scattering variations. After preprocessing, the corrected spectra are insensitive to light scattering variations and respond linearly to analyte concentration.

[0166] In some cases, it is impossible to clearly identify the maximum and minimum values of each peak, for example due to overlap, making it impossible to locate the position of the signal peak in the spectrum. When the minimum and maximum values of the peak are easier to identify, it is easier to locate the individual peaks in the spectrum. Taking the first derivative of the spectrum helps to identify the peaks in the spectrum because it causes the peak maximum or peak minimum to have a zero crossing. Taking the second derivative will cause the peak minimum to appear exactly at that position, where the peak maximum is in the original spectrum, and vice versa. Obtaining the first or second derivative of the spectrum also helps to identify outliers in the spectrum group of known feed raw materials and / or feeds. Therefore, in the context of the present invention, it is preferred to adopt the first and / or second derivative of the spectrum.

[0167] Continuous Piecewise Direct Standardization (CPDS) is used to account for spectral variations caused by continuous external factors such as temperature. The first step of the PDS and CPDS algorithms is to calculate the spectral distribution of the spectrum X(t ref ) that is, the order of m×n and the specific temperature t ref The PLS model is constructed between the response matrix Y. According to the piecewise direct standardization (PDS) algorithm, the transformation is then calculated at t ref The matrix of spectral measurements recorded at temperature values other than :

[0168] X(t ref )=X(t k )×Q(t k ),(k=1,2,…,k;t k ≠t ref )

[0169] Where Q(t k) is the m×m band transformation matrix obtained from X(t k )(i.e. X(t k ) corresponds to the spectrum of the wavelength window from jw to j+k) on the submatrix X(t ref ) is obtained by linear regression of the jth column of , where the window size 2w+1 is determined by adjusting the parameter w. Q(t k ). For each temperature value t k ,(k=1,2,…,k;t k ≠t ref ) Repeat this process and mark Q(t k ) is the unit matrix to obtain the K transformation matrix. In order to consider the influence of continuous temperature values, the polynomial function is calculated for the temperature difference Δt k Fitting into a band matrix Q(t k ) non-zero elements, where Δt k =t k -t ref :

[0170]

[0171] where q ij (t k ) is the Q(t k ). Once the parameters of the polynomial function are estimated, the ij (t k ) is used to calculate the transformation matrix t of the spectrum measured at the invisible temperature value during the test phase test Therefore, the temperature effect can be removed by transforming the spectral matrix as if it were measured at the reference temperature:

[0172] X(t ref |t test )=X(t test )×Q(t test )

[0173] The response can then be predicted by a PLS model constructed at the reference temperature.

[0174] Preferably, the infrared spectrum of step a) is subjected to a standard normal variate transformation, detrending and taking the first derivative of the spectrum, preferably in the given order.

[0175] When only signals with wavelengths equidistant from one another in the spectrum are considered, it not only facilitates but, more importantly, also improves the prediction quality of the prediction function generated by the method according to the present invention. This approach significantly reduces the risk of missing any relevant signals and information in the spectrum. This is particularly relevant for concentration series of complex mixtures with multiple infrared-active components, where the signal intensity increases significantly with the concentration of each component. Therefore, the signal of a component with a low concentration in the complex mixture may be (partially) superimposed by the signal of another component with a higher concentration. In this case, ignoring or missing individual signals in the spectrum may already lead to erroneous data correlations and / or the generation of erroneous prediction functions, thereby leading to erroneous prediction results.

[0176] Therefore, in a further embodiment of the method according to the invention, signals which are equidistant in wavelength from one another in the spectrum are taken into account in steps b) and / or c).

[0177] If data preprocessing results in different wavelength scales for the spectrum before and after preprocessing, it is beneficial to perform a correction on the entire spectrum. This is particularly relevant to the correct location of the wavelengths of the extreme points, i.e., signals with relative absorption maxima or minima. Therefore, it is preferable to perform a correction on spectra with unequal wavelength spacing to obtain a correct spectrum with equidistant wavelength spacing. The correct wavelengths of the extreme points, i.e., the points of relative absorption maxima and minima, provide a good basis for correcting the entire spectrum. First, the position of one or more extreme points in the spectrum is shifted from the shifted (i.e., incorrect) wavelength to the correction wavelength where the corresponding signal is expected. For example, a specific signal from a functional group in a known compound is visible at a specific wavelength in the infrared spectrum. In the next step, a graph is plotted through the extreme points at the correction wavelength and, where appropriate, through other points in the spectrum whose positions are not affected by the wavelength shift. This can be accomplished by interpolation. However, plotting a single graph through all the extreme points in the spectrum to be corrected can be difficult or even impractical. Instead, better results can be achieved by using multiple polynomials to connect adjacent points in the spectrum to be corrected, as they can be smoothly combined into a single graph. This can be done by (cubic) spline interpolation. A (cubic) spline interpolation can be visualized as a curved flexible strip that passes through each extreme point in the spectrum whose graphical representation is to be interpolated. From a mathematical point of view, this strip can be described by a series of cubic solid polynomials that must satisfy the following requirements: i) when the graphical representation of the spectrum is considered as a function or a piecewise defined function, the spline obtained thereby must pass through all function values, ii) the first and second derivatives of the cubic polynomial are continuous, and iii) the curvature is forced to zero at the end points of the spectral interval. Once the correction wavelength of the extreme points is known, the cubic polynomial can be easily calculated by any appropriate commercial mathematical program. Therefore, the low requirement for computing power is another advantage of (cubic) spline interpolation. Spline interpolation or cubic spline interpolation is therefore very useful for interpolating data in a spectrum to a new wavelength and for generating a continuous spectrum in the corrected spectrum.

[0178] In yet another embodiment, the method according to the present invention further comprises the following steps before step b):

[0179] a3) Performing cubic spline interpolation on spectra with non-equidistant wavelength distances to provide spectra with equidistant wavelength distances.

[0180] Preferably, step a3) is performed after step a2) and before step b).

[0181] The method according to the present invention is not subject to any restrictions regarding the number of prediction functions to be generated. If a specific number of different spectral groups are provided in step a), a corresponding number of prediction functions will be generated for each group. Alternatively, if the infrared spectrum group of step a) includes infrared spectra of different sample classes, the method according to the present invention will generate a corresponding prediction function for each of these sample classes.

[0182] In the method according to the invention, in step d), an infrared spectrum of a sample of unknown origin and / or composition or of the same origin and / or composition as the sample in step a) is provided, and then, by means of a prediction function in step c), the value of the property of interest is predicted from said spectrum for the spectral sample in step d). To this end, it is necessary to select an appropriate prediction function that is suitable for the sample spectrum in question. In principle, the prediction function can be selected based on available information about the sample whose infrared spectrum is provided in step d), such as information about a bag of feed, feed, etc. However, this information may be incorrect, thus leading to the selection of an incorrect prediction function, or the user of the method may select a prediction function at his or her discretion, but this prediction function may be the wrong prediction function for the sample spectrum in question.

[0183] Therefore, it is preferred that the method according to the present invention select a prediction function that is most suitable for the infrared spectrum of the sample in question. One suitable method is to identify the similarity between the sample spectrum provided in step d) and any clean spectrum group or the centroid of the clean spectrum group in step b7). The first step of the method includes converting the absorption intensity of the wavelength or wavenumber in the spectrum of step d) into a query vector and comparing the query vector with a database vector obtained by converting the absorption intensity of the wavelength or wavenumber in the clean spectrum in step b7) into a set of database vectors, or by converting the absorption intensity or wavenumber of the wavelength or wavenumber of the centroid of the clean spectrum population in step b7) into a database vector. Preferably, the query vector is compared with each database vector in the database vector set. In its broadest sense, a vector is a geometric object with size (or length) and direction. In a Cartesian coordinate system, a vector can be represented by identifying the coordinates of its initial point and end point. Therefore, a vector is suitable for representing the absorption intensity of a specific wavelength or wavenumber in a two-dimensional near-infrared spectrum. In addition, vectors are not limited to describing two-dimensional systems. Instead, vectors can describe multidimensional spaces, such as near-infrared spectra with multiple absorption intensities at multiple different wavelengths or wavenumbers. In this case, each dimension of the vector corresponds to a single absorption intensity at a specific wavelength or wavenumber. The next step is to calculate a similarity measure and / or distance measure between the query vector and each of the database vectors or a set of database vectors, and to rank the similarity measures and / or distance measures with the best values obtained thereby, i.e., the measure with the highest similarity is ranked first or the measure with the lowest distance is ranked first. The prediction function corresponding to the clean spectrum group with the highest similarity to the database vector is then selected in step d4) for use in the prediction in step e).

[0184] Preferably, step d) of the method according to the invention therefore further comprises the following steps:

[0185] d1) converting the absorption intensity of the wavelength or wave number in the spectrum of step d) into a query vector,

[0186] d2) converting the absorption intensity of the wavelength or wavenumber in the clean spectrum group in step b7) into a set of database vectors, or converting the absorption intensity of the wavelength or wavenumber of the centroid of the clean spectrum group in step b7) into a database vector,

[0187] d3) calculating a similarity metric and / or a distance metric between the query vector of step d1) and the database vector of step d2) to provide a similarity value between the query vector and the database vector,

[0188] d4) arranging the similarity values obtained in step d3) in descending order when calculating the similarity metric in step d3) or in ascending order when calculating the distance metric in step d3), wherein in any case, the top-ranked database vectors have the highest similarity to the query vector, and

[0189] d5) selecting the prediction function corresponding to the group of clean spectra having the highest similarity to the database vector in step d4) for the prediction in step e).

[0190] Preferably, the absorption intensities at equidistant wavelengths or wavenumbers in the spectrum of step d) and in the population of clean spectra in step b7) are converted into vectors, in particular the absorption intensities at the same equidistant wavelengths or wavenumbers in the spectrum of step d) and in the clean spectra. This enables an optimal similarity analysis. Preferably, the distance between wavelengths in steps d1) and / or d2) of the method according to the present invention is 0.1+ / –10% to 10+ / –10% nm, 0.1+ / –10% to 5+ / –10% nm, or 0.1+ / –10% to 2+ / –10% nm. Correspondingly, the distance between wavenumbers in steps d1) and / or d2) of the method according to the present invention is 108+ / –10% to 10610+ / –10%, 108+ / –10 to 5*106+ / –10% nm, or 108+ / –10% to 2*106+ / –10% nm. In the context of the present invention, the term + / −10% is used for explicitly mentioned values to indicate that deviations from these explicitly mentioned values are still within the scope of the present invention as long as they essentially result in the effect of the present invention.

[0191] The similarity analysis in step d3) is not subject to any restriction on using a specific distance or similarity metric. In principle, the similarity analysis is a search for nearest neighbors. Therefore, any similarity metric suitable for determining the nearest neighbors of the query vector can be used in step d3).

[0192] Cosine similarity is an example of a similarity measure suitable for use in the context of the present invention. It enables the calculation of the similarity between two vectors very quickly and with high accuracy. The cosine similarity C between two observations is A,B , that is, two spectral vectors, can be calculated by the following formula:

[0193]

[0194] Among them A i and B iare the components of vectors A and B, where one vector is the query vector and the other vector is the database vector, and n is the number of vectors considered. The value of the similarity measure can range from -1 (indicating complete opposites of each other), through 0 (indicating orthogonality (decorrelation)) to +1 (indicating identity), with intermediate values indicating intermediate similarity or dissimilarity.

[0195] The Euclidean distance is an example of a distance metric suitable for use in the context of the present invention. It enables the dissimilarity between two vectors to be calculated very quickly and with high accuracy. If desired, the distance metric thus obtained can be easily calculated as a similarity metric for use in the method according to the present invention. The Euclidean distance E between two observations is A,B , that is, two spectral vectors, can be calculated by the following formula:

[0196]

[0197] Among them A i and B i are the components of vectors A and B, where one vector is the query vector and the other vector is the database vector, and n is the number of vectors considered. The value of the distance metric can range from +1 (indicating complete opposites of each other), through 0 (indicating orthogonality (decorrelation)) to –1 (indicating identity), with intermediate values indicating intermediate similarity or dissimilarity.

[0198] Another object of the present invention is a system for predicting the value of a property of interest of a material, the system comprising a processing unit adapted to execute the computer-implemented method for predicting the value of a property of interest of a material according to the present invention, and / or any embodiment thereof.

[0199] The one or more prediction functions generated in the method according to the invention are stored in the processing unit of the system. This enables the system according to the invention to communicate with one or more infrared spectrometers located at the same and / or different locations from the system.

[0200] In one embodiment of the system according to the invention, the processing unit forms a network with one or more infrared spectrometers.

Claims

1. A computer-implemented method for predicting a value of a property of interest of a material, the method comprising the steps of: a) providing a population of infrared spectra of a sample, wherein the spectra form an m×n input data matrix X, Where m is the number of samples in rows, and n is the number of data points in columns, b) removing spectral outliers from the spectral group of step a), comprising the following steps: b1) obtaining principal components by performing principal component analysis on the matrix X, b2) Generate a diagonal matrix Σ from the input data matrix X, which contains the singular values σ of the matrix X m and the loading matrix V, b3) calculating the score x for each spectrum by multiplying each data point of said input data matrix X by said loading of each component of step b2) m , forming the mean of each column of the X matrix to provide B 0,m The score index si is calculated by the following formula: b4) Determine the component N whose eigenvalue causes the regression of X to converge for at least 99% of the scores C The number of, and the distance metric threshold D of each spectrum in step a) is calculated by the following formula i : b5) calculating the average of all scores for each principal component of each spectrum of step a), and calculate the distance metric between the mean and each score of each principal component, b6) when the distance metric value of the score of the principal component obtained in step b5) is greater than the distance metric threshold of step b4), the sample spectrum is regarded as a spectral outlier, b7) removing the spectral outliers of step b6) from the spectral group of step a) to obtain a clean spectral group, c) generating a prediction function of the cleaned spectral population of step b7), d) providing an infrared spectrum of a sample of unknown origin and / or composition or a sample of the same origin and / or composition as the sample in step a), and e) predicting the value of the attribute of interest from the spectrum of step d) using the prediction function of step c).

2. The method according to claim 1, wherein a principal component analysis is performed on the input data matrix X, wherein the input data matrix X is decomposed into two mutually orthogonal matrices, U m×n Matrix and The matrix is described mathematically by: in, X m×n is the input data matrix, where m is the number of samples in row form, and n is the number of data points in column form, is the transposed matrix or so-called loading matrix, which contains the loadings, and U m×n is a unitary m×n matrix.

3. The method according to claim 1 or 2, wherein the input data matrix X is decomposed into the product of several matrices by singular value decomposition, which is mathematically described as: in, X m×n is the input data matrix, where m is the number of samples in row form, and n is the number of data points in column form, U m×m is a unitary m×m matrix, ∑ m×n is the matrix X containing the singular values σ m A real m×n matrix, and is the adjoint matrix of the unitary n×n matrix V containing the load; wherein the real number of the adjoint matrix is equivalent to the transposed matrix 4. The method of claim 1 or 2, wherein the number of spectra in the spectral group ranges from 50 to 10,000, from 50 to 5,000, from 50 to 2,500, from 50 to 2,000, from 50 to 1,500, from 50 to 1,000, from 100 to 1,000, from 50 to 500, from 100 to 500, from 50 to 250, from 100 to 250, or from 50 to 100. 5 . The method according to claim 1 , wherein the distance metric is a Euclidean distance metric, a Pearson distance metric, a Mahalanobis distance metric, or a distance metric obtained from a similarity metric.

6. The method according to claim 1 or 2, wherein step b) further comprises the following steps: b5.1) increasing the distance metric threshold obtained in step b4) by +1, b5.2) determining the two distance metrics obtained in step b5) with the highest values using the distance metric threshold obtained in step b5.1), b5.3) determining the difference between the distance metric values determined in step b5.2), and b5.4) repeating steps b5.1) to b5.3) with the distance metric threshold of step b5.1) until the difference determined in step b5.3) is at least 1 and the maximum value of the distance metric is at most 8.

7. The method according to claim 1 or 2, wherein the generation of the prediction function of the clean spectral population of step b7) involves using a reference data set obtained from determining the value of the attribute of interest in each sample of step a) in a quantitative analysis.

8. The method according to claim 1 or 2, further comprising the steps of: a1) determining the value of the attribute of interest in each sample of step a) in a quantitative analysis to provide a reference data set. 9 . The method according to claim 8 , wherein the generation of the prediction function in step c) comprises analyzing data correlation between the reference data set of step a1) and the clean spectrum group of step b7) to provide the prediction function.

10. The method according to claim 8, wherein the generation of the prediction function in step c) includes validation of the prediction function. The method according to claim 10 , wherein the validation is a leave-one-out cross validation, a k-fold validation and / or a validation on circular test data.

12. The method according to claim 11, wherein when the validation is a leave-one-out cross-validation, step c) of the method further comprises the following steps: cV1) dividing the reference dataset of step a1) and the clean spectrum group of step b7) into a training set S0 and a test set S1, wherein the size of the test set S1 is smaller than the size of the training set S0, cV2) generating a preliminary prediction function of the training set S0 of step cV1), cV3) predicting a value of a property of interest by applying the preliminary prediction function of step cV2) to the clean spectrum group, cV4) calculating the average prediction error of the preliminary function of step cV2) using the predicted attribute value of step cV3) and the corresponding reference data set of step cV1), cV5) repeating steps cV1) to cV5) when the average prediction error of the preliminary prediction function is outside the limit, or when the average prediction error of the preliminary prediction function is within the limit, Continue with step cV6), and cV6) Approving said preliminary prediction function of step cV2) as a prediction function.

13. The method according to claim 11, wherein when the validation is k-fold cross validation, step c) of the method further comprises the following steps: cV1) dividing the set of clean spectra from step b7) and the corresponding reference dataset from step a1) into n subsets of equal size, where n is an integer with a minimum value of at least 2 and a maximum value less than the number of spectra, one of the subsets serving as a test set and the remaining subsets serving as training sets, cV2) generating a preliminary prediction function of the training set of step cV1), cV3) predicting a value of a property of interest by applying the preliminary prediction function of step cV2) to the clean spectral population, cV4) calculating an average prediction error of the preliminary prediction function using the predicted attribute values of step cV3) and the corresponding reference data set of step cV1), cV5) repeating steps cV1) to cV5) when the average prediction error of the preliminary prediction function is outside the limit, or when the average prediction error of the preliminary prediction function is within the limit, Continue with step cV6), cV6) approving the preliminary prediction function of step cV2) as the prediction function, and performing steps cV2) to cV6) k times, using each of the n subsets once as a test set to give k prediction functions, and cV7) Averaging the parameters of each approved prediction function in step cV6) to give a calibration function.

14. The method according to claim 11, wherein when the verification is verification on ring test data, step c) of the method further comprises the following steps: cV1) providing a set of external reference spectra and corresponding reference data, cV2) generating a preliminary prediction function of the cleaned spectral group of step b7) and the reference data set of step a1), cV3) predicting a value of the property of interest by applying the preliminary prediction function of step cV2) to the external reference spectrum of step cV1), cV4) calculating an average prediction error of the preliminary prediction function of step cV2) using the predicted property values of step cV3) and the corresponding external reference data of step cV1), cV5) repeating steps cV1) to cV5) when the average prediction error of the preliminary prediction function is outside a limit, or when the average prediction error of the preliminary prediction function is within a limit, Continue with step cV6), and cV6) Approving said preliminary prediction function of step cV2) as a prediction function.

15. The method according to claim 9, wherein step c) further comprises the following steps: c1) plotting the predicted property values of the clean spectrum group of step b7) and the prediction function of step c) as a regression line into a measurement-prediction diagram to give a scatter plot, c2) drawing an angle bisector between the axes of the graph of step c1), c3) determining the distance of each predicted attribute value of step c1) from the angle bisector in the scatter plot of step c2), c4) forming an overall ranking of the distances obtained in step c3) with the highest distances ranked at the top, and removing spectra corresponding to at least three top-ranked distances from the group of clean spectra of step b7), and c5) Modifying the prediction function of step c) to compensate for the removal of the spectrum in step c4).

16. The method according to claim 15, wherein the modification of the prediction function is a new prediction function for the clean spectral population generated due to the additional removal of spectra in step c4).

17. The method according to claim 15, wherein the method further comprises: - in step c2), drawing a perpendicular line through the angle bisector at half the length of the angle bisector to divide the plane between the axes of the scatter plot into four planes, and - In step c4), when the spectra corresponding to at least two top-ranked distances are present in the same plane, they are weighted with a factor of 0.

5.

18. The method according to claim 15, further comprising the steps of: c6) Repeating steps c1) to c5) using the modified prediction function and the clean spectral population obtained from the previous correction, until said plot of the modified prediction function thus obtained as a regression line is as close as possible to said angle bisector.

19. The method according to claim 9, further comprising the steps of: c7) plotting the predicted property values of the clean spectrum group of step b7) and the prediction function of step c) as a regression line into a measurement-prediction graph to give a scatter plot, c8) drawing an angle bisector between two axes of the graph of step c7), c9) drawing a confidence interval having a predetermined width around the angle bisector into the graph of step c8), c10) removing spectra corresponding to points outside the confidence interval in the graph of step c9) from the cleaned spectrum group of step b7), and c11) Modifying the prediction function of step c) to compensate for the removal of the spectrum in step c10).

20. The method according to claim 19, wherein the modification of the prediction function is a new prediction function for the clean spectral population generated due to the additional removal of spectra in step c10).

21. The method of claim 19, wherein the method further comprises: - in step c8), drawing a perpendicular line through the angle bisector at half the length of the angle bisector to divide the plane between the axes of the scatter plot into four planes, and - In step c10), when the spectra corresponding to at least two top-ranked distances are present in the same plane, they are weighted with a factor of 0.

5.

22. The method according to claim 19, further comprising the steps of: c12) Repeating steps c7) to c11) using the modified prediction function and the clean spectral population obtained from the previous correction, until said plot of the modified prediction function thus obtained as a regression line is as close as possible to said angle bisector.

23. The method according to claim 1 or 2, further comprising the following steps before step b): a2) performing data preprocessing on the infrared spectrum of step a), wherein the data preprocessing is one or more selected from the following group: smoothing, multiplicative scatter correction, standard normal variate transformation, detrending, spectral derivative and continuous segmented direct standardization.

24. The method according to claim 23, wherein the infrared spectrum of step a) is subjected to standard normal variate transformation, detrending and taking the first derivative of the spectrum in a given order.

25. The method according to claim 1 or 2, further comprising the following steps before step b): a3) Performing cubic spline interpolation on spectra with non-equidistant wavelength distances to provide spectra with equidistant wavelength distances.

26. The method according to claim 25, wherein step a3) is performed after step a2) and before step b).

27. The method according to claim 1 or 2, wherein the method further comprises the following steps: d1) converting the absorption intensity of wavelengths or wavenumbers in the spectrum of step d) into a query vector, d2) converting the absorption intensity of wavelengths or wavenumbers in the clean spectrum group of step b7) into a set of database vectors, or converting the absorption intensity of wavelengths or wavenumbers of the centroid of the clean spectrum group of step b7) into a database vector, d3) calculating a similarity metric and / or a distance metric between the query vector of step d1) and the database vector of step d2) to provide a similarity value between the query vector and the database vector, d4) arranging the similarity values obtained in step d3) in descending order when the similarity metric is calculated in step d3), or in ascending order when the distance metric is calculated in step d3), wherein in either case, the top-ranked database vectors have the highest similarity to the query vector, and d5) Selecting the prediction function corresponding to the clean spectrum group having the highest similarity to the database vector in step d4) for prediction in step e).

28. The method of claim 27, wherein the distance between the wavelengths in steps d1) and / or d2) is 0.1 + / - 10% to 10 + / - 10% nm, 0.1 + / - 10% to 5 + / - 10% nm, or 0.1 + / - 10% to 2 + / - 10% nm.

29. The method according to claim 1 or 2, wherein the material whose property value of interest is to be predicted is an organic compound, an amino acid, a peptide, a protein, a carboxylic acid, a fatty acid, a saturated and / or unsaturated fatty acid, a polyunsaturated fatty acid or a mixture thereof, or a substance of human, animal or plant origin, a digest, a body part, an animal part or a plant part, or a product of technological production based on products of plant and / or animal origin, wet distiller's grains with solubles (DDGS), hydrolyzed fish meal or hydrolyzed feather meal, and / or an inorganic compound, carbonates, phosphates, nitrates and sulfates, or a mixture thereof.

30. The method according to claim 1 or 2, wherein the material whose attribute value of interest is to be predicted is feed and / or feed ingredients.

31. The method of claim 1 or 2, wherein the material for which the value of the property of interest is to be predicted is animal feed.

32. A system for predicting a property value of interest of a material, the system comprising a processing unit adapted to perform the computer-implemented method for predicting a property value of interest of a material according to claim 1.

33. The system of claim 32, wherein the processing unit is networked with one or more infrared spectrometers.

Citation Information

Patent Citations

  • Method for searching analog tobacco leaf based on tobacco leaf near infrared spectra

    CN101251471A

  • Near infrared spectrum quantitative model simplification method based on principal component analysis

    CN104165861A