Method for predicting feed and / or feed ingredients
By counting and weighting the similarity values of feed and feed raw materials after similarity analysis, the problem of inaccurate spectral matching in the prior art is solved, and higher precision feed and feed raw materials type identification is achieved.
Patent Information
- Application Number
- CN202080046434.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-24
- Filing Date
- 2020-06-23
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-06-23
AI Technical Summary
When using near-infrared spectroscopy to analyze feed and feed raw materials, the prior art is susceptible to misclassification and false positive results, which leads to inability to accurately match the spectra and reference spectrum of the sample substance, resulting in incorrect prediction results.
After similarity analysis, the similarity values of feed and feed raw materials are counted and weighted according to the sorting position to form the weighted ranking position of each feed and feed to obtain the final score, thereby determining the type of sample.
Improve the prediction accuracy of feed and feed raw materials, reduce false positive results, and ensure accurate classification and type identification of sample substances.
Smart Images

Figure BDA0003431188480000071 
Figure BDA0003431188480000072 
Figure BDA0003431188480000161
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting unknown types of feed ingredients and / or feeds by near-infrared spectroscopy and similarity analysis. Background Art
[0002] Animal feeds typically contain a variety of different feeds and / or feed ingredients. Therefore, it is necessary to know as precisely and quickly as possible the characteristics and types of feeds and / or feed ingredients. This is particularly important when different feeds and / or feed ingredients are to be mixed to produce a diet with a specific composition for a specific species. Qualitative analysis methods for feeds and feed ingredients in principle allow for the precise identification of unknown types of feeds and / or feed ingredients, i.e., unknown characteristics, origin, etc. However, these methods require high-cost and high-maintenance laboratory equipment. Other disadvantages of these methods are the high standards for the time required and the expertise and experience of the operators. In principle, near-infrared spectroscopy would be a suitable method for the identification and determination of feeds and / or feed ingredients. The article "Near-infrared Spectroscopy in Food Analysis" (Osborne B.G., Encyclopedia of Analytical Chemistry, Wiley & Sons, 2006, pages 1-14) gives a review of the application of near-infrared (NIR) spectroscopy in this field. The article states that due to different sample presentation techniques, wide application of NIR in food analysis is possible. There are techniques available for any type of liquid, slurry, powder or solid sample. However, when used as a routine method, near-infrared spectroscopy requires information on the characteristics and types of feeds and / or feed ingredients. However, human error in the selection of feeds and / or feed ingredients may have led to incorrect classification of feeds and / or feed ingredients with respect to their characteristics and form of existence. Based on the incorrect classification, the wrong calibration method will be selected for the near-infrared analysis of the components and their specific amounts in feeds and / or feed ingredients. Therefore, the data obtained from an incorrectly calibrated NIR spectrometer will be incorrect. Therefore, these data will mislead any other operating steps involving the corresponding feed ingredients and / or feeds.
[0003] One option to overcome this problem is to perform a similarity analysis of the near-infrared spectrum of the sample substance and the reference spectrum to identify a match, thereby determining what the sample substance is. The general principle of this method is described in WO 2016 / 141198 A1 and the article "Algorithms, Strategies and Application Process of Spectral Searching Methods" (Chu X.-L., Li J.-Y., Chen P., Xu Y.-P., Chinese Journal of Analytical Chemistry, 2014, 42(9), 1379-1386). Specifically, the similarity analysis involves calculating a similarity measure or a distance measure between the spectrum of the sample substance and the reference spectrum. A high value of the similarity measure between the spectrum of the sample and the reference spectrum indicates a high similarity between the sample substance and the reference substance corresponding to the reference spectrum. By comparison, when the similarity analysis involves a distance measure, a low similarity value indicates a high similarity between the sample substance and the reference substance corresponding to the reference spectrum. The similarity values obtained in the similarity analysis are usually ranked using the reference substance of the highest ranked entry having the greatest similarity to the sample substance.
[0004] The article "Spectral Library Searching: Mid-Infrared Versus Near-Infrared Spectra for Classification of Powdered Food Ingredient" (Reeves J.B. et al., Applied Spectroscopy, 1999, vol. 53, no. 7, pp. 836-844) discloses a method for predicting food ingredients, which includes the step of providing a near-infrared (NIR) spectrum of a sample of a powdered food ingredient and constructing and searching a spectral library based on full spectrum algorithms. The results of various searches are expressed as the distance between the unknown or test sample and the samples in the algorithm library. When the sample closest to the test sample (the minimum distance) is a sample from the same group, the match is considered successful.
[0005] However, when false positive entries are at the top of the sorted list of similarity values, the prior art methods are very susceptible to false determination or prediction. The reason for false positive entries in the said sorting may be the incorrect assignment of spectra to the wrong reference substance or wrong reference substance category, the heterogeneity or confusion of the reference substance categories to which the spectra are recorded, or the similarity of some reference substances to each other, which makes exact matching quite difficult. Any of these situations makes it difficult or impossible to accurately and reliably match the spectra of samples of sample substances with the spectra of reference substances.
[0006] The article "Novel Search Algorithms for a Mid-Infrared Spectral Library of Cotton Contaminants" (Loudermilk J.B. et al, Applied Spectroscopy, June 2008, vol.62, no.6, pp. 661 - 670) discloses voting scheme algorithms used in the search of the MIR library of cotton contaminants. Counting is performed in the so-called "group" algorithms, such as the number of times the substance / category "seed coat" appears in the hit list. The "weighted frequency" algorithm sorts individual spectra and their occurrences in the hit lists of multiple search algorithms, and then sums them to obtain a score. Therefore, the said "weighted frequency" algorithm is applicable to multiple search algorithms for a single spectrum. Summary of the Invention
[0007] According to the present invention, the method for solving this problem is to first perform a similarity analysis, and then count the occurrences of feed materials and / or feeds in the sorted list of similarity values. Next, the number of similarity values of the thus determined feed materials and / or feeds is weighted according to their sorting positions to give weighted sorting positions, and the weighted sorting positions are summed to obtain a score for the feed materials and / or feeds, and the highest score represents the feed materials and / or feeds having the greatest similarity to the sample substance.
[0008] Therefore, the object of the present invention is a computer-implemented method for predicting feeds and / or feed materials, the method comprising the following steps:
[0009] a) providing a near-infrared (NIR) spectrum of a sample of an unknown feed material and / or feed,
[0010] b) converting the absorption intensity of wavelengths or wave numbers in the spectrum of step a) into a query vector,
[0011] c) Provide a set of database vectors of spectral groups of known feed ingredients and / or feeds, wherein the spectral groups of the known feeds and / or feed ingredients in step c) include at least 50 spectra for each sample of each feed and / or feed ingredient from each of its global growing regions,
[0012] d) Analyze the similarity between the query vector in step b) and the set of database vectors in step c), including the following steps:
[0013] d1) Calculate a similarity measure and / or a distance measure between each database vector in step c) and the query vector in step b) to give a similarity value for each database vector and the query vector,
[0014] d2) When calculating the similarity measure in step d1), sort the similarity values obtained in step d1) in descending order, or when calculating the distance measure in step d1), sort them in ascending order, where the database vector ranked first has the greatest similarity to the query vector,
[0015] d3) Count the number of occurrences of each feed ingredient and / or feed in the top-ranked database vectors in the sorting of step d2), where the number of occurrences is represented by the variable N,
[0016] d4) Weight the top N similarity values for each feed ingredient and / or feed according to its position in the sorting of step d2) to give a weighted ranked position for each feed ingredient and / or feed,
[0017] d5) Form the sum of the weighted ranked positions of each feed ingredient and / or feed in step d4) to give a score for each feed ingredient and / or feed, and
[0018] e) Assign the feed ingredient and / or feed of the database vector with the highest score to the sample in step a).
[0019] In the context of the present invention, the terms unknown feed ingredient and / or feed refer to any kind of feed and / or feed ingredient whose characteristics, composition, origin, and / or form, i.e., whether it is ground or unground, are unknown. By comparison, in the context of the present invention, the terms known feed ingredient and / or feed refer to any kind of feed and / or feed ingredient whose characteristics, composition, origin, and / or form, i.e., whether it is ground or unground, are known. Thus, the spectral group of a known feed ingredient and / or feed is a plurality or multitude of spectra that are known to belong to a specific feed and / or feed ingredient with known characteristics, composition, origin, and / or form.
[0020] The weighting of the similarity values in step d4) is typically carried out by taking the reciprocal ranking position of a specific feed ingredient and / or feed in the ranking of step d2). For example, if a specific feed ingredient appears at position 2 in the ranking, its weighting gives a weighted ranking position value of 1 / 2. The top N similarity values of each feed ingredient and / or feed are weighted according to their positions in the ranking of step d2) to give the weighted ranking position of each feed ingredient and / or feed in step d4). Next, in step d5), the sum of the weighted ranking positions of step d4) is formed for each feed ingredient and / or feed to give the score of each feed ingredient and / or feed.
[0021] In its broadest sense, a vector is a geometric object having a magnitude (or length) and a direction. In a Cartesian coordinate system, a vector can be represented by identifying the coordinates of its starting and ending points. Thus, a vector is suitable for representing the absorption intensity at a specific wavelength or wavenumber in a two-dimensional near-infrared spectrum. In addition, a vector is not limited to the description of a two-dimensional system. Instead, a vector can describe a multi-dimensional space, such as a near-infrared spectrum having multiple absorption intensities at multiple different wavelengths or wavenumbers. In this case, each dimension of the vector corresponds to a single absorption intensity at a specific wavelength or wavenumber. In the context of the present invention, a database vector also contains the characteristics of its feed ingredient and / or feed and optionally its source or other information, such as whether it is ground or unground, or the season at harvest, in the case of corn. Alternatively, each database vector has an identification number, and the aforementioned information is stored under this identification number on a processing unit such as a computer, cloud, or server.
[0022] In one embodiment of the computer-implemented method according to the present invention, the vectors in steps b) and c) are multi-dimensional vectors, where each dimension corresponds to the absorption intensity at a specific wavelength or wavenumber.
[0023] Similar to the query vector of the spectrum of an unknown feed ingredient and / or feed, a set of database vectors provided in step c) of the computer-implemented method according to the present invention is also obtained by converting each spectrum of a group of spectra of known feed ingredients and / or feeds into a corresponding vector. If the set of database vectors does not exist, step c) also includes converting each spectrum of a group of spectra of known feed ingredients and / or feeds into a corresponding vector to give a set of database vectors. In this case, step c) of the computer-implemented method according to the present invention includes the steps of: converting each spectrum of a group of spectra of known feed ingredients and / or feeds into a corresponding vector to give a set of data set vectors, and providing the obtained set of database vectors of the group of spectra of known feed ingredients and / or feeds.
[0024] According to the present invention, in step a), a near-infrared spectrum of a sample of an unknown feed material and / or feed is provided. In the context of the present invention, this means that the location where the spectrum to be provided is recorded and the location where the computer-implemented method according to the present invention is performed can be different or the same. For example, the near-infrared spectrum of a sample of an unknown feed material and / or feed can be recorded at one location and sent in any way to a remote location where the computer-implemented method according to the present invention is performed. Alternatively, based on the spectrum, the recording of the spectrum and the prediction of the feed material and / or feed can be performed at the same location.
[0025] In one embodiment of step a) of the computer-implemented method described above, this step includes recording a near-infrared spectrum of a sample of an unknown feed material and / or feed.
[0026] In principle, the near-infrared (NIR) spectrum in step a) can be recorded at a wavelength of 700 to 2,500 nm using any suitable near-infrared spectrometer, which operates either as a monochromator or based on the Fourier transform principle. However, it has been found that in the computer-implemented method according to the present invention, only the near-infrared spectral portion from 1,100 to 2,500 nm is necessary for the prediction of feed materials and / or feeds. Therefore, preferably, in step a) of the computer-implemented method according to the present invention, the spectrum is recorded in the range of 1,100 to 2,500 nm. Accordingly, the spectral group of known feed materials and / or feeds in computer-implemented step c) only needs to cover the range of 1,100 to 2,500 nm. Since the wavelength can be easily converted into the respective wave number, the near-infrared spectrum in step a) and / or the spectrum in step c) can also be recorded at the corresponding wave number. When the sample of the feed material and / or feed in step a) is not translucent, the reflectance of the emitted light from the sample is measured, and the difference between the emitted light and the reflected light is given as absorption. The absorption thus obtained, i.e., their intensity and their wavelength or wave number, is used to generate vectors in step b) and / or step c). Accordingly, the near-infrared spectrometer applicable to the method according to the present invention can operate in transmission mode or reflection mode.
[0027] In an embodiment of the computer-implemented method according to the present invention, the spectrum in step a) and / or step c) is recorded in the range of 1,100 to 2,500 nm.
[0028] The method according to the present invention is not limited to a specific distance or similarity measure for analyzing the similarity between the query vector in step b) and the database vector in step c). Thus, any distance or similarity measure suitable for determining the similarity between the vector in step b) and the vector in step c) can be used in the method according to the present invention. In principle, the similarity analysis is based on nearest neighbor search. It has been found that the cosine coefficient is a particularly suitable similarity measure for nearest neighbor search in the method according to the present invention. For example, the cosine coefficient is particularly suitable for the method according to the present invention, which allows the similarity between two vectors to be calculated extremely quickly with high precision. The cosine coefficient S of two vectors A and B A,B is represented by the following formula:
[0029]
[0030] where x jA and x jB are the components of vectors A and B respectively, and n is the number of spaces, here the number of absorption intensities at a specific wavelength or wave number. The value range of the similarity is from -1 (meaning exactly opposite to each other) to 1 (meaning the same), where 0 represents orthogonality (decorrelation), and intermediate values represent intermediate similarity or dissimilarity.
[0031] Alternatively, the similarity between vectors can also be calculated by a distance measure. For example, the Euclidean distance also allows the similarity between two vectors to be calculated very quickly and precisely, and it is particularly suitable for the method according to the present invention. The Euclidean distance D of two vectors A and B A,B is represented by the following formula:
[0032]
[0033] where x jA and x jB are the components of vectors A and B respectively, and n is the number of spaces, here the number of absorption intensities at a specific wavelength or wave number.
[0034] Therefore, preferably, in step d1) of the computer-implemented method according to the present invention, the similarity measure is the cosine coefficient and the distance measure is the Euclidean distance.
[0035] According to step b) of the present invention, the absorption intensity at the wavelength or wavenumber in the spectrum is converted to give a query vector. In principle, the strongest and thus most significant absorption intensity in the spectrum can be selected, and only the said absorption intensity is converted to give a vector. However, this would require a thorough analysis of each individual spectrum of the sample substance, which is not only time-consuming but also requires a good knowledge of near-infrared spectra. Therefore, this method is not suitable for routine analysis. In addition, the disadvantage of this method is that the meaningful but relatively weak absorption intensities in the spectrum may be ignored, resulting in loss of information. This may ultimately lead to incorrect assignment of unknown feed materials and / or feeds. Therefore, it is preferable to consider as much information as possible in the spectrum without the aforementioned in-depth analysis of the spectrum. Therefore, it is preferred to convert the absorption intensities at equidistant wavelengths or wavenumbers in the spectrum, i.e., in step b) and / or c), to give a vector of the said spectrum. In order to allow the best possible similarity analysis between the query vector and the database vector, it is preferred to convert the absorption intensities at equidistant wavelengths or wavenumbers in the spectra of step b) and step c) of the method according to the present invention to give a vector of the said spectrum. Preferably, the distance between the absorption intensities converted to a vector in step b) is the same as the distance between the absorption intensities converted to a vector in step c). This allows for a higher precision prediction of the computer-implemented method according to the present invention, even without any specific knowledge of the sample substance and its spectrum.
[0036] In an embodiment of step b) and / or step c) of the computer-implemented method according to the present invention, the absorption intensities at equidistant wavelengths and / or wavenumbers in the spectrum are converted to give a vector of the spectrum of step a) and / or step b).
[0037] In another embodiment of the computer-implemented method according to the present invention, the distance between the absorption intensities converted to a vector in step b) is the same as the distance between the absorption intensities converted to a vector in step c).
[0038] Preferably, the absorption intensities at the wavelengths or wavenumbers in the spectrum that are converted to a vector of the said spectrum are at a small distance from each other. This has the advantage that most, if not all, of the relevant absorption intensities, i.e., information, of the spectrum is converted to a vector of the said spectrum. It is believed that this allows for a very precise conversion of all relevant information of the spectrum into a vector, even without knowing the feed and / or feed material for which the spectrum was recorded, in particular its properties, composition, origin, and / or form. Preferably, in step b) of the method according to the present invention, the distance between the wavelengths is 0.1 + / – 10% to 10 + / – 10% nm, 0.1 + / – 10% to 5 + / – 10% nm, or 0.1 + / – 10% to 2 + / – 10% nm. Therefore, in step b) of the method according to the present invention, the distance between the wavenumbers is 10 8 + / – 10% to 10 6+ / –10%, 10 8 + / –10% to 5×10 6 + / –10% nm, or 10 8 + / –10% to 2×10 6 + / –10% nm. In the context of the present invention, the term + / –10% is used relative to a value explicitly mentioned to indicate that the deviation from the explicitly mentioned value is still within the scope of the present invention, provided that they substantially result in the effects of the present invention. In step c) of the method according to the present invention, the distance between the wavelengths or wave numbers is preferably the same as that in step b) to provide the best possible comparison between the recorded spectrum of the unknown feed and / or feedstuff and the spectrum of the known feed and / or feedstuff.
[0039] In an embodiment of the computer-implemented method according to the present invention, the distance between the wavelengths or wave numbers in step b) and / or step c) is 0.1 nm + / –10% to 10 nm + / –10% or 10 8 cm –1 + / –10% to 10 6 cm –1 + / –10%.
[0040] In principle, the computer-implemented method according to the present invention is not subject to any limitation regarding the number of absorption intensities to be converted to give a vector. On the contrary, the amount of relevant information in the spectrum of the feedstuff and / or feed depends to a large extent on the individual feedstuff and / or feed, in particular on its composition and components. The more complex the feed and / or feedstuff, i.e., the more components the feed and / or feedstuff contains, the more information is required to predict the unknown feed and / or feedstuff from the near-infrared spectrum. In addition, it is not practical to perform an in-depth analysis to find out the absorption intensities that must be converted to give a vector. A suitable choice for the number of absorption intensities to be converted to give a vector is to associate them with the distance between the corresponding wavelengths or wave numbers, for example, 0.1 nm + / –10% to 10 nm + / –10% or 0.1 + / –10% to 2 + / –10% nm, and the recording range of the spectrum, for example, 1,100 to 2,500 nm. Preferably, the number of absorption intensities in each spectrum to be converted into a vector is at least 100, and in particular, the number is in the range of 150 + / –10% to 15,000 + / –10% or 700 + / –10% to 15,000 + / –10%.
[0041] In another embodiment of the method according to the present invention, the number of absorption intensities in each spectrum converted into a vector is 100 + / –10% or more.
[0042] In principle, apart from the number of entries in the ranking of step d2), step d3) of the method according to the invention is not subject to any limitation regarding the number of vectors of the highest ranking. However, when step d3) involves using a relatively low number of database vectors, the accuracy of the prediction in the method according to the invention is neither affected nor improved in any way compared to when step d3) involves using a relatively high number of database vectors. On the contrary, when d3) involves using a larger number of database vectors in the ranking of step d2), this will lead to an increase in the number of database vectors that are less meaningful, apart from the database vectors actually being of the highest ranking, which will reduce the accuracy of the prediction in the method according to the invention. Therefore, preferably, the number N of the first ranked database vectors in step d3) is in the range of 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30 or 10 to 20. This results in a very accurate prediction equivalent to using a large number of database vectors, however, without unnecessarily increasing the computational effort as in the case of more database vectors.
[0043] In one embodiment of the method according to the invention, the number of the highest ranked database vectors considered in step d3) is in the range of 10 to 100, preferably in the range of 10 to 50.
[0044] The spectral groups of the known feed materials and / or feeds used in the method according to the invention are not limited to specific feed materials and / or feeds. On the contrary, it preferably encompasses the range of all feed materials and / or feeds for animal nutrition, preferably for the nutrition of poultry, pigs, pigs in aquaculture and / or animals such as fish and / or crustaceans. The spectra of the feed materials and / or feeds can vary significantly depending on their form or appearance, for example when they are in ground or unground form. Therefore, the spectral groups preferably also include the spectra recorded from the above-mentioned ground and / or unground feed materials and / or feeds.
[0045] In another embodiment of the method according to the invention, the spectral groups of the known feeds and / or feed materials in step c) of the method include the spectral groups of all feeds and / or feed materials in ground and / or unground form used in animal nutrition.
[0046] In principle, the method according to the invention does not limit in any way the quantity and type of feed and / or feedstuffs, the spectra of which, in ground and / or unground form, form a spectral group. Nevertheless, it is preferred that the spectral group of known feed and / or feedstuffs in step c) of the method comprises the spectra of all feed and / or feedstuffs in ground and / or unground form for animal nutrients, preferably for the nutrition of pigs and / or animals such as fish and / or crustaceans reared in aquaculture. Particularly preferably, the feed and / or feedstuffs are unprocessed and / or processed feedstuffs and / or feeds. Processed feedstuffs and / or feeds are those that have been subjected to any type of heat or pressure treatment in order to remove or detoxify antinutritional factors. Preferred feed and / or feedstuffs are oilseeds, in particular soybean meal and press cake, full-fat soybeans, rapeseed meal and press cake, cottonseed meal, peanut meal, sunflower meal, coconut meal and / or palm kernel meal; legumes, in particular roasted guar gum powder; brewing and distillation by-products, in particular dried distillers grains with solubles (DDGS), by-products of cereal processing and feed production, in particular corn gluten, corn germ meal and / or baking by-products; animal by-products, in particular fish meal, meat meal, poultry meal, blood meal and / or bone meal; and any type of grain. In particular, the feedstuff is soybeans, soya beans or soybean products.
[0047] Depending on factors such as climate, soil and plant genetics, the composition and the content of the composition of feedstuffs and / or feeds from different growing regions of the world may vary. In order to allow a reliable and reproducible prediction of feed and / or feedstuffs, the feedstuffs and / or feeds, the spectra of which are part of the spectral group, thus originate from all their global growing regions.
[0048] The number of spectra of the spectral group of feed materials and / or feeds used in the method according to the invention should be representative to allow reliable and reproducible prediction of the feed materials and / or feeds in question. Thus, the spectral group includes at least 50 spectra of samples of each feed and / or feed material from each of its global growing regions, i.e., each feed and / or feed material in ground and / or unground form is used in animal nutrients, preferably in the nutrients of poultry, pigs, pigs reared in aquaculture, and / or animals such as fish and / or crustaceans. The method according to the invention is not limited by any number of spectra of samples of feeds and / or feed materials from their global growing regions. Thus, the number of spectra of samples of feeds and / or feed materials from their global growing regions can be in the range of 50 to 10,000, 50 to 5,000, 50 to 2,500, 50 to 2,000, 50 to 1,500, 50 to 1,000, 100 to 1,000, 50 to 500, 100 to 500, 50 to 250, 100 to 250, or 50 to 100.
[0049] When the spectral group of feeds and / or feed materials of a known type takes into account each global growing region of the feed and / or feed material and the number of spectra from each global growing region is representative, the method according to the invention allows not only reliable and reproducible prediction of the feed materials and / or feeds in question, but also prediction of the origin of the feed materials and / or feeds in question.
[0050] Thus, the spectral group of the known feeds and / or feed materials in step c) preferably includes at least 50 spectra of samples of each feed and / or feed material from each of its global growing regions. The number of spectra of samples of feeds and / or feed materials from each global growing region is not limited by any means. Thus, the number of spectra of samples of any feed and / or feed material from each global growing region can be in the range of 50 to 10,000, 50 to 5,000, 50 to 2,500, 50 to 2,000, 50 to 1,500, 50 to 1,000, 100 to 1,000, 50 to 500, 100 to 500, 50 to 250, 100 to 250, or 50 to 100.
[0051] The spectral group of known feeds and / or feed materials can be an existing database, preferably a database that is continuously updated, and / or it can be a new database that is continuously updated or further developed. A spectral database of feeds and / or feed materials of a known type suitable for the method according to the invention is, for example, Evonik's 5.0.
[0052] It may occur that the position of the signal peak in the spectrum cannot be located in step b) and / or step c) of the method according to the invention, because the maximum and minimum values of the individual peaks cannot be clearly identified in such a spectrum. When the minimum and maximum values of the peaks are more easily identifiable, the individual peaks in the spectrum can be more easily located. Taking the first derivative of the spectrum facilitates the identification of the peaks in the spectrum, because it gives the zero crossing of the maximum value of the peak or the minimum value of the peak. Taking the second derivative gives the minimum value of the peak exactly at that position, where the maximum value of the peak is in the original spectrum, and vice versa. Taking the first derivative or the second derivative of the spectrum also helps to identify outliers in the spectrum group of known feed materials and / or feeds.
[0053] In another embodiment of the method according to the invention, derivatives of the spectrum of the unknown feed material and / or feed in step a) and / or the spectrum of the known feed material and / or feed in step c) are formed.
[0054] Preferably, the first derivative of the spectrum of the unknown type of feed and / or feed material in step a) and / or the spectrum of the known feed and / or feed material in step c) is formed.
[0055] Preferably, a set of database vectors of the spectrum group of known feed materials and / or feeds is directly provided in step c) of the computer-implemented method according to the invention. Nevertheless, it is also possible to first provide only the spectrum group of known feed materials and / or feeds, which is converted into a set of database vectors for the similarity analysis in step d) in the next step. In this case, step c) of the computer-implemented method according to the invention also includes the step of converting the absorption intensity of the wavelength or wave number in each spectrum of the spectrum group of known feed materials and / or feeds into a vector. As described above, the multiple vectors of the spectrum group of known feed materials and / or feeds thus obtained are then a set of database vectors. In any case, it is preferred to store the spectrum group of known feed materials and / or feeds or the set of database vectors of the spectrum group of known feed materials and / or feeds on a processing unit, such as a computer or the cloud. The processing unit on which the spectrum group or the database vectors are stored can be the same as or different from the processing unit that executes the computer-implemented method according to the invention. In the second case, the first processing unit that executes the computer-implemented method according to the invention and the second processing unit on which the spectrum group or the database vectors are stored form a network. For example, the spectrum group or the database vectors can also be stored on the cloud. In this case, the first processing unit, such as a computer, that executes the computer-implemented method according to the invention and the second processing unit, such as the cloud, that stores the spectrum group or the database vectors form a network.
[0056] Accordingly, another object of the present invention is to provide a system for predicting feed ingredients and / or feeds, which comprises i) a processing unit adapted to perform the computer-implemented method according to the present invention.
[0057] In the case where the spectral groups of known feed ingredients are stored on a computer, preferably, the spectral groups are stored on a second computer, i.e., a computer on which the computer-implemented method according to the present invention has not yet been performed. Thus, the workload is evenly distributed among the computers, which can result in faster performance of the computer-implemented method according to the present invention. This also allows communication between the user and the database provider with the database vectors used in the method according to the present invention, such as the update of the database vectors.
[0058] In an embodiment of the system according to the present invention, a processing unit adapted to perform the computer-implemented method according to the present invention forms a network with at least one other processing unit, and the database vectors are stored on the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 The computer-implemented method according to the present invention is shown. Starting from a query vector of the spectrum of an unknown feed ingredient and / or feed, a similarity value is calculated for each database vector of the spectral groups of known feed ingredients and / or feeds. The similarity values thus obtained are either increasing, with the highest at the top if the similarity values are calculated using a similarity metric, or decreasing, with the lowest at the top if the similarity values are calculated using a distance metric, to give a ranking. The top n ranked vectors representing the most adjacent to the query vector in the similarity search are selected from the ranking, and the occurrences of the corresponding categories, i.e., feed ingredients and / or feeds, are determined. The categories are weighted and added according to the ranking of the vectors of the categories to give a category score. The category scores are ranked according to the values having the highest values at the top. The feed ingredients and / or feeds of the category ranked highest have the highest similarity to the query vector, and thus, the samples assigned to the query vector are determined. EXAMPLES
[0060] The following are examples to illustrate the computer-implemented method according to the present invention in comparison with a simple prediction method of the prior art and a prediction method involving majority voting.
[0061] In the first step, the NIR spectrum of a sample of feed ingredient FRM3 is recorded. The relevant information of the spectrum, i.e., the absorption intensity of the wavelength, is converted to give the query vector of the spectrum. In the next step, a distance metric is used to calculate the similarity values of all database vectors with the query vector. The similarity values thus obtained are arranged in descending order according to their values including the indication of the corresponding feed ingredients and / or feeds to give the ranking of the database vectors.
[0062] In a simple prediction model according to the prior art (hereinafter also referred to as prediction method 1: 1-NN), the first n vectors in this ranking represent the nearest neighbors in the similarity analysis between the query vector of the sample of the unknown feed material and / or unknown feed and the database vectors of the spectral groups of the known feed materials and / or feeds.
[0063] The method with majority voting also starts from the ranking of the above database, but also includes the occurrence count of the material situation, that is, the feed materials and / or feeds corresponding to the database vectors.
[0064] Finally, the method according to the present invention is used, which includes majority voting and weighting of the results obtained therefrom.
[0065] Table 1: Ranking of database vectors according to their ranking, and description of their corresponding types of feed materials and / or feeds, namely FRM1, FRM2 or FRM3
[0066] Sorting Similarity value Vector name Feed raw material / feed 1 0.93 V1 FRM1 2 0.91 V3 FRM3 3 0.74 V5 FRM3 4 0.69 V4 FRM3 5 0.42 V7 FRM2 6 0.33 V2 FRM3 7 0.27 V12 FRM2 8 0.15 V6 FRM2 9 0.04 V8 FRM1 10 0.03 V9 FRM1 11 0.02 V10 FRM2 12 0.01 V11 FRM2
[0067] In prediction method 1 (1-NN), the feed material with the highest similarity value is ranked according to the ranking in Table 1, that is, FRM1 is indicated as the feed material with the highest similarity to the material of the query vector.
[0068] By comparison, in prediction method 2 (majority voting), the feed material that appears the largest number of times in the ranking of Table 1, that is, relative to FRM1 with 3 votes and FRM2 with 4 votes, FRM2 with 5 votes is indicated as the feed material with the highest similarity to the material of the query vector.
[0069] In prediction method 3, that is, the method according to the present invention, the feed material FRM3 is indicated as the feed material with the highest similarity to the material of the query vector.
[0070]
[0071]
[0072]
[0073] The results of these three methods are extremely different. More importantly, both test methods also give false positive results. Only the method according to the present invention gives the correct result.
Claims
1. A computer-implemented method for predicting feeds and / or feed ingredients, comprising the steps of: a) providing a near-infrared (NIR) spectrum of a sample of an unknown feed ingredient and / or feed; b) transforming the absorption intensity of wavelengths or wave numbers in the spectrum of step a) to give a query vector; c) providing a set of database vectors of spectral groups of known feed ingredients and / or feeds, wherein c) the spectral groups of the known feeds and / or feed ingredients include at least 50 spectra of samples of each feed and / or feed ingredient from each of its global growing regions; d) analyzing the similarity between the query vector of step b) and the set of database vectors of step c), including the steps of: d1) calculating a similarity measure and / or a distance measure between each database vector of step c) and the query vector of step b) to give a similarity value of each database vector to the query vector; d2) sorting the similarity values obtained in step d1) in descending order when calculating the similarity measure in step d1), or in ascending order when calculating the distance measure in step d1), wherein the database vectors ranked higher have the greatest similarity to the query vector; d3) counting the number of occurrences of each feed ingredient and / or feed in the database vectors ranked higher in the sorting of step d2), wherein the number of occurrences is represented by the variable N for each feed ingredient and / or feed; d4) weighting the top N similarity values for each feed ingredient and / or feed according to its position in the sorting of step d2) to give a weighted ranking position for each feed ingredient and / or feed, wherein the weighting of the similarity values is performed by taking the reciprocal of the sorting position of the feed ingredient and / or feed in the sorting of step d2); d5) forming the sum of the weighted ranking positions of each feed ingredient and / or feed in step d4) to give a score for each feed ingredient and / or feed, and e) assigning the feed ingredient and / or feed of the database vector with the highest score to the sample of step a).
2. The computer-implemented method according to claim 1, wherein the vectors in steps b) and c) are multi-dimensional vectors, wherein each dimension corresponds to the absorption intensity at a specific wavelength or wave number.
3. The computer-implemented method according to claim 1 or 2, wherein the spectra in step a) and / or step c) are recorded in the range of 1,100 to 2,500 nm.
4. The computer-implemented method according to claim 1 or 2, wherein in step d1), the similarity measure is the cosine coefficient and the distance measure is the Euclidean distance.
5. The computer-implemented method according to claim 1 or 2, wherein the absorption intensities at equidistant wavelengths and / or wave numbers in the spectrum are transformed to give the vectors of the spectra of step a) and / or step b).
6. The computer-implemented method according to claim 1 or 2, wherein the distance of the absorption intensity transformed into a vector in step b) is the same as the distance of the absorption intensity transformed into a vector in step c).
7. The computer-implemented method according to claim 1 or 2, wherein the absorption intensities of the wavelengths or wave numbers in the spectrum that are converted to give the vectors of the spectrum have small distances from each other.
8. The computer-implemented method according to claim 1 or 2, wherein the distance between the wavelengths in step b) and / or step c) is from 0.1 nm ± 10% to 10 nm ± 10%.
9. The computer-implemented method according to claim 1 or 2, wherein the distance between the wave numbers in step b) and / or step c) is 10 8 cm –1 + / – 10% to 10 6 cm –1 + / – 10%.
10. The computer-implemented method according to claim 1 or 2, wherein the distance between the wavelengths in step b) is from 0.1 nm ± 10% to 5 nm ± 10%.
11. The computer-implemented method according to claim 1 or 2, wherein the distance between the wave numbers in step b) is 10 8 cm –1 + / – 10% to 5 × 10 6 cm -1 + / – 10%.
12. The computer-implemented method according to claim 1 or 2, wherein the distance between the wavelengths or wave numbers in step c) is the same as the distance in step b) to provide the best possible comparison between the recorded spectrum of the unknown feed and / or feed ingredient and the spectrum of the known feed and / or feed ingredient.
13. The computer-implemented method according to claim 1 or 2, wherein the number of absorption intensities in each spectrum that are converted to vectors is 100 ± 10% or more.
14. The computer-implemented method according to claim 1 or 2, wherein the number N of the first ranked database vectors in step d3) is from 10 to 100.
15. The computer-implemented method according to claim 1 or 2, wherein the number of the ranked database vectors to be considered in step d3) is from 10 to 100.
16. The computer-implemented method according to claim 1 or 2, wherein the number of the ranked database vectors to be considered in step d3) is from 10 to 50.
17. The computer-implemented method according to claim 1 or 2, wherein the spectrum group of the known feed and / or feed ingredient in step c) of the method includes the spectra of all feeds and / or feed ingredients in ground and / or unground form used in animal nutrition.
18. The computer-implemented method according to claim 1 or 2, wherein the spectrum group includes at least 50 spectra of samples of each feed and / or feed ingredient in ground and / or unground form used in animal nutrition.
19. The computer-implemented method according to claim 1 or 2, wherein the spectrum group of the known feed and / or feed ingredient in step c) includes the spectra of all feed ingredients and / or feeds used in the nutrition of poultry, pigs, and / or animals raised in aquaculture.
20. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is an oilseed, a brewing and distillation by-product, a by-product of grain processing and feed production, a baking by-product, and / or an animal by-product.
21. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is soybeans.
22. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is soybean, soy product, soybean extracted meal and press cake, full-fat soybean, rapeseed meal and press cake, cotton extracted meal, peanut extracted meal, sunflower extracted meal, coconut extracted meal, and / or palm kernel extracted meal.
23. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is legume.
24. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is roasted guar gum powder.
25. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is dried distillers grains with solubles (DDGS).
26. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is corn gluten and / or corn seed flour.
27. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is meat meal.
28. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is fish meal, poultry meal, blood meal, and / or bone meal.
29. The computer-implemented method according to claim 1 or 2, wherein the feed and / or feed ingredient is grain.
30. The computer-implemented method according to claim 1 or 2, wherein the spectral group of the known feed and / or feed ingredient in step c) includes at least 50 spectra of samples of each feed and / or feed ingredient from each of its global growing regions.
31. The computer-implemented method according to claim 1 or 2, wherein the number of spectra of samples of any feed and / or feed ingredient from any of its global growing regions is from 50 to 10,000.
32. The computer-implemented method according to claim 1 or 2, wherein forming the first or second derivative of the spectrum of the unknown feed ingredient and / or feed in step a) and / or the first or second derivative of the spectrum of the known feed ingredient and / or feed in step c).
33. The computer-implemented method according to claim 1 or 2, wherein forming the first derivative of the spectrum of the unknown type of feed and / or feed ingredient in step a) and / or the first derivative of the spectrum of the known feed and / or feed ingredient in step c).
34. A system for predicting feed ingredients and / or feeds, comprising: i) A processing unit adapted to execute the computer-implemented method according to any one of claims 1 to 33.
35. The system according to claim 34, wherein the processing unit forms a network with at least one other processing unit on which the database vector is stored.
Citation Information
Patent Citations
Optimized spectral matching and display
WO2016141198A1
Near infrared spectrum similarity calculation method and device and qualitative analysis system of substances
CN108362662A