Methods, apparatus and computer-readable media for detecting anomalies in substances.
By constructing chemical fingerprints of substances and utilizing PCA graphs and machine learning methods, the efficiency and accuracy issues of food adulteration detection in existing technologies have been resolved, enabling rapid and low-cost food adulteration detection.
Patent Information
- Application Number
- CN202210685683.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-14
- Filing Date
- 2022-06-16
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-06-16
AI Technical Summary
Existing methods for detecting food adulteration are mainly target-oriented, making it difficult to detect adulterants that have not been encountered before. They are also costly, require specialized facilities and are labor-intensive, failing to meet the need for rapid and effective detection.
A non-target-oriented approach is adopted to obtain the chemical fingerprint of substances, construct multidimensional principal component analysis (PCA) diagrams and build contour pattern prediction models, and use machine learning methods to optimize the prediction models to identify anomalous samples.
It enables rapid and accurate identification of unknown adulterants in food without the need for expensive facilities and a large number of personnel, improving detection efficiency and accuracy, and can identify the presence of a variety of abnormal components.
Smart Images

Figure CN116794245B_ABST
Abstract
Description
Technical Field
[0001] This specification relates broadly, but not exclusively, to methods, apparatus, and computer-readable media for the detection of anomalies in substances. Background Technology
[0002] Food adulteration poses a significant risk to public health. Some common food products susceptible to adulteration include olive oil, milk, honey, saffron, orange juice, coffee, apple juice, wine, vanilla extract, and maple syrup. Examples include the addition of nitrogen-rich compounds to milk proteins, the discovery of undeclared horse meat in products advertised as containing beef, and the sale of counterfeit olive oil. Therefore, an effective protocol is needed to detect previously unseen adulterants.
[0003] Despite global collaboration, current testing methods are largely objective-oriented, as mandated by local legal authorities. For example, because all analyses seek specific chemicals and concentrations, newly designed and previously unknown adulterants can circumvent existing objective-oriented (or targeted, as interchangeable in this application) testing methods, posing a serious threat. However, it is impossible to test products using all available testing methods, not to mention that quality assessment methods currently used in the food industry are often expensive, require specialized infrastructure, and are labor-intensive.
[0004] Therefore, an effective non-targeted (or non-specific, as interchangeable in this application) protocol is desired to detect previously unencountered dopants (i.e., anomalies in the material). Summary of the Invention
[0005] According to one aspect, a method for anomaly detection of a substance is provided. The method includes: obtaining a first set of chemical fingerprints, wherein each chemical fingerprint in the first set indicates multiple physicochemical properties of each sample in a set of normal samples of the substance; converting the first set of chemical fingerprints into a cluster of data points in a multidimensional principal component analysis (PCA) plot, wherein each dimension of the multidimensional PCA plot is based on principal components (PCs), each PC corresponding to one of multiple physicochemical properties; constructing a contour pattern of the data point clusters as a predictive model, the predictive model being configured to identify new samples having chemical fingerprints falling outside the contour pattern as anomalies; and optimizing the predictive model using a second set of chemical fingerprints, wherein the second set of chemical fingerprints indicates multiple physicochemical properties of a set of test samples comprising multiple normal test samples and multiple anomalous test samples of the substance.
[0006] According to another aspect, an apparatus for anomaly detection of a substance is provided. The apparatus includes: at least one processor; and a memory including computer program code for execution by the at least one processor, the computer program code instructing the at least one processor to: obtain a first set of chemical fingerprints, wherein each chemical fingerprint in the first set indicates multiple physicochemical properties of each sample in a set of normal samples of the substance; convert the first set of chemical fingerprints into a cluster of data points in a multidimensional principal component analysis (PCA) plot, wherein each dimension of the multidimensional PCA plot is based on principal components (PCs), each PC corresponding to one of multiple physicochemical properties; construct a contour pattern of the data point clusters into a predictive model, the predictive model being configured to identify new samples having chemical fingerprints falling outside the contour pattern as anomalies; and optimize the predictive model using a second set of chemical fingerprints, wherein the second set of chemical fingerprints indicates multiple physicochemical properties of a set of test samples including multiple normal test samples and multiple anomalous test samples of the substance. Attached Figure Description
[0007] The embodiments and implementations are provided by way of example only, and those skilled in the art will better understand and readily comprehend these embodiments and implementations by reading the following written description in conjunction with the accompanying drawings.
[0008] Figure 1 This is a schematic diagram of an apparatus 100 for detecting anomalies in a substance according to one embodiment.
[0009] Figure 2 This is a flowchart illustrating a method 200 for anomaly detection of a substance according to one embodiment.
[0010] Figure 3A This is a principal component analysis (PCA) diagram of the data point clusters, shown in Figure 300. The data point clusters are derived from the first set of chemical fingerprints of a group of normal samples of a substance.
[0011] Figure 3B It is a histogram that shows the corresponding value of the square root of the sum of all calculated principal components (PCs) based on the data points, with the data point clusters arranged at a predetermined number of intervals.
[0012] Figure 3C The outline pattern 350 of the data point cluster is shown. The outline pattern 350 shows the "core" and "skin" of the data point cluster.
[0013] Figure 4A and Figure 4B This is an alternative perspective showing the new data points in PCA diagram 300. The new data points are derived from a second set of chemical fingerprints of a set of test samples, including multiple normal test samples and multiple abnormal test samples of the substance.
[0014] In particular, Figure 4A The diagram shows a contour pattern (depicted as points forming an envelope shape 402) surrounding a normal test sample (depicted as points 404 enclosed by the envelope shape 402) from the test sample, while Figure 4B The sample found outside the outline pattern (depicted as a dot such as point 406) is shown to be an anomaly test sample.
[0015] Figure 4C A steep slope plot of the cumulative percentage variance of profile patterns with different numbers (from 1 to 8 in this example) of principal components (PC) is shown.
[0016] Figure 5 Chart 500 is shown, which depicts the distribution of the squared Mahalanobis distance (MD) score for each sample in a set of test samples. Line 502 represents the threshold squared MD score of 5.4 to distinguish whether a sample falls inside (normal) or outside (abnormal) the contour pattern.
[0017] Figure 6A This is a graph showing the overall accuracy when iterating through a series of squared MD scores to determine the threshold squared MD score.
[0018] Figure 6B This is a graph showing the sensitivity values when iterating through a series of squared MD scores to determine the threshold squared MD score.
[0019] Figure 7A This is a graph showing the spectral data of raw milk retrieved from Fourier transform infrared (FTIR) spectroscopy. In this graph, the spectral data are plotted as absorbance versus wavenumber. Figure 7A In the text, the numbers 1 to 8 with arrows refer to eight (8) spectral regions in which absorbance values are extracted for hierarchical cluster learning to define real normal samples.
[0020] Figure 7B It is a tree diagram that shows the sample clusters within the main branch in hierarchical cluster learning.
[0021] Figure 7C It is an unrooted tree, which shows two main clusters observed by hierarchical cluster learning. In the two main clusters, one cluster includes branches #1-3, and the other cluster includes branches #4-7.
[0022] Figure 8 This is a heatmap of hierarchical clustering learning, visually classifying 113 anomalous samples (i.e., abnormal) and 198 truly normal samples within the national standard. Figure 8 In the figure, the numbers on the left indicate the branch numbers of the sample's dendrogram, while the numbers at the bottom indicate the regions from which absorbance values are extracted.
[0023] Figure 9 The distribution of different anomalous components in the seven branches is shown (TN: true normal, AD: dopant, N: nitrogen, Carb: carbohydrate).
[0024] Figure 10A It is a graph showing the ratio of real normal samples within the corresponding branch. Figure 10B This is a graph showing the ratio of nitrogen-rich anomalous samples within the corresponding branches. Figure 10C This is a graph showing the ratio of carbohydrate-based anomalous samples within the corresponding branch. Figure 10D This is a graph showing the ratio of buffer reagents used as anomalous samples within the corresponding branches. Figures 10A to 10D In the middle, the dashed line and the accompanying percentage are... Figure 10A The percentage of truly normal individuals is represented by (198 / 311 = 63.7%). Figure 10B The figure represents the total percentage of nitrogen-rich anomalous samples (26 / 311 = 8.4%). Figure 10C The figure represents the total percentage of abnormal samples based on carbohydrates (65 / 311 = 20.9%) and in... Figure 10D The value in the figure represents the total percentage of the buffer reagent (22 / 311 = 7.1%).
[0025] Figure 11A The boosting tree of an XGBoost model learned from a current prediction model is shown according to one embodiment.
[0026] Figure 11B This is a graph showing the relative importance of each component feature (i.e., physicochemical properties) f1-f8 according to the learned XGBoost model. f1-f8 represent fat, protein, total solids, non-fat solids, lactose, relative density, freezing point, and acidity, respectively.
[0027] Figure 12 A block diagram of a computer system 1200 is shown, which is suitable for use as a device for detecting anomalies in substances.
[0028] Figure 13 This is a diagram illustrating an embodiment of a data flow framework 1300 for converting chemical and biological fingerprints into a standardized format. The standardized format includes patterns regarding numerical and classification test results (num_cat_testing), spectral test results (spec_testing), numerical and classification test templates (num_cat_testing_template), and spectral test result templates (spec_testing_template). Those skilled in the art will understand that the standardized format may include other patterns regarding other physicochemical properties of the sample.
[0029] Figure 14 It is shown Figure 13 A diagram illustrating an embodiment of the patterns depicted. In this embodiment, an exemplary table is shown having data columns captured in the respective patterns num_cat_testing, spec_testing, num_cat_testing_template, and spec_testing_template.
[0030] Those skilled in the art will understand that the elements in the figures are shown for simplicity and clarity and are not necessarily depicted to scale. For example, the dimensions of some elements in the illustrations, block diagrams, or flowcharts may be exaggerated relative to other elements to aid in understanding the present embodiment. Detailed Implementation
[0031] Embodiments will be described by way of example only with reference to the accompanying drawings. The same reference numerals and characters in the drawings refer to the same elements or equivalents.
[0032] The following sections of the specification are explicitly or implicitly presented based on the algorithms and functional or symbolic representations of operations on data in computer memory. These algorithmic descriptions and functional or symbolic representations are the means by which those skilled in the art of data processing most effectively communicate their work to others skilled in the art. In this document, an algorithm is generally considered to be a self-consistent sequence of steps that leads to a desired result. These steps are those that require physical manipulation of physical quantities, such as electrical, magnetic, or optical signals that can be stored, transmitted, combined, compared, and otherwise manipulated.
[0033] Unless otherwise expressly stated and apparent from the following, it should be understood that throughout this specification, discussions using terms such as “obtain,” “convert,” “build,” “optimize,” “calculate,” “arrange,” “determine,” “grade,” “iterate,” and “train” refer to the actions and processes of a computer system or similar electronic device that manipulate and convert data represented as physical quantities within the computer system into other data similarly represented as physical quantities within the computer system or other information storage, transmission, or display device.
[0034] This specification also discloses apparatus for performing the methods described herein. Such apparatus may be specifically constructed for the desired purpose or may include a computer or other means selectively activated or reconfigured by a computer program stored in the computer. The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various machines may be used with the program in accordance with the teachings herein. Alternatively, it may be appropriate to construct more specialized equipment to perform the required method steps. The architecture of a computer suitable for performing the various methods / processes described herein will be apparent from the following description.
[0035] Furthermore, this specification implicitly discloses a computer program, as it will be apparent to those skilled in the art that the various steps of the methods described herein can be implemented by computer code. The computer program is not intended to be limited to any particular programming language or its implementation. It should be understood that the teachings of the specification contained herein can be implemented using a variety of programming languages and their encoding. Moreover, the computer program is not intended to be limited to any particular control flow. The computer program has many other variations that may use different control flows without departing from the spirit or scope of the invention.
[0036] Furthermore, one or more steps of a computer program can be executed in parallel rather than sequentially. Such a computer program can be stored on any computer-readable medium. Computer-readable media can include storage devices such as disks or optical discs, memory chips, or other storage devices suitable for interfacing with a computer. Computer-readable media can also include hardwired media, as exemplified in Internet systems, or wireless media, as exemplified in GSM mobile phone systems. When the computer program is loaded and executed on such a computer, it effectively creates a device for implementing the steps of the preferred method.
[0037] This specification uses the term "configured as" in conjunction with systems, apparatuses, and computer program components. For a system of one or more computers to be configured to perform a specific operation or action, this means that software, firmware, hardware, or a combination thereof are installed on the system, which, in operation, causes the system to perform those operations or actions. For one or more computer programs to be configured to perform a specific operation or action, this means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform those operations or actions. For a dedicated logic circuit to be configured to perform a specific operation or action, this means that the circuit has electronic logic for performing those operations or actions.
[0038] Embodiments of this application provide methods for non-targeted detection of anomalies in substances. These methods utilize machine learning approaches to learn the chemical fingerprint of a substance, thereby constructing and optimizing predictive models to identify anomalies, whether previously encountered or not. In this application, anomalies are interchangeably referred to as anomalous samples or dopants (ADs).
[0039] The predictive model provided in this application offers several advantages. First, it eliminates the need for labor-intensive sample pretreatment or testing using expensive specialized infrastructure and machinery, thus reducing the chance of human error and associated time costs. Second, after the training phase of the predictive model, prior knowledge of the presence of anomalies is unnecessary, as the chemical fingerprints of longitudinal inspection data help detect changes in milk composition and identify anomalous patterns over time. Third, it eliminates the need to manually establish tolerance levels for known real reference products (i.e., genuine normal samples), as this application provides efficient mathematical formulas for establishing threshold scores (interchangeably referred to as cutoff scores) that define anomalies, thereby saving time and improving the accuracy of anomaly detection. Finally, and most importantly, the predictive model is capable of identifying more than one anomaly in each sample without prior knowledge of its composition. In reality, multiple unknown anomalies can coexist in a single sample, and their interactions can lead to signal interference and entirely different chemical fingerprints. Newly encountered chemical fingerprints identified by the predictive model can be learned for further characterization, which can address the practical problem of food adulteration in real-world scenarios.
[0040] Figure 1 A schematic diagram of an apparatus 100 for detecting anomalies in a substance is shown.
[0041] The device 100 includes at least one or more processors 102 and a memory 104. The at least one processor 102 and the memory 104 are interconnected. The memory 104 includes computer program code. Figure 1 (not shown), which is used by the at least one processor 102 to perform, as Figure 2 The steps for anomaly detection shown and described in this application.
[0042] In step 202, computer program code instructs the at least one processor 102 to obtain a first set of chemical fingerprints. Each chemical fingerprint in the first set indicates multiple physicochemical properties of each sample in a set of normal samples of a substance.
[0043] In one embodiment, the substance is milk. In one example, a set of normal samples of the substance consists of 63,473 normal raw milk samples collected by Mengniu Dairy (Group) Co., Ltd. (Helin, Inner Mongolia, China) from 2017 to 2019 using a MilkoScan FT120 (FOSS Analytical, Denmark). The chemical fingerprints of the normal raw milk samples were acquired using a Fourier transform infrared (FTIR) spectrometer and stored in a database. Such a database can be established in the internal memory or storage component of device 100 or in external storage space.
[0044] Each chemical fingerprint indicates the variability of each sample. For example, a chemical fingerprint can include two data formats: (1) each sample in the range of 1000-3550 cm⁻¹. -1 (1) spectral data of absorbance relative to wavenumber within the range; and (2) component data for each sample having multiple component characteristics (interchangeably referred to as physicochemical properties). In one embodiment, the multiple physicochemical properties include eight (8) physicochemical properties, such as fat, protein, non-fat solids, total solids, lactose, relative density, freezing point depression, and acidity of milk.
[0045] Those skilled in the art will understand that compositional data can include physicochemical properties for different groups of different types of milk. Therefore, a chemical fingerprint will reflect these physicochemical properties.
[0046] Furthermore, since the current method can be used to detect anomalies in substances other than milk, those skilled in the art will understand that the composition data can include physicochemical properties for various groups of substances (food or non-food). Therefore, a chemical fingerprint will reflect these physicochemical properties.
[0047] In this embodiment, the first set of chemical fingerprints includes 63,473 chemical fingerprints from 63,473 normal raw milk samples. Each chemical fingerprint indicates eight physicochemical properties and / or spectral data for each sample in the group of 63,473 normal raw milk samples. In step 202, the first set of 63,473 chemical fingerprints is obtained from the database.
[0048] In step 204, computer program code instructs the at least one processor 102 to convert the first set of chemical fingerprints into a cluster of data points in a multidimensional principal component analysis (PCA) plot. Each dimension of the multidimensional PCA plot is based on a principal component (PC). In this application, each PC corresponds to one of a plurality of physicochemical properties. In this sense, PCs can be interchangeably referred to as physicochemical properties in the following description. In this application, for ease of reference, the data cluster can be referred to as the first envelope.
[0049] An example of the transformed PCA diagram is in Figure 3A As shown in the image. Figure 3A As shown, PCA Figure 300 includes a cluster of data points for all 63,473 chemical fingerprints from a set of 63,473 normal raw milk samples. However, this cluster is dense, and including all data points in the cluster to build a predictive model would require a significant amount of computation. This technical challenge is advantageously addressed through the following steps of this application.
[0050] In step 206, computer program code instructs the at least one processor 102 to construct a predictive model from the contour pattern of the data point clusters. This predictive model can identify new samples with chemical fingerprints falling outside the contour pattern as anomalies.
[0051] For example, when the substance is milk, anomalies include one or more of the following: potassium sulfate, potassium dichromate, citric acid, ammonium sulfate, melamine, urea, lactose, glucose, sucrose, maltodextrin and fructose, water, sodium citrate, and real-life scenarios (such as milk having a cow odor and milk being improperly stored for 36 hours). Those skilled in the art will understand that when anomalies change, the changes are based on different substances. In one embodiment, if the predictive model receives a chemical fingerprint of a new milk sample that, when converted into new data points in a PCA graph, falls outside the outline pattern in the PCA graph, the new milk sample will be identified as an anomaly by the predictive model. Similarly, if the predictive model receives a chemical fingerprint of a new milk sample that, when converted into new data points in a PCA graph, falls on or within the outline pattern in the PCA graph, the new milk sample will be identified as normal by the predictive model.
[0052] Figure 3B An embodiment of step 206 in constructing the contour pattern is illustrated. As shown, during the construction of the contour pattern, computer program code instructs at least one processor 102 to perform the following sub-steps:
[0053] First, in substep 206a, for each data point in the cluster, calculate the square root of the sum of all squared PCs.
[0054] Subsequently, in substep 206b, the data points are clustered into a predetermined number of intervals based on the corresponding value of the square root of the calculated sum of the data points.
[0055] Subsequently, in substep 206c, the contour pattern in the PCA diagram is obtained by removing data points from the cluster that fall into one or more permutations with more than a predetermined number of data points.
[0056] exist Figure 3BIn the illustrated embodiment, the predetermined number of intervals is 20, and the predetermined number of data points is 1000. Those skilled in the art will understand that the predetermined number can vary and can be determined according to actual needs. In this embodiment, a total of 61,782 data points are removed from the PCA graph according to sub-steps 206a, 206b, and 206c, and the remaining 1691 data points in the PCA graph are used to construct a contour pattern 350. This contour pattern 350 shows the "core" and "skin" of the data point cluster, maintaining the structure of the data point cluster. In this application, for ease of reference, the constructed contour pattern may be referred to as the second envelope.
[0057] In the above-described manner, the constructed contour pattern advantageously enables this application to significantly reduce the required computation from 63,473 data points to only 1,691 data points, while preserving the structure of all 63,473 chemical fingerprint data point clusters. In other words, it improves computational efficiency while maintaining the accuracy of the prediction model.
[0058] In step 208, the computer program code instructs the at least one processor 102 to optimize the prediction model using a second set of chemical fingerprints of a set of test samples, including multiple normal test samples and multiple abnormal test samples.
[0059] In one embodiment, the test sample set comprises 1087 test samples, including 976 normal raw milk samples and 111 abnormal raw milk samples. The abnormal raw milk samples were spiked with various concentrations of the following chemicals to simulate real-world adulterants (ADs): potassium sulfate, potassium dichromate, citric acid, ammonium sulfate, melamine, urea, lactose, glucose, sucrose, maltodextrin and fructose, water, sodium citrate, and real-world scenarios (such as milk having a cow odor and milk being improperly stored for 36 hours).
[0060] Table 1 shows the number of abnormal samples in this group of test samples that were spiked with the corresponding concentration of dopant.
[0061] Added dopant (AD) Number of abnormal samples Common chemicals Potassium dichromate* 9 Potassium sulfate* 9 Sodium citrate 4 Citric acid* 9 Nitrogen-based doping Ammonium sulfate* 5 Urea* 5 Melamine* 10 Carbohydrate-based dopants sucrose# 5 glucose# 5 lactose# 5 fructose# 5 Maltodextrin# 5 Improper storage for 36 hours 23 water 2 cow smell 10 total 111
[0062] Table 1 shows the number of abnormal samples with added spiked dopants.
[0063] *The concentration of adulterants added is 0.01, 0.02, 0.05, 0.1, and 0.2 grams per 100 grams of raw milk; #The concentration of adulterants added is 0.1, 0.2, 0.5, 1, and 2 grams per 100 grams of raw milk.
[0064] Similar to the first group of chemical fingerprints, each chemical fingerprint in the second group indicates the variability of each normal or abnormal test sample in that group of test samples. For example, chemical fingerprints may include two data formats: (1) each sample in the range of 1000-3550 cm⁻¹ -1 (1) spectral data in terms of absorbance relative to wavenumber; and (2) component data with multiple physicochemical properties for each normal or abnormal test sample. In one embodiment, the multiple physicochemical properties include eight (8) physicochemical properties for each normal raw milk sample or each abnormal raw milk sample, such as fat, protein, non-fat solids, total solids, lactose, relative density, freezing point depression, and acidity.
[0065] The embodiment of step 208 in optimizing the prediction model is reflected in Figures 4A to 4C As shown in the figure, during the process of optimizing the prediction model, the computer program code instructs the at least one processor 102 to perform the following sub-steps:
[0066] First, in sub-step 208a, for this set of test samples, the second set of chemical fingerprints is converted into new data points in the PCA graph. One embodiment of the new data points converted from the second set of chemical fingerprints is depicted as points enclosed by envelope shape 402 and points found outside envelope shape 402 in PCA graphs 400, 450, such as... Figure 4A and Figure 4B As shown in the alternative perspective. In PCA figures 400 and 450, the contour pattern obtained in step 206 is depicted as points forming the envelope shape 402.
[0067] Subsequently, in substep 208b, for each sample in the group of test samples, the squared Mahalanobis distance (MD) score between the sample and the centroid of the profile pattern in the PCA plot is calculated, and a threshold squared MD score is determined to distinguish whether the sample falls inside or outside the profile pattern.
[0068] In one embodiment, the squared MD score can be calculated using an MD score calculation script on R or RStudio as follows, where the first X principal components (PCs) are considered. It should be understood that the italicized names can be replaced with your own names as needed:
[0069] name1<-prcomp(matrixdataframe,scale=TRUE)
[0070] name2<-as.data.frame(name1$x[,1:X])
[0071] MD<-mahalanobis(name2,center=colMeans(name2),cov=cov(name2))
[0072] name2$MD<-round(MD,3)
[0073] like Figure 4A As shown in the embodiments, the sample (described by points enclosed by the envelope shape 402) wrapped within the contour pattern (depicted as points formed by the envelope shape 402) is a normal test sample from a set of test samples, while Figure 4B The sample found outside the outline pattern (e.g., depicted as a dot such as point 406) is shown to be an anomalous test sample from the group of test samples.
[0074] Figure 4C A steep slope plot showing the cumulative percentage variance of profile patterns with different numbers (from 1 to 8 in this example) of principal components (PCs) is shown. Figure 4C The diagram shows that the four PCs are responsible for 94.9% of the variance of the profile pattern. Therefore, this application can advantageously improve computational efficiency by using only the four most important PCs to calculate the squared MD fraction in substep 208b. Regarding... Figure 11A and 11B Details of the four most important PCs are described.
[0075] Furthermore, as shown in Table 4 below, compared to the data point clusters used for the first set of chemical fingerprints (i.e., the first envelope), the contour pattern (i.e., the second envelope) produces better accuracy (95.22% vs. 89.97%), specificity (99.45% vs. 91.68%), precision (96.21% vs. 64.98%), and CSI (70.95% vs. 56.40%), with comparable BA (86.22% vs. 86.36%), but with lower sensitivity (72.99% vs. 81.03%). This demonstrates that the prediction model calculated using the squared MD scores of the four most important PCs in sub-step 208b can achieve accurate prediction results by effectively distinguishing between normal and abnormal samples in an industrial dairy environment, outperforming national standards.
[0076] In one embodiment, in sub-step 208b, computer program code instructs the at least one processor 102 to perform the following sub-steps to determine the threshold squared MD score.
[0077] First, in substep 208bl, the squared MD scores of the test samples are sorted.
[0078] Subsequently, in substep 208b2, a series of squared MD scores are iterated as thresholds in the sensitivity calculation.
[0079] Subsequently, in substep 208b3, the squared MD score that has the highest sensitivity value and an overall accuracy of over 95% among the iterated series of squared MD scores is determined as the threshold squared MD score.
[0080] In some embodiments, the sensitivity value is calculated based on the following equation:
[0081]
[0082] In some embodiments, the overall accuracy is calculated based on the following equation:
[0083]
[0084] Figure 5 Chart 500 is shown, depicting the distribution of the squared Mahalanobis distance (MD) score for each sample in a set of test samples according to one embodiment. In this embodiment, line 502 represents a threshold squared MD score of 5.4 to distinguish whether a sample falls inside (normal) or outside (abnormal) the contour pattern.
[0085] Table 2 shows the overall accuracy, sensitivity and specificity, precision (or positive predictive value), negative predictive value (NPV), critical success index (CSI), and balanced accuracy (BA) when a threshold squared MD score of 5.4 is used in sub-step 208b to distinguish whether each test sample in the group of test samples falls inside (normal) or outside (abnormal) of the contour pattern. In this embodiment, the threshold squared MD score of 5.4 provides overall accuracy, sensitivity, and specificity of 95.86%, 72.99%, and 97.39%, respectively. Furthermore, the precision (or positive predictive value) is 65.13%, the negative predictive value (NPV) is 98.18%, the CSI is 52.48%, and the balanced accuracy (BA) is 85.19%.
[0086]
[0087] Table 2
[0088] Figure 6A A graph depicting the overall accuracy when iterating a series of MD squared scores to determine the threshold squared MD score is shown. Figure 6BA graph depicting the sensitivity values as a series of squared MD scores are iterated to determine the threshold squared MD score is shown. Sample grouping indicates that a threshold squared MD score of 5.4 is highly effective in identifying normal raw milk samples, with a correct prediction rate of 99.5%. This threshold squared MD score is also effective in detecting anomalous samples, providing a correct prediction rate of 94.6%, as... Figure 6A and 6B As shown.
[0089] Table 3 shows the detection limit for each anomalous (dopant) in a set of test samples when the threshold MD score used in sub-step 208b is 5.4. The squared MD scores between normal and anomalous samples were compared using a Welch two-sample t-test (2-sided), yielding 6.523 × 10⁻⁶. -8 The p-values were calculated, and all squared MD scores of normal samples in the test group lay within the profile pattern. Furthermore, the squared MD values could even detect anomalous samples that had passed national standards: a t-test comparing “spiked anomalous samples within national standards” (N=44) with normal samples yielded a p-value of 0.01359. Among the spiked anomalous samples, raw milk diluted with water (N=2) and inferior raw milk (with a “cow odor”, N=10) were easily detected, with a correct prediction rate of 100%. Most raw milk spoiled due to improper storage conditions (N=23) and raw milk spiked with various buffer reagents or inorganic chemicals (potassium sulfate, potassium dichromate, citric acid, sodium citrate) were also detected, resulting in correct prediction rates of 87.0% and 71.0%, respectively. Glucose, maltodextrin, and fructose were detectable at 2%, while sucrose was detectable only at 1%. While melamine can be detected at concentrations as low as 0.02% (w / w), ammonium sulfate and urea are detectable at 0.1% and 0.2%, respectively. Nitrogen-rich (ammonium sulfate, melamine, urea) or carbohydrate-based (lactose, glucose, sucrose, maltodextrin, fructose) adulterants are difficult to detect. Lactose is undetectable even at a concentration of 2%.
[0090]
[0091]
[0092] Table 3
[0093] To compare the performance of the second “enveloping” with the first “enveloping”, a cutoff squared MD score of 8.6 was obtained by applying sub-step 208b to the centroids in the data point clusters (first envelope) instead of the centroids in the pattern profile (second envelope). While the first envelope contains more samples, resulting in higher overall accuracy and sensitivity, precision and CSI are significantly reduced, and sucrose becomes undetectable. Therefore, the prediction model using the pattern profile (second envelope) according to this application outperforms the prediction model using the data point clusters (first envelope) of the first set of chemical fingerprints in terms of accuracy, specificity, precision, and CSI, as shown in Table 4 below.
[0094]
[0095]
[0096] Table 4 Comparison of prediction performance between the first and second "envelopes"
[0097] In some embodiments, to facilitate the calculation of overall accuracy, computer program code may further instruct at least one processor 102 to train a prediction model to define true normal samples based on hierarchical cluster learning of spectral data of a set of test samples.
[0098] Figure 7A This is a graph showing the spectral data of raw milk retrieved from Fourier transform infrared (FTIR) spectroscopy. As mentioned above, the spectral data can be included in the chemical fingerprint of each sample. In this graph, the spectral data is plotted as absorbance versus wavenumber. Figure 7A In the text, the numbers 1 to 8 with arrows refer to eight (8) spectral regions in which absorbance values are extracted for hierarchical cluster learning to define real normal samples.
[0099] In one embodiment, to test the effectiveness of classifying strictly anomalous samples, unsupervised hierarchical clustering was performed on 40 anomalous samples and 68 normal samples with low concentrations of AD that met national standards. For example... Figure 7A As shown, the spectral regions 1000-1100, 1500-1600, 1730-1800, 2840-2940, and 3450-3550 cm⁻¹ were extracted from the spectral data of each sample. -1 The absorbance values of the seven inner peaks and 1250-1450 cm⁻¹ -1 The average absorbance value was calculated. Unsupervised hierarchical clustering was performed using Euclidean distance and Ward's linkage method. The clustering was visualized using dendrograms and heatmaps, such as... Figure 7B , Figure 7C and Figure 8 As shown.
[0100] Figure 7B It is a tree diagram that shows the sample clusters within the main branch in hierarchical cluster learning. Figure 7B The results show that carbohydrate-based dopants cluster in branches #1 and #6, nitrogen-rich dopants cluster in branch #7, buffer reagents or inorganic chemicals cluster in branch #4, while normal samples cluster in branches #2, #3, and #5.
[0101] Figure 7C It is an unrooted tree, which shows two main clusters observed through hierarchical cluster learning. In the two main clusters, one cluster includes branches #1 to #3, and the other cluster includes branches #4 to #7.
[0102] Figure 8 A heatmap showing hierarchical clustering learning is presented, visually classifying 113 anomalous samples (i.e., abnormal) and 198 truly normal samples within the national standard. Figure 8 In the figure, the numbers on the left side represent the branch numbers # of the sample's dendrogram, while the numbers at the bottom of the figure represent the regions from which absorbance values are extracted. Figure 8 Subclasses were observed in normal raw milk samples and carbohydrate-based anomalous samples within clusters #2, #3, and #5, and branches #1 and #6. On the other hand, nitrogen-rich dopants and buffer reagents had a stronger impact on FTIR spectra, with only one major branch observed.
[0103] Figure 9 The distribution of different anomalous components in the seven branches is shown (TN: true normal, AD: dopant, N: nitrogen, Carb: carbohydrate). Figure 9 Further, it was shown that the first cluster contained mostly normal samples (true normal % = 81.2%; total = 101, true normal = 82, carbohydrate-based AD = 19), while the second cluster contained more different types of abnormal samples (true normal % = 55.2%; total = 210, true normal = 116, nitrogen-rich AD = 26, carbohydrate-based AD = 46, buffer reagent = 22).
[0104] Figure 10A It is a graph showing the ratio of real normal samples within the corresponding branch. Figure 10B This is a graph showing the ratio of nitrogen-rich anomalous samples within the corresponding branches. Figure 10C This is a graph showing the ratio of carbohydrate-based anomalous samples within the corresponding branch. Figure 10D This is a graph showing the ratio of buffer reagents used as anomalous samples within the corresponding branches. Figures 10A to 10D In the middle, the dashed line and the accompanying percentage are... Figure 10A The percentage of truly normal individuals is represented by (198 / 311 = 63.7%). Figure 10BThe figure represents the total percentage of nitrogen-rich anomalous samples (26 / 311 = 8.4%). Figure 10C The figure represents the total percentage of abnormal samples based on carbohydrates (65 / 311 = 20.9%) and in... Figure 10D The value in the figure represents the total percentage of the buffer reagent (22 / 311 = 7.1%).
[0105] In some embodiments, to optimize overall accuracy, the computer program code may further instruct the at least one processor 102 to train a prediction model based on Extratree or XGBoost learning of multiple physicochemical properties of a set of test samples. The prediction model learns that certain physicochemical properties are more indicative of the profile pattern than others.
[0106] Figure 11A The boosting tree of an XGBoost model learned from a current prediction model is shown according to one embodiment.
[0107] Figure 11B This is a graph showing the relative importance of each component feature (i.e., physicochemical property) f1-f8 according to the learned XGBoost model, where f1-f8 represent fat, protein, total solids, non-fat solids, lactose, relative density, freezing point, and acidity, respectively.
[0108] Figure 11A and 11B Both studies showed that non-fat solids are the most indicative factor in determining whether a raw milk sample has been processed; while protein and total lactose solids are the most indicative factors in detecting potential adulteration or contamination in samples.
[0109] Figure 12 A block diagram of a computer system 1200 is shown, which is suitable for use as an apparatus 100 for anomaly detection of substances as described herein.
[0110] The following description of the computer system / computing device 1200 is provided by way of example only and is not intended to be limiting.
[0111] like Figure 12 As shown, the example computing device 1200 includes a processor 1204 for executing software routines. Although a single processor is shown for clarity, the computing device 1200 may also include a multiprocessor system. The processor 1204 is connected to a communication infrastructure 1206 for communicating with other components of the computing device 1200. The communication infrastructure 1206 may include, for example, a communication bus, a crossbar switch, or a network.
[0112] The computing device 1200 also includes a main memory 1208 such as random access memory (RAM) and an auxiliary memory 1210. The auxiliary memory 1210 may include, for example, a hard disk drive 1212 and / or a removable storage drive 1214, which may include a magnetic tape drive, an optical disc drive, etc. The removable storage drive 1214 reads from and / or writes to the removable storage unit 1218 in a well-known manner. The removable storage unit 1218 may include magnetic tapes, optical discs, etc., read from and written to by the removable storage drive 1214. As those skilled in the art will understand, the removable storage unit 1218 includes a computer-readable storage medium in which computer-executable program code instructions and / or data are stored.
[0113] In alternative implementations, auxiliary memory 1210 may additionally or alternatively include other similar means for allowing computer programs or other instructions to be loaded into computing device 1200. Such means may include, for example, removable storage unit 1222 and interface 1220. Examples of removable storage unit 1222 and interface 1220 include removable storage chips (such as EPROM or PROM) and associated slots, as well as other removable storage units 1222 and interface 1220 that allow software and data to be transferred from removable storage unit 1222 to computer system 1200.
[0114] The computing device 1200 also includes at least one communication interface 1224. The communication interface 1224 allows software and data to be transferred between the computing device 1200 and external devices via a communication path 1226. In various embodiments, the communication interface 1224 allows data to be transferred between the computing device 1200 and a data communication network (such as a public or private data communication network). The communication interface 1224 can be used to exchange data between different computing devices 1200 that form part of an interconnected computer network. Examples of the communication interface 1224 may include a modem, a network interface (such as an Ethernet card), a communication port, an antenna with associated circuitry, etc. The communication interface 1224 may be wired or wireless. Software and data transmitted through the communication interface 1224 are in the form of signals, which may be electronic signals, electromagnetic signals, optical signals, or other signals that can be received by the communication interface 1224. These signals are provided to the communication interface via the communication path 1226.
[0115] Optionally, the computing device 1200 also includes a display interface 1202 and an audio interface 1232, the display interface 1202 performing operations for rendering an image to an associated display 1230, and the audio interface 1232 performing operations for playing audio content through one or more associated speakers 1234.
[0116] As used herein, the term "computer program product" may refer in part to removable storage unit 1218, removable storage unit 1222, hard disk installed in hard disk drive 1212, or a carrier wave carrying software connected to communication interface 1224 via communication path 1226 (wireless link or cable). A computer-readable storage medium is any non-transitory tangible storage medium that provides recorded instructions and / or data to computing device 1200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, and Blu-ray discs. TM Optical discs, hard disk drives, ROMs or integrated circuits, USB storage devices, magneto-optical discs, or computer-readable cards (such as PCMCIA cards), whether these devices are internal or external to the computing device 1200. Examples of temporary or intangible computer-readable transmission media that may also be involved in providing software, applications, instructions, and / or data to the computing device 1200 include radio or infrared transmission channels and network connections to another computer or networked device, as well as the Internet or intranet including email transmissions and information recorded on websites, etc.
[0117] The computer program (also referred to as computer program code) is stored in main memory 1208 and / or secondary memory 1210. The computer program may also be received via communication interface 1224. Such a computer program, when executed, enables computing device 1200 to perform one or more features of the embodiments discussed herein. In various embodiments, the computer program, when executed, enables processor 1204 to perform the features of the embodiments described above. Therefore, such a computer program represents a controller of computer system 1200.
[0118] The software may be stored in a computer program product and loaded into a computing device 1200 using a removable storage drive 1214, a hard disk drive 1212, or an interface 1220. Alternatively, the computer program product may be downloaded to the computer system 1200 via communication path 1226. When executed by the processor 1204, the software causes the computing device 1200 to perform the functions of the embodiments described herein.
[0119] It should be understood that Figure 12The embodiments described are presented by way of example only. Therefore, in some embodiments, one or more features of the computing device 1200 may be omitted. Furthermore, in some embodiments, one or more features of the computing device 1200 may be combined together. Additionally, in some embodiments, one or more features of the computing device 1200 may be divided into one or more components.
[0120] The techniques described in this specification produce one or more technical effects. As described above, embodiments of this application provide methods for non-targeted detection of anomalies in substances. These methods utilize machine learning methods to learn the chemical fingerprint of a substance, thereby constructing and optimizing predictive models to identify anomalies that have been encountered or not encountered before. In this application, anomalies may be interchangeably referred to as anomalous samples or dopants (ADs).
[0121] Furthermore, as those skilled in the art will understand, predictive models can be trained using continuous updates of the chemical fingerprints of new milk samples. Updates can include as many factors as possible, including key chemometrics and all known characteristics of the cattle, such as geographic seasonal logistics variations, breed, feed, age, and information regarding the erroneous addition of specific ingredients (such as non-dairy proteins, illegal preservatives, high levels of legal preservatives, antibiotics, pesticides). Chemical fingerprints should be extended beyond FTIR to also cover mass spectrometry, MMR, infrared spectroscopy, liquid chromatography, gas chromatography, etc. Chemical fingerprints can also be extended to include biological fingerprints, such as data from next-generation sequencing (NGS), to provide valuable information on fraud, such as different types of milk, plant and animal fats; fraudulent claims about breed and geographic breeding; and contamination by pathogenic microorganisms. Chemical and biological fingerprints can be reported in various heterogeneous formats, including numerical values, data points, graphs, images, numerical representations, etc. These chemical and biological fingerprints can be converted into standardized formats in schema form for easy data storage in databases. Figure 13 and 14 An example of the pattern is shown in the diagram. Figure 13 In this example, chemical fingerprints and biological fingerprints are referred to as "data" in the exemplary data flow framework 1300. Figure 14 In the example, a data table is shown as having, for example, Figure 13 The data columns captured in the corresponding patterns num_cat_testing, spec_testing, num_cat_testing_template, and spec_testing_template are shown.
[0122] These data can be processed in data processing module 1302 and converted into a standardized format for storage in database 1304. The standardized format includes patterns regarding numerical and classification test results (num_cat_testing), spectral test results (spec_testing), numerical and classification test templates (num_cat_testing_template), and spectral test result templates (spec_testing_template). Those skilled in the art will understand that the standardized format may include other patterns regarding other physicochemical properties of the sample. The stored data in the standardized format can be provided to machine learning module 1306 to train a predictive model for anomaly detection of substances based on the embodiments described in the preceding paragraphs.
[0123] With continuous machine learning, the predictive power of the predictive model will be further enhanced, enabling the method provided in this application to be used as a standard screening for a wide range of food commodities.
[0124] Those skilled in the art will understand that various changes and / or modifications can be made to the invention as illustrated in the specific embodiments without departing from the spirit or scope of the invention as broadly described. Therefore, the present embodiments are to be considered illustrative rather than restrictive in all respects.
Claims
1. A method for detecting anomalies in a substance, the method comprising: A first set of chemical fingerprints is obtained, wherein each chemical fingerprint in the first set of chemical fingerprints indicates multiple physicochemical properties of each sample in a set of normal samples of the substance; The first set of chemical fingerprints is converted into a cluster of data points in a multidimensional principal component analysis (PCA) diagram, wherein each dimension of the multidimensional PCA diagram is based on a principal component (PC), and each PC corresponds to one of the multiple physicochemical properties. The contour pattern of the data point clusters is used to construct a prediction model, which is configured to identify new samples with chemical fingerprints falling outside the contour pattern as anomalies; and The prediction model is optimized using a second set of chemical fingerprints, wherein the second set of chemical fingerprints indicates multiple physicochemical properties of a set of test samples comprising multiple normal test samples and multiple abnormal test samples of the substance. The construction of the contour pattern includes: For each data point in the cluster, calculate the square root of the sum of all squared PCs; Based on the corresponding value of the square root of the calculated sum of the data points, the data points are clustered into a predetermined number of intervals; and The contour pattern in the PCA diagram is obtained by removing data points from the cluster that fall into one or more permutations with more than a predetermined number of data points.
2. The method according to claim 1, characterized in that, The predetermined number of intervals is 20, and the predetermined number of data points is 1000.
3. The method according to claim 1, characterized in that, The optimization of the prediction model includes: For the set of test samples, the second set of chemical fingerprints is converted into new data points in the PCA graph; and For each sample in the set of test samples, the squared Mahalanobis distance (MD) score between the sample and the centroid of the contour pattern in the PCA diagram is calculated, and a threshold squared MD score is determined to distinguish whether the sample falls inside or outside the contour pattern.
4. The method according to claim 3, characterized in that, Determining the threshold squared MD score includes: Arrange the squared MD scores of the aforementioned set of test samples; Iterate through a series of squared MD scores as the threshold in sensitivity calculation; and The squared MD score with the highest sensitivity and an overall accuracy of over 95% among the iterative series of squared MD scores is determined as the threshold squared MD score.
5. The method according to claim 4, characterized in that, The sensitivity value is calculated based on the following equation:
6. The method according to claim 4, characterized in that, The overall accuracy is calculated based on the following equation:
7. The method according to claim 4, characterized in that, Also includes: The prediction model is trained based on hierarchical clustering learning of the spectral data of the set of test samples to define the true normal sample.
8. The method according to claim 4, characterized in that, Also includes: The prediction model is trained to optimize the overall accuracy based on Extratree or XGBoost learning of multiple physicochemical properties of the set of test samples.
9. The method according to claim 1, characterized in that, The substance in question is milk.
10. An apparatus for detecting anomalies in a substance, the apparatus comprising: At least one processor; as well as The memory includes computer program code for execution by the at least one processor, wherein the computer program code instructs the at least one processor to: A first set of chemical fingerprints is obtained, wherein each chemical fingerprint in the first set of chemical fingerprints indicates multiple physicochemical properties of each sample in a set of normal samples of the substance; The first set of chemical fingerprints is converted into a cluster of data points in a multidimensional principal component analysis (PCA) diagram, wherein each dimension of the multidimensional PCA diagram is based on a principal component (PC), and each PC corresponds to one of the multiple physicochemical properties. The contour pattern of the data point clusters is used to construct a prediction model, which is configured to identify new samples with chemical fingerprints falling outside the contour pattern as anomalies; and The prediction model is optimized using a second set of chemical fingerprints, wherein the second set of chemical fingerprints indicates multiple physicochemical properties of a set of test samples comprising multiple normal test samples and multiple abnormal test samples of the substance. During the construction of the contour pattern, the computer program code further instructs the at least one processor to: For each data point in the cluster, calculate the square root of the sum of all squared PCs; Based on the corresponding value of the square root of the calculated sum of the data points, the data points are clustered into a predetermined number of intervals; and The contour pattern in the PCA diagram is obtained by removing data points from the cluster that fall into one or more permutations with more than a predetermined number of data points.
11. The apparatus according to claim 10, characterized in that, The predetermined number of intervals is 20, and the predetermined number of data points is 1000.
12. The apparatus according to claim 10, characterized in that, During the optimization of the prediction model, the computer program code further instructs the at least one processor to: For the set of test samples, the second set of chemical fingerprints is converted into new data points in the PCA diagram; as well as For each sample in the set of test samples, the squared Mahalanobis distance (MD) score between the sample and the centroid of the contour pattern in the PCA diagram is calculated, and a threshold squared MD score is determined to distinguish whether the sample falls inside or outside the contour pattern.
13. The apparatus according to claim 12, characterized in that, During the determination of the threshold squared MD score, the computer program code further instructs the at least one processor to: Arrange the squared MD scores of the aforementioned set of test samples; A series of squared MD scores are used as the threshold in sensitivity calculation; as well as The squared MD score with the highest sensitivity and an overall accuracy of over 95% among the iterative series of squared MD scores is determined as the threshold squared MD score.
14. The apparatus according to claim 13, characterized in that, The sensitivity value is calculated based on the following equation:
15. The apparatus according to claim 13, characterized in that, The overall accuracy is calculated based on the following equation:
16. The apparatus according to claim 13, characterized in that, The computer program code further instructs the at least one processor to: The prediction model is trained based on hierarchical cluster learning of the spectral data of the set of test samples to define the true normal sample.
17. The apparatus according to claim 10, characterized in that, The substance in question is milk.
18. A non-transitory computer-readable storage medium having instructions encoded thereon, which, when executed by a processor, cause the processor to perform one or more steps of the method for anomaly detection of a substance according to claim 1.
Citation Information
Patent Citations
Water quality abnormal event detection method based on spectral statistical characteristics
CN104777115A
Anomaly detection apparatus, method and computer-readable medium
WO2020044495A1