Element traceability system and method

By calculating the nucleotide sequence distance and statistical analysis, the shortcomings of element traceability in existing technologies are solved, and accurate traceability and safety control of product ingredients are achieved.

CN120641985APending Publication Date: 2025-09-12GENETIC IDENTIFICATION LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480009872.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-30
Filing Date
2024-01-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing element traceability solutions are not effective in detecting the presence in a given product of materials originating from a given source, materials under specific conditions, or materials that should not be present.

Method used

Determine whether a given source contributed to the formation of a product or detect non-compliance in a product by calculating nucleotide-based sequence distances and applying two-sample t-tests, linear regression models, and QQ plots.

Benefits of technology

It achieves accurate traceability of product ingredients, improves the effectiveness of product safety control, ensures that products meet market and regulatory requirements, and reduces the risk of security vulnerabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120641985A_ABST
    Figure CN120641985A_ABST
Patent Text Reader

Abstract

The subject of the invention is to provide an element traceability system and method. The element traceability system includes processing circuitry configured to perform at least one of: (i) detect a presence of a material originating from a given source in a given product; (ii) detecting the presence of a material in a given product originating from a source subjected to specific conditions, such as a disease; and / or (iii) detecting the presence or absence of a material that should not be contained in a given product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of element tracing. Background Art

[0002] Element traceability is a general term encompassing systems and methods used to identify, track and / or trace the movement of an object, characteristic or information through a process (e.g., tracking the movement of raw materials through a production process where they are used to produce a product). Traceability is particularly useful as a key tool for implementing standards and regulations, improving product safety controls (e.g., enabling public and private actors to verify that products meet market and regulatory requirements), and helping to address security breaches.

[0003] Existing element traceability solutions remain insufficient because they fail to effectively address the following problems: (i) detecting the presence of materials originating from a given source in a given product; (ii) detecting the presence of materials originating from a source that has been subjected to specific conditions (e.g., disease, etc.) in a given product; and / or (iii) detecting the presence of materials that should not be present in a given product.

[0004] Therefore, there is a need in the art for new and improved element tracing systems and methods. Summary of the Invention

[0005] According to a first aspect of the disclosed subject matter, there is provided a system for determining whether a given source contributed to the formation of a given product, the system comprising processing circuitry configured to: obtain: (i) a first nucleotide-based sequence derived from the given source, (ii) a first set of nucleotide-based sequences, each of which is derived from a source from a plurality of sources that are part of a group of individuals, not all of which are sources, (iii) a second set of nucleotide-based sequences, each of which is derived from the given product, and (iv) a second nucleotide-based sequence derived from a source that is part of the group of individuals that is not known to be included in the given product; calculate a first distance associated with the given source, wherein the first distance is a distance comprised of: (i) a distance of the first nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the first nucleotide-based sequence from the second set of nucleotide-based sequences; calculating a second distance associated with a source not included in the plurality of sources, wherein the second distance is comprised of: (i) a distance of the second nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the second nucleotide-based sequence from the second set of nucleotide-based sequences; and determining that the given source contributed to the formation of the product when a difference between the first distance and the second distance is greater than a threshold value.

[0006] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, a two-sample t-test is applied after calculating the first distance and the second distance.

[0007] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, the first set of nucleotide-based sequences and the second set of nucleotide-based sequences each consist of single nucleotide polymorphism DNA sequences.

[0008] According to a second aspect of the presently disclosed subject matter, a system for detecting non-compliance in a given product produced by a source group is provided, the system comprising processing circuitry configured to: obtain: (i) a group of nucleotide-based sequences from each given source in the source group, wherein the group of nucleotide-based sequences comprises a collection of nucleotide-based sequences common to all sources in the source group, and (ii) a group of nucleotide-based sequences from each given source in the source group in the given product; for each given nucleotide-based sequence in the group of nucleotide-based sequences, perform the following operations: (a) generate a linear regression model based on a comparison of the allele frequencies of the given nucleotide-based sequences in each given source in the source group and the allele frequencies of the given nucleotide-based sequences in each given source in the given product; (b) obtain residuals from the linear regression model, wherein the residuals are random errors of the model; and (c) determine that non-compliance exists in the given product when the residuals exceed a threshold.

[0009] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, the random error of the model is defined as the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

[0010] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, said groups of nucleotide-based sequences each consist of a single nucleotide polymorphism DNA sequence.

[0011] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, non-compliance is also determined to exist when the residuals of all nucleotide-based sequences in the set of nucleotide-based sequences are non-normally distributed.

[0012] In one embodiment of the presently disclosed subject matter and / or embodiments thereof: (i) the threshold is defined according to a tolerance threshold, and (ii) the tolerance threshold is the percentage of sources in the source group that lack at least one nucleotide-based sequence in the group of nucleotide-based sequences.

[0013] According to a third aspect of the subject matter of the present disclosure, a system for detecting whether a product produced from a source group contains a substance from a disease source is provided, the system comprising a processing circuit configured to: obtain: (i) a group of nucleotide-based sequences from each given source in the source group, wherein the group of nucleotide-based sequences comprises a collection of nucleotide-based sequences common to all sources in the source group, and (ii) a group of nucleotide-based sequences from each given source in the source group of a given dairy product; for each given nucleotide-based sequence in the group of nucleotide-based sequences, perform the following operations: (a) generate a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source in the source group with the allele frequency of the given nucleotide-based sequence in each given source in the given dairy product; (b) obtain residuals from the linear regression model, wherein the residuals are random errors of the model; generate a QQ plot based on the residuals of the nucleotide-based sequences in the group of nucleotide-based sequences; and determine that the product contains a substance from a disease source when at least one side of the QQ plot is nonlinear.

[0014] In one embodiment of the disclosed subject matter and / or embodiments thereof, prior to generating the QQ-plot, the processing circuit is configured to determine whether the residuals are normally distributed and when the distribution is non-normal, the processing circuit moves to the QQ-plot generation step.

[0015] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, said groups of nucleotide-based sequences each consist of a single nucleotide polymorphism DNA sequence.

[0016] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, the random error of the model is defined as the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

[0017] According to a fourth aspect of the presently disclosed subject matter, there is provided a method for determining whether a given source contributed to the formation of a given product, the method comprising: obtaining: (i) a first nucleotide-based sequence derived from the given source, (ii) a first set of nucleotide-based sequences, each of which is derived from a source from a plurality of sources that are part of a group of individuals, not all of which are sources, (iii) a second set of nucleotide-based sequences, each of which is derived from the given product, and (iv) a second nucleotide-based sequence derived from a source known not to be included in the given product, which is part of the group of individuals; calculating a first distance associated with the given source, wherein the first distance a distance comprised of: (i) a distance of the first nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the first nucleotide-based sequence from the second set of nucleotide-based sequences; calculating a second distance associated with a source not included in the plurality of sources, wherein the second distance is comprised of: (i) a distance of the second nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the second nucleotide-based sequence from the second set of nucleotide-based sequences; and determining that the given source contributed to the formation of the product when the difference between the first distance and the second distance is greater than a threshold value.

[0018] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, a two-sample t-test is applied after calculating the first distance and the second distance.

[0019] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, the first set of nucleotide-based sequences and the second set of nucleotide-based sequences each consist of single nucleotide polymorphism DNA sequences.

[0020] According to a fifth aspect of the subject matter of the present disclosure, a method for detecting non-compliance in a given product produced by a source group is provided, the method comprising: obtaining: (i) a group of nucleotide-based sequences from each given source in the source group, wherein the group of nucleotide-based sequences comprises a collection of nucleotide-based sequences common to all sources in the source group, and (ii) a group of nucleotide-based sequences from each given source in the source group of the given product; for each given nucleotide-based sequence in the group of nucleotide-based sequences, performing the following operations: (a) generating a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source of the source group and the allele frequency of the given nucleotide-based sequence in each given source in the given product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; and (c) determining that non-compliance exists in the given product when the residuals exceed a threshold value.

[0021] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, the random error of the model is defined as the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

[0022] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, said groups of nucleotide-based sequences each consist of a single nucleotide polymorphism DNA sequence.

[0023] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, non-compliance is also determined to exist when the residuals of all nucleotide-based sequences in the set of nucleotide-based sequences are non-normally distributed.

[0024] In one embodiment of the presently disclosed subject matter and / or embodiments thereof: (i) the threshold is defined according to a tolerance threshold, and (ii) the tolerance threshold is the percentage of sources in the source group that lack at least one nucleotide-based sequence in the group of nucleotide-based sequences.

[0025] According to a sixth aspect of the subject matter of the present disclosure, a method for detecting whether a product produced from a source group contains a substance from a disease source is provided, the method comprising: obtaining: (i) a group of nucleotide-based sequences from each given source in the source group, wherein the group of nucleotide-based sequences comprises a collection of nucleotide-based sequences common to all sources in the source group, and (ii) a group of nucleotide-based sequences from each given source in the source group of a given dairy product; for each given nucleotide-based sequence in the group of nucleotide-based sequences, performing the following operations: (a) generating a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source in the source group with the allele frequency of the given nucleotide-based sequence in each given source in the given dairy product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; generating a QQ plot based on the residuals of the nucleotide-based sequences in the group of nucleotide-based sequences; and determining that the product contains a substance from a disease source when at least one side of the QQ plot is nonlinear.

[0026] In one embodiment of the disclosed subject matter and / or embodiments thereof, prior to generating the QQ-plot, the processing circuit is configured to determine whether the residuals are normally distributed, and when the distribution is non-normal, the processing circuit moves to the QQ-plot generation step.

[0027] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, said groups of nucleotide-based sequences each consist of a single nucleotide polymorphism DNA sequence.

[0028] In one embodiment of the presently disclosed subject matter and / or embodiments thereof, the random error of the model is defined as the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

[0029] According to a seventh aspect of the presently disclosed subject matter, there is provided a non-transitory computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by at least one processor to implement a method for determining whether a given source contributed to the formation of a given product, the method comprising: obtaining: (i) a first nucleotide-based sequence derived from the given source, (ii) a first set of nucleotide-based sequences, each of which is derived from a source from a plurality of sources that are part of a group of individuals, not all of which are sources, (iii) a second set of nucleotide-based sequences, each of which is derived from the given product, and (iv) a second nucleotide-based sequence derived from a source known to be not included in the given product, which is a source of the group of individuals portion; calculating a first distance associated with the given source, wherein the first distance is comprised of: (i) a distance of the first nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the first nucleotide-based sequence from the second set of nucleotide-based sequences; calculating a second distance associated with a source not included in the plurality of sources, wherein the second distance is comprised of: (i) a distance of the second nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the second nucleotide-based sequence from the second set of nucleotide-based sequences; and determining that the given source contributed to the formation of the product when a difference between the first distance and the second distance exceeds a threshold value.

[0030] According to an eighth aspect of the presently disclosed subject matter, there is provided a non-transitory computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by at least one processor to implement a method for detecting non-compliance in a given product produced by a source group, the method comprising: obtaining: (i) a group of nucleotide-based sequences from each given source in the source group, wherein the group of nucleotide-based sequences comprises a collection of nucleotide-based sequences common to all sources in the source group, and (ii) a group of nucleotide-based sequences from each given source in the source group for the given product; for each given nucleotide-based sequence in the group of nucleotide-based sequences, performing the following operations: (a) generating a linear regression model based on a comparison of the allele frequencies of the given nucleotide-based sequences in each given source in the source group with the allele frequencies of the given nucleotide-based sequences in each given source in the given product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; and (c) determining that the given product is non-compliant when the residuals exceed a threshold.

[0031] According to a ninth aspect of the presently disclosed subject matter, there is provided a non-transitory computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by at least one processor to implement a method for detecting whether a product produced from a source group contains material from a disease source, the method comprising: obtaining: (i) a set of nucleotide-based sequences from each given source in the source group, wherein the set of nucleotide-based sequences comprises a collection of nucleotide-based sequences common to all sources in the source group, and (ii) a set of nucleotide-based sequences from each given source in the source group for a given dairy product; For each given nucleotide-based sequence in the group of nucleotide-based sequences, the following operations are performed: (a) generating a linear regression model based on a comparison of the allele frequencies of the given nucleotide-based sequence in each given source in the source group and the allele frequencies of the given nucleotide-based sequence in each given source in the given dairy product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; generating a QQ plot based on the residuals of the nucleotide-based sequences in the group of nucleotide-based sequences; and determining that the product contains material from a disease source when at least one side of the QQ plot is nonlinear. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to understand the disclosed subject matter and to see how it may be carried into effect, it will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:

[0033] Figure 1is a schematic diagram of an environment in which an element tracing system according to the presently disclosed subject matter operates;

[0034] Figure 2 A block diagram schematically illustrating an embodiment of an element tracing system according to the present disclosure; and

[0035] Figure 3 is a flow chart illustrating an embodiment of a sequence of operations performed by an element traceability system according to the presently disclosed subject matter;

[0036] Figure 4 is a schematic diagram of an element tracing process performed by an element tracing system according to the presently disclosed subject matter;

[0037] Figure 5 is a flow chart illustrating another embodiment of a sequence of operations performed by an element tracing system according to the presently disclosed subject matter;

[0038] Figure 6 A schematic diagram of a linear regression model for a given nucleotide-based sequence generated by an element traceability system according to the presently disclosed subject matter;

[0039] Figure 7 is a flow chart illustrating yet another embodiment of a sequence of operations performed by an element tracing system according to the presently disclosed subject matter; and

[0040] Figure 8A and 8B A schematic diagram of an exemplary QQ (quantile-quantile) plot generated by an element traceability system according to the presently disclosed subject matter. DETAILED DESCRIPTION

[0041] In the following detailed description, numerous specific details are set forth to provide a full understanding of the disclosed subject matter. However, it will be understood by those skilled in the art that the disclosed subject matter can be practiced without these specific details. In other cases, well-known methods, procedures, and components have not been described in detail to avoid obscuring the disclosed subject matter.

[0042] In the accompanying drawings and description, like reference numerals indicate those components that are common to the different embodiments or configurations.

[0043] Unless expressly stated otherwise, it will be apparent from the following discussion that terms such as "obtain," "calculate," "determine," "execute," "generate," etc., used throughout the specification include computer actions and / or processes that manipulate and / or transform data into other data, the data being represented as physical quantities, such as electronic quantities, and / or the data representing physical objects. The terms "computer," "processor," "processing resource," "processing circuitry," and "controller" should be broadly interpreted to encompass any electronic device with data processing capabilities, including, by way of non-limiting example, personal desktop / laptop computers, servers, computing systems, communication devices, smartphones, tablets, smart TVs, processors (such as digital signal processors (DSPs), microcontrollers, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.), a group of multiple physical machines sharing multi-tasking capabilities, virtual servers co-resident on a single physical machine, any other electronic computing device, and / or any combination thereof.

[0044] Operations according to the teachings herein may be performed by a computer specially constructed for the required purpose, or by a general-purpose computer specially configured for the required purpose using a computer program stored in a non-transitory computer-readable storage medium. The term "non-transitory" is used herein to exclude transient, propagating signals, but otherwise includes any volatile or non-volatile computer memory technology suitable for the application.

[0045] As used herein, the phrases "for example," "such as," "for instance," and variations thereof are used to describe non-limiting embodiments of the presently disclosed subject matter. Reference in the specification to "one embodiment," "some embodiments," "other embodiments," or variations thereof indicates that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the presently disclosed subject matter. Thus, the appearance of the phrases and "one embodiment," "some embodiments," "other embodiments," or variations thereof do not necessarily refer to the same embodiment.

[0046] It will be appreciated that, unless expressly stated otherwise, certain features of the disclosed subject matter described in the context of separate embodiments for clarity may also be provided in combination in a single embodiment. Conversely, various features of the disclosed subject matter described in the context of a single embodiment for brevity may also be provided individually or in any suitable subcombination.

[0047] In an embodiment of the presently disclosed subject matter, Figure 3 、 5 Fewer, more, and / or different stages than those shown in FIG. 7 may be performed. In the disclosed subject matter, Figure 3 、 5One or more of the stages shown in Figures 1 and 7 may be performed in a different order and / or one or more groups of stages may be performed simultaneously. Figure 1 A general schematic diagram of a system architecture according to an embodiment of the disclosed subject matter is presented. Figure 2 Each module may be composed of any combination of software, hardware and / or firmware that performs the functions defined and explained herein. Figure 2 The modules in the system can be centralized in one location or dispersed across multiple locations. In other embodiments of the disclosed subject matter, the system can include Figure 2 Fewer, more and / or different modules than shown.

[0048] Any reference in the specification to a method should apply mutatis mutandis to a system capable of performing the method and should apply mutatis mutandis to a non-transitory computer-readable medium storing instructions that, once executed by a computer, result in performance of the method.

[0049] Any reference in the specification to a system shall apply mutatis mutandis to a method that may be performed by the system and to a non-transitory computer-readable medium storing instructions that may be executed by the system.

[0050] Any reference in the specification to a non-transitory computer-readable medium shall apply mutatis mutandis to a system capable of executing instructions stored in the non-transitory computer-readable medium and shall apply mutatis mutandis to a method executable by a computer reading instructions stored in the non-transitory computer-readable medium.

[0051] Any reference in this specification to "element traceability" shall apply mutatis mutandis to the ability to identify, track and trace a product, ingredient, raw material, part of a product, characteristic, etc. as it moves through a process, as described below in conjunction with Figure 1 Described in detail.

[0052] For this reason, see Figure 1 , which shows a schematic diagram of an environment in which an element tracing system according to the presently disclosed subject matter operates.

[0053] As shown schematically, environment 100 includes a process 102 comprised of a plurality of stages, labeled "1" through "N" (where N is an integer representing any number of stages). Process 102 represents a path taken by a given element (e.g., a product, ingredient, raw material, portion of a product, etc.), and may extend between an initial stage (stage "1"), where the given element may be derived from a source 104, and a final stage (stage "N"), where the given element (either by itself or in the form of a product produced therefrom) reaches its intended destination (e.g., a point of sale where the given element is used / utilized / consumed, etc.).

[0054] Note that source 104 can be a single source (e.g., a single animal (e.g., a cow, goat, etc.), a single human, etc.) or a group of sources (e.g., a herd of animals (e.g., a herd of cows, goats, etc.), a farm, a group of farms, a group of people, etc.). Furthermore, note that source 104 can be of different types, such as human, animal (e.g., cows, goats, etc.), plant, etc., and therefore, the derivation method for a given element may need to be adjusted depending on the source type.

[0055] The stages comprising process 102 can be divided into two or more subroutes, and at the end of each subroutes, a given element can be located at a designated point. For example, process 102 can be divided into: (i) a first subroutes, referred to as production subroutes, during which the given element can undergo various production or manufacturing processes under specialized conditions, such as pasteurization, sterilization, fermentation, grinding, etc., until it reaches its designated form (e.g., a product form); and (ii) a second subroutes, referred to as transportation subroutes, during which the designated form of the given element undergoes various transportation processes, such as sorting, packaging, shipping, etc., until the given element in its designated form reaches its designated destination (e.g., a point of sale where the designated form of the given element is used / utilized / consumed, etc.).

[0056] As a non-limiting example (which is only used to better understand the disclosed subject matter and is not intended to limit its scope in any way), Figure 1 As shown, process 102 is a supply chain that represents the path taken by dairy products (e.g., cheese, butter, yogurt, ice cream, milk, condensed milk, and milk powder, etc.) or parts of dairy products (i.e., products containing dairy ingredients, such as cheese, milk, etc.). Process 102 extends from an initial stage (stage "1"), in which raw milk is milked from a herd of cows (represented by an image of a pair of cows), to a final stage (stage "N"), in which dairy products produced from the raw milk from the herd arrive at a dairy store 106 and are sold there.

[0057] The dairy product's path includes: (i) a dairy product production sub-route, which extends from the initial stage (stage "1") to the stage where the dairy product is bottled and ready for sale (stage "3"); and (ii) a dairy product transportation sub-route, which extends from the stage where the bottles ready for sale are packed into boxes (stage "4") to the final stage where the bottles arrive at the dairy store 106 and are available for purchase by end users (stage "N").

[0058] During the dairy production sub-route, the raw milk obtained from the herd of cows undergoes a sterilization and pasteurization process, at the end of which the sterilized and pasteurized milk is bottled in the bottles, ready for transportation.

[0059] During the product transportation sub-route, the bottles are packed into designated boxes and transported by designated trucks to the dairy store 106 where they can be purchased by the end user.

[0060] It is noted that the supply chain 102 and each of its subroutes may include fewer or more stages than those described above, depending on the number of steps involved therein. It is further noted that the supply chain 102 and each of its subroutes may include different stages than those described above, which shall apply mutatis mutandis.

[0061] Throughout the process 102 described above, fraud, failures, errors, etc. may occur at any stage, which may potentially affect the composition of the specified form of a given element. The specific stage at which such a situation (which may endanger the well-being of the end user) occurs may be difficult to detect and therefore difficult to resolve. In order to ensure that the composition of the specified form of a given element only contains ingredients derived from sources that are expected to contribute to its composition, an element traceability system according to the subject matter of the present disclosure will be operated, as described below in conjunction with Figure 3 and Figure 5 As stated.

[0062] Attention is now directed to a description of the components of the element tracing system 200 .

[0063] Figure 2 is a block diagram schematically illustrating one embodiment of an element tracing system 200 according to the disclosed subject matter.

[0064] According to the disclosed subject matter, an element traceability system 200 (also interchangeably referred to herein as "system 200") may include a network interface 206. The network interface 206 (e.g., a network card, a Wi-Fi client, a Li-Fi client, a 3G / 4G client, or any other component) enables the system 200 to communicate with external systems over a network and to process inbound and outbound communications from such systems. For example, the system 200 may receive, via the network interface 206, one or more sets of nucleotide-based sequences / a collection of one or more nucleotide-based sequences, the nucleotide-based sequences being derived from one or more sources and / or one or more products.

[0065] The system 200 may further include or be associated with a data repository 204 (e.g., a database, a storage system, a memory including read-only memory (ROM), random access memory (RAM), or any other type of memory, etc.) configured to store data. Some examples of data that may be stored in the data repository 204 include:

[0066] one or more first nucleotide-based sequences derived from one or more sources that are expected to contribute to the formation of one or more finished products;

[0067] one or more second nucleotide-based sequences derived from one or more sources not intended to contribute to the formation of one or more finished products;

[0068] one or more first distances associated with one or more given sources expected to contribute to the formation of one or more finished products;

[0069] one or more second distances associated with one or more given sources that unintentionally contribute to the formation of one or more finished products;

[0070] one or more linear regression models associated with the one or more nucleotide-based sequences;

[0071] One or more residuals, each associated with a linear regression model, representing the random error of that model;

[0072] One or more tolerance thresholds;

[0073] • One or more QQ (quantile-quantile) plots of the residuals of the respective linear regression models; etc.

[0074] The data repository 204 may be further configured to enable retrieval and / or updating and / or deletion of stored data. Note that in some cases, the data repository 204 may be distributed, and the system 200 (e.g., via a wired or wireless network to which the system 200 can connect, using its network interface 206) may be able to access information stored thereon.

[0075] The system 200 also includes processing circuitry 202. The processing circuitry 202 may be one or more processing units (e.g., central processing units), microprocessors, microcontrollers (e.g., microcontroller units (MCUs)), or any other computing device or module, including multiple and / or parallel and / or distributed processing units, adapted to process data independently or collaboratively to control relevant system 200 resources and implement operations related to the resources of the system 200.

[0076] The processing circuit 202 includes an element tracing module 208, which is configured to perform an element tracing process as described below (especially with reference to Figure 3 、 Figure 5 and Figure 7 ) is further described.

[0077] Go to Figure 3 , which shows a flow chart illustrating one example of operations performed by an element tracing system 200 according to the disclosed subject matter.

[0078] Thus, the element tracing system 200 (hereinafter also interchangeably referred to as "system 200") can be configured to perform an element tracing process 300, for example, using the element tracing module 208. The element tracing process 300 is intended to detect the presence of raw materials originating from a given source in a given finished product, and thereby determine whether the given source contributed to the formation of the given finished product.

[0079] To this end, the system 200 obtains: (i) a first nucleotide-based sequence derived from a given source (e.g., a vector of allele frequencies for the given source), (ii) a first set of nucleotide-based sequences, each of which is derived from a source in a plurality of sources (e.g., a vector of allele frequencies for the plurality of sources, such as an aggregated sum of the plurality of sources), the plurality of sources optionally being part of a larger group of individuals, not all of which serve as sources of raw materials for producing the given finished product, (iii) a second set of nucleotide-based sequences, each of which is derived from a given product (e.g., a vector of allele frequencies for the given product), and (iv) a second nucleotide-based sequence derived from a source known not to be included in the given product (e.g., a vector of allele frequencies for the source known not to be included in the given product), which is part of the larger group of individuals (box 302).

[0080] Note that for the vectors of allele frequencies for (i) to (iv), we only need their genetic profiles. More specifically, we only need: the genetic profile of a given source, the genetic profiles of multiple sources (a single vector; we do not need access to the genetic profiles of all of the multiple sources), the genetic profile of a given product (again, a single vector), and the genetic profiles of individuals known to be present in the group but not in the given product.

[0081] In one embodiment, each of the sequences based on Nucleotide of (i) to (iv) can be a DNA sequence dna with a germline replacement of a single Nucleotide at a specific position in the sequence (i.e., a single nucleotide polymorphism (SNP) DNA sequence). In another embodiment, each of the sequences based on Nucleotide can be a DNA sequence dna with a germline replacement of two or more Nucleotide at a specific position in the sequence.

[0082] In one non-limiting example, the plurality of sources from which the first set of nucleotide-based sequences is obtained may be about Figure 1 The source group mentioned above is obtained from which the raw materials for the production of the finished product are obtained; and the second set of nucleotide-based sequences can be obtained from about Figure 1Furthermore, a given source from which the first nucleotide-based sequence is obtained may or may not be a member of the group of sources from which the first set of nucleotide-based sequences is obtained.

[0083] As non-limiting examples (which are only used to better understand the disclosed subject matter and are not intended to limit its scope in any way), Figure 4 As shown, the system 200 obtains: (i) SNP sequences derived from a cow (denoted as yj), (ii) a first set of SNP sequences, each of which is derived from a cow in a plurality of cows (denoted as Popj), the plurality of cows being part of a larger group of cows, and (iii) a second set of SNP sequences, each of which is derived from a dairy product, denoted as M j , and (iv) SNP sequences, denoted as y'j, derived from cows known not to be included in the milk product, which cows are included in the larger group of cows.

[0084] The system 200 calculates a first distance associated with a given source, the first distance consisting of: (i) a distance of a first nucleotide-based sequence from a first set of nucleotide-based sequences, and (ii) a distance of the first nucleotide-based sequence from a second set of nucleotide-based sequences (block 304).

[0085] The first nucleotide-based sequence can be, for example, a Euclidean distance (although other forms of mathematical distances may also be applicable) from the first set to effectively measure the similarity or dissimilarity of an individual or colony to another individual or colony. For example, in the case of two individuals, "individual A" and "individual B", the genotypes of the two individuals can be converted to a digital form of 0, 0.5 or 1, which represents the number of minor alleles that the individual has at each gene position divided by 2. After conversion, the absolute difference of the digital genetic value at each gene position is calculated, and all absolute differences at each gene position are used to calculate the mean difference, which represents the distance metric (ranging from 0 to 1) between "individual A" and "individual B". The closer the distance metric is to 1, the greater the difference between "individual A" and "individual B". The closer the distance metric is to 0, the more similar "individual A" and "individual B" are.

[0086] According to us Figure 4 In a non-limiting example, the system 200 calculates a first distance associated with cow yj. The first distance is composed of: (i) the distance of the SNP sequence from cow yj to the first set of SNP sequences from multiple cows Popj, and (ii) the distance of the SNP sequence from cow yj to the first set of SNP sequences from milk product M. j distances to the second set of SNP sequences.

[0087] The system 200 then calculates a second distance associated with a source known not to be included in the given product, the distance being comprised of: (i) the distance of the second nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) the distance of the second nucleotide-based sequence from the second set of nucleotide-based sequences (block 306).

[0088] according to Figure 4 In a non-limiting example, the system 200 calculates a second distance associated with cow y'j. The second distance is composed of: (i) the distance of the SNP sequence from cow y'j to the first set of SNP sequences from multiple cows Popj, and (ii) the distance of the SNP sequence from cow y'j to the first set of SNP sequences from dairy product M. j distances to the second set of SNP sequences.

[0089] When the difference between the first distance and the second distance is greater than a threshold, the system 200 determines that the given source contributed to the formation of the product (block 308).

[0090] according to Figure 4 In a non-limiting example, the system 200 determines that cow yj contributes to the milk product M j is formed because the difference between the first distance associated with cow yj and the second distance associated with cow y'j is greater than a predefined threshold.

[0091] In some cases, after calculating the first distance and the second distance, to assess the likelihood that a genomic profile from a given source is present in the given product, a two-sample t-test may be applied.

[0092] It is important to note that a second nucleotide-based sequence derived from a source known not to be present in the given product can be used not only to generate an optional multidimensional test statistic, as explained above, but can also be used as a control group to improve the sensitivity and effectiveness of the system. In more detail, by incorporating the variability of the two groups, the optional two-sample t-test can provide a more comprehensive understanding of the discreteness and overlap of the two distributions and improve the sensitivity of detecting true differences between the groups, thereby providing deeper insight into whether a particular individual is present. Comparing two different conditions (present in a given finished product versus not present in a given finished product) can make the analysis more robust to individual-specific variation and can potentially reduce the tendency to make type I and type II errors. In addition, comparing the genotypic differences of two different individuals (relative to multiple sources and a given finished product) can provide more context and insight into the nature and significance of the observed differences.

[0093] Transfer back Figure 5 , which shows a flow chart illustrating another embodiment of operations performed by an element tracing system 200 according to the disclosed subject matter.

[0094] Thus, the element tracing system 200 (hereinafter also interchangeably referred to as "system 200") can be configured to perform the element tracing process 500, for example, using the element tracing module 208. The element tracing process 500 is intended to detect non-compliance in a given finished product produced by a source group, and thereby detect the presence of materials that should not be included in the given finished product.

[0095] To this end, the system 200 obtains: (i) a set of nucleotide-based sequences from each given source in a source group, and (ii) a set of nucleotide-based sequences from each given source in the source group for the given product (block 502). The set of nucleotide-based sequences comprises a collection of nucleotide-based sequences common to all sources in the source group. Each nucleotide-based sequence can be: a DNA sequence having a germline substitution of a single nucleotide at a specific position in the sequence (i.e., a single nucleotide polymorphism (SNP) DNA sequence), a DNA sequence having a germline substitution of two or more nucleotides at a specific position in the sequence, a DNA sequence of a specific DNA fragment with different copy numbers (copy number variation (CNV)), a DNA sequence with different numbers of short tandem repeats (STRs), etc.

[0096] As a non-limiting example (for the purpose of providing a better understanding of the disclosed subject matter only and not intended to limit its scope in any way), system 200 obtains: (i) a set of SNP sequences from each cow in a plurality of cows used as a source of raw materials for producing dairy products, and (iii) a set of SNP sequences from each cow in the plurality of cows from which the dairy products are produced.

[0097] Next, for the set of nucleotide-based sequences, the system 200 generates a multiple linear regression model based on a comparison of the allele frequencies of the given nucleotide-based sequence in each given source of the set of sources with the allele frequencies of the given nucleotide-based sequence in the given finished product from each given source (box 504(a)).

[0098] Multiple linear regression is a generalization of simple linear regression with more than one independent variable, and is a special case of the general linear model with only one dependent variable. Therefore, to simplify the explanation, Figure 6An exemplary simple linear regression model for a given nucleotide-based sequence in the set of nucleotide-based sequences is shown, represented by graph 600. Graph 600 includes: (i) an x-axis labeled 602 representing the values ​​of the allele frequency of the given nucleotide-based sequence in each given source in the set of sources, and (ii) a y-axis labeled 604 representing the values ​​of the allele frequency of the given nucleotide-based sequence in a given finished product for each source in the set of sources. Furthermore, graph 600 includes a plurality of points, each representing a point at which the values ​​of the allele frequency of a given nucleotide-based sequence for a given source in the set of sources intersect; and a line 606 representing the potential relationship between the two allele frequencies of the given nucleotide-based sequence. Line 606 is associated with a linear equation comprising a slope value and a y-intercept value.

[0099] According to our non-limiting embodiment, for each given SNP sequence in the group of SNP sequences, system 200 generates a linear regression model, represented by a graph similar to graph 600, which is based on a comparison of the allele frequency of the given SNP sequence in each given source in the source group and the allele frequency of the given SNP sequence in the produced dairy product for each given source.

[0100] In some cases, instead of implementing the multiple linear regression model approach described above, the system 200 may implement other approaches, such as Bayesian or non-parametric approaches.

[0101] Back to Figure 5 , the system 200 obtains a residual from the linear regression model, which is a random error of the model (e.g., the y-intercept value of the linear equation of the linear regression model) (block 504(b)). And when the residual exceeds a threshold, the system 200 determines that non-conformance exists in the given finished product (block 504(c)).

[0102] In some cases, the random error of the model can be defined as, for example, the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source in the given finished product.

[0103] In some cases, the threshold value can be a customized threshold value for each given nucleotide-based sequence, or a standard threshold value that applies to all nucleotide-based sequences. In addition, in some cases, the threshold value can be defined, for example, based on a tolerance threshold value, which is the percentage of sources in the source group that lack at least one given nucleotide-based sequence in the group of nucleotide-based sequences.

[0104] In some cases, in addition to or in lieu of the foregoing, non-compliance may also be determined when the residuals of all nucleotide-based sequences in the set of nucleotide-based sequences are non-normally distributed (in a mathematical sense).

[0105] According to our non-limiting embodiment, for each SNP sequence in the set of SNP sequences, the system 200 obtains the corresponding residual (i.e., the y-intercept value of the linear equation of the linear regression model) and compares it with a standard threshold. When one of the corresponding residuals associated with a given SNP sequence exceeds the predefined threshold, the system 200 determines that non-compliance exists in the given finished product.

[0106] Go to Figure 7 , which shows a flow chart illustrating yet another embodiment of operations performed by the element tracing system 200 according to the disclosed subject matter.

[0107] Accordingly, the element tracing system 200 (hereinafter also interchangeably referred to as "system 200") can be configured to perform the element tracing process 700, for example, using the element tracing module 208. The element tracing process 700 is intended to detect whether the finished product produced by the source group contains substances from one or more sources under specific conditions (e.g., disease, etc.).

[0108] For this purpose, system 200 obtains: (i) from the group of sequences based on nucleotides of each given source in the source group, and (ii) from the group of sequences based on nucleotides of each given source in the source group of a given finished product (frame 702). The group of sequences based on nucleotides can comprise the set of sequences based on nucleotides that are common to all sources of the source group. In one embodiment, each sequence based on nucleotides can be a DNA sequence (i.e., single nucleotide polymorphism (SNP) DNA sequence) with a germline substitution of a single nucleotide at a specific position in the sequence. In another embodiment, each sequence based on nucleotides can be a DNA sequence with a germline substitution of two or more nucleotides at a specific position in the sequence.

[0109] As a non-limiting example (for the purpose of providing a better understanding of the disclosed subject matter only and not intended to limit its scope in any way), system 200 obtains: (i) a set of SNP sequences from each cow in a plurality of cows used as a source of raw materials for producing dairy products, and (iii) a set of SNP sequences from each cow in the plurality of cows from which the dairy products are produced.

[0110] For each given nucleotide-based sequence in the set of nucleotide-based sequences, system 200 generates a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source in the set of sources with the allele frequency of the given nucleotide-based sequence in a given finished product for each given source (in a manner similar to Figure 5 and Figure 6 the manner described) (block 704(a)).

[0111] In accordance with our non-limiting embodiment, for each given SNP sequence in the group of SNP sequences, system 200 generates a linear regression model, represented by a graph similar to graph 600, which is based on a comparison of the allele frequency of the given SNP sequence in each given source in the source group and the allele frequency of the given SNP sequence in each given source in the produced dairy product.

[0112] From the linear regression model for each given nucleotide-based sequence, the system 200 obtains a residual, which can be, for example, a random error of the model (e.g., a y-intercept value of the linear equation of the linear regression model) (block 704(b)). In some cases, the random error of the model can be defined as, for example, the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the set of sources and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

[0113] Based on the residuals of all nucleotide-based sequences of the set of nucleotide-based sequences, the system 200 generates a QQ plot (quantile-quantile plot) (block 706). Figure 8A and Figure 8B Exemplary QQ plots 800 and 800' are shown, respectively, each representing the distribution of residuals for a group of nucleotide-based sequences. QQ plots 800 and 800' each include: (i) an x-axis labeled 802 representing the actual value of the residual for each nucleotide-based sequence, and (ii) a y-axis labeled 804 representing the expected value of the residual for each nucleotide-based sequence. In addition, QQ plots 800 and 800' include a plurality of points, each representing the intersection of the actual value and the expected value of the residual for each nucleotide-based sequence; and a line 806 representing the potential relationship between two residual values. Figure 8A As shown, most of the points in the QQ graph 800 are arranged along line 806, forming a normal distribution. Figure 8B As shown, a plurality of points of the QQ plot 800 deviate from the line 806 at its two edges, designated 808a and 808b, forming a non-normal distribution.

[0114] return Figure 7When at least one edge of the QQ graph is non-linear, the system 200 determines that the product contains a substance originating from a source that is subject to a particular condition (eg, a disease, etc.) (block 708 ).

[0115] In some cases, before generating the QQ plot, the system 200 determines whether the residuals for the set of nucleotide-based sequences are normally distributed; if the distribution is non-normal, the system 200 moves to the QQ plot generation step.

[0116] In some cases, alternative methods to generating a QQ plot can be used to determine whether a product contains a substance derived from a source subjected to specific conditions. In one embodiment, generating a QQ plot can be replaced by comparing the results to a predefined threshold representing the maximum absolute value of the residuals. In another embodiment, the use of this predefined threshold can be combined with an adjusted R-squared multiple regression analysis.

[0117] It should be noted that in addition to the uses of system 200 described above, system 200 can be used in any application involving complex sample mixtures. In one embodiment, system 200 can be used to identify and / or track contributors for research purposes (e.g., clinical trials, etc.). In another embodiment, system 200 can be used to identify and / or track contributors in the production of biopharmaceutical products, which may use animal by-products (e.g., gelatin, bovine serum albumin (BSA), etc.) from complex mixed sources.

[0118] Please note, refer to Figure 3 、 Figure 5 and Figure 7 , some of these blocks may be integrated into a combined block, or may be decomposed into several blocks, and / or additional blocks may be added. It should also be noted that some of these blocks are optional. It should also be noted that although the flowcharts are described with reference to system components that implement them, this is by no means binding, and these blocks may be performed by other components than those described herein.

[0119] It should be understood that the application of the disclosed subject matter is not limited to the details set forth in the specification or the contents shown in the drawings. The disclosed subject matter can be used in other embodiments and can be practiced and implemented in various ways. Therefore, it should be understood that the wording and terminology used in this specification are for descriptive purposes and should not be considered restrictive. Therefore, it will be understood by those skilled in the art that the concepts on which this disclosure is based can be easily used as the basis for designing other structures, methods and systems to achieve the several purposes of the currently disclosed subject matter.

[0120] It should also be understood that a system according to the presently disclosed subject matter can be implemented at least in part as a suitably programmed computer. Likewise, the presently disclosed subject matter contemplates a computer program readable by a computer to perform the disclosed methods. The presently disclosed subject matter further contemplates a machine-readable memory tangibly embodying a program of instructions executable by a machine for performing the disclosed methods.

Claims

1. A system for determining whether a given source contributes to a given product, the system comprising processing circuitry configured to: obtaining: (i) a first nucleotide-based sequence derived from the given source, (ii) a first collection of nucleotide-based sequences, each of which is derived from a source among a plurality of sources that are part of a group of individuals, not all of which are sources, (iii) a second collection of nucleotide-based sequences, each of which is derived from the given product, and (iv) a second nucleotide-based sequence derived from a source known not to be included in the given product that is part of the group of individuals; calculating a first distance associated with the given source, wherein the first distance is comprised of: (i) a distance of the first nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the first nucleotide-based sequence from the second set of nucleotide-based sequences; calculating a second distance associated with a source not included in the plurality of sources, wherein the second distance is comprised of: (i) a distance of the second nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the second nucleotide-based sequence from the second set of nucleotide-based sequences; as well as When a difference between the first distance and the second distance is greater than a threshold, it is determined that the given source contributed to forming the product. 2 . The system of claim 1 , wherein after calculating the first and second distances, a two-sample t-test is applied.

3. The system of claim 1, wherein the first set of nucleotide-based sequences and the second set of nucleotide-based sequences each consist of single nucleotide polymorphism DNA sequences.

4. A system for detecting non-conformance in a given product produced by a source group, the system comprising processing circuitry configured to: obtaining: (i) a set of nucleotide-based sequences from each given source in said set of sources, wherein the set of nucleotide-based sequences comprises the set of nucleotide-based sequences common to all sources in said set of sources, and (ii) a set of nucleotide-based sequences from each given source in said set of sources for said given product; For each given nucleotide-based sequence in the set of nucleotide-based sequences, the following operations are performed: (a) generating a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source of the set of sources with the allele frequency of the given nucleotide-based sequence in each given source of the given product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; and (c) When the residual exceeds a threshold, determining that non-compliance exists in the given product.

5. The system of claim 4, wherein the random error of the model is defined as the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

6. The system of claim 4, wherein the groups of nucleotide-based sequences each consist of a single nucleotide polymorphism DNA sequence.

7. The system of claim 4, wherein non-compliance is also determined to exist when the residuals of all nucleotide-based sequences in the group of nucleotide-based sequences are non-normally distributed.

8. The system of claim 4, wherein: (i) the threshold is defined according to a tolerance threshold, and (ii) the tolerance threshold is the percentage of sources in the set of sources that lack at least one nucleotide-based sequence in the set of nucleotide-based sequences.

9. A system for detecting whether a product produced by a source group contains material from a disease source, the system comprising processing circuitry configured to: obtaining: (i) a set of nucleotide-based sequences from each given source in said group of sources, wherein the set of nucleotide-based sequences comprises the set of nucleotide-based sequences common to all sources in said group of sources, and (ii) a set of nucleotide-based sequences from each given source in said group of sources of a given dairy product; For each given nucleotide-based sequence in the set of nucleotide-based sequences, the following operations are performed: (a) generating a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source of the source group and the allele frequency of the given nucleotide-based sequence in each given source of the given dairy product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; generating a QQ plot based on the residuals of the nucleotide-based sequences in the set of nucleotide-based sequences; and, When at least one side of the QQ plot is nonlinear, it is determined that the product contains substances from a disease source.

10. The system of claim 9, wherein before generating the QQ plot, the processing circuit is configured to determine whether the residuals are normally distributed, and when the distribution is non-normal, the processing circuit moves to the QQ plot generation step.

11. The system of claim 9, wherein the groups of nucleotide-based sequences each consist of a single nucleotide polymorphism DNA sequence.

12. The system of claim 9, wherein the random error of the model is defined as the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

13. A method for determining whether a given source contributes to the formation of a given product, the method comprising: obtaining: (i) a first nucleotide-based sequence derived from the given source, (ii) a first collection of nucleotide-based sequences, each of which is derived from a source among a plurality of sources that are part of a group of individuals, not all of which are sources, (iii) a second collection of nucleotide-based sequences, each of which is derived from the given product, and (iv) a second nucleotide-based sequence derived from a source known not to be included in the given product that is part of the group of individuals; calculating a first distance associated with the given source, wherein the first distance is comprised of: (i) a distance of the first nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the first nucleotide-based sequence from the second set of nucleotide-based sequences; calculating a second distance associated with a source not included in the plurality of sources, wherein the second distance is comprised of: (i) a distance of the second nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the second nucleotide-based sequence from the second set of nucleotide-based sequences; and, When a difference between the first distance and the second distance is greater than a threshold, it is determined that the given source contributed to forming the product. The method according to claim 13 , wherein after calculating the first distance and the second distance.

15. The method of claim 13, wherein the first set of nucleotide-based sequences and the second set of nucleotide-based sequences each consist of single nucleotide polymorphism DNA sequences.

16. A method for detecting non-compliance in a given product produced by a source group, the method comprising: obtaining: (i) a set of nucleotide-based sequences from each given source in said set of sources, wherein the set of nucleotide-based sequences comprises the set of nucleotide-based sequences common to all sources in said set of sources, and (ii) said set of nucleotide-based sequences from each given source in said set of sources for said given product; For each given nucleotide-based sequence in the set of nucleotide-based sequences, the following operations are performed: (a) generating a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source of the set of sources with the allele frequency of the given nucleotide-based sequence in each given source of the given product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; and, (c) When the residual exceeds a threshold, determining that non-compliance exists in the given product.

17. The method of claim 16, wherein the random error of the model is defined as the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

18. The method of claim 16, wherein the groups of nucleotide-based sequences each consist of a single nucleotide polymorphism DNA sequence.

19. The method of claim 16, wherein non-compliance is also determined to exist when the residuals of all nucleotide-based sequences in the group of nucleotide-based sequences are non-normally distributed.

20. The method of claim 16, wherein: (i) the threshold is defined according to a tolerance threshold, and (ii) the tolerance threshold is the percentage of sources in the set of sources that lack at least one nucleotide-based sequence in the set of nucleotide-based sequences.

21. A method for detecting whether a product produced by a source group contains material from a disease source, the method comprising: obtaining: (i) a set of nucleotide-based sequences from each given source in said group of sources, wherein the set of nucleotide-based sequences comprises the set of nucleotide-based sequences common to all sources in said group of sources, and (ii) a set of nucleotide-based sequences from each given source in said group of sources of a given dairy product; For each given nucleotide-based sequence in the set of nucleotide-based sequences, the following operations are performed: (a) generating a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source of the source group and the allele frequency of the given nucleotide-based sequence in each given source of the given dairy product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; generating a QQ plot based on the residuals of the nucleotide-based sequences in the set of nucleotide-based sequences; as well as, When at least one side of the QQ plot is nonlinear, it is determined that the product contains substances from a disease source.

22. The method of claim 21, wherein before generating the QQ-plot, the processing circuit is configured to determine whether the residuals are normally distributed, and when the distribution is non-normal, the processing circuit moves to the QQ-plot generation step.

23. The method of claim 21, wherein the groups of nucleotide-based sequences each consist of a single nucleotide polymorphism DNA sequence.

24. The method of claim 21, wherein the random error of the model is defined as the minimum sum of squares of the differences between the allele frequencies of the given nucleotide-based sequence in each given source of the source group and the allele frequencies of the given nucleotide-based sequence in each given source of the given finished product.

25. A non-transitory computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by at least one processor to implement a method for determining whether a given source contributes to the formation of a given product, the method comprising: obtaining: (i) a first nucleotide-based sequence derived from the given source, (ii) a first collection of nucleotide-based sequences, each of which is derived from a source among a plurality of sources that are part of a group of individuals, not all of which are sources, (iii) a second collection of nucleotide-based sequences, each of which is derived from the given product, and (iv) a second nucleotide-based sequence derived from a source known not to be included in the given product that is part of the group of individuals; calculating a first distance associated with the given source, wherein the first distance is comprised of: (i) a distance of the first nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the first nucleotide-based sequence from the second set of nucleotide-based sequences; calculating a second distance associated with a source not included in the plurality of sources, wherein the second distance is comprised of: (i) a distance of the second nucleotide-based sequence from the first set of nucleotide-based sequences, and (ii) a distance of the second nucleotide-based sequence from the second set of nucleotide-based sequences; and, When a difference between the first distance and the second distance is greater than a threshold, it is determined that the given source contributed to forming the product.

26. A non-transitory computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by at least one processor to implement a method for detecting non-conformance in a given product produced by a source group, the method comprising: obtaining: (i) a set of nucleotide-based sequences from each given source in said set of sources, wherein the set of nucleotide-based sequences comprises the set of nucleotide-based sequences common to all sources in said set of sources, and (ii) said set of nucleotide-based sequences from each given source in said set of sources for said given product; For each given nucleotide-based sequence in the set of nucleotide-based sequences, the following operations are performed: (a) generating a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source of the set of sources with the allele frequency of the given nucleotide-based sequence in each given source of the given product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; and, (c) When the residual exceeds a threshold, determining that non-compliance exists in the given product.

27. A non-transitory computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by at least one processor to implement a method for detecting whether a product produced by a source group contains material from a disease source, the method comprising: obtaining: (i) a set of nucleotide-based sequences from each given source in said group of sources, wherein the set of nucleotide-based sequences comprises the set of nucleotide-based sequences common to all sources in said group of sources, and (ii) a set of nucleotide-based sequences from each given source in said group of sources of a given dairy product; For each given nucleotide-based sequence in the set of nucleotide-based sequences, the following operations are performed: (a) generating a linear regression model based on a comparison of the allele frequency of the given nucleotide-based sequence in each given source of the source group and the allele frequency of the given nucleotide-based sequence in each given source of the given dairy product; (b) obtaining residuals from the linear regression model, wherein the residuals are random errors of the model; generating a QQ plot based on the residuals of the nucleotide-based sequences in the set of nucleotide-based sequences; as well as, When at least one side of the QQ plot is nonlinear, it is determined that the product contains substances from a disease source.