System for and method of defect detection

The system automates defect detection in granular food materials using machine learning models to analyze collected spectra, reducing manual labor and efficiently identifying both visual and non-visual defects.

WO2026106548A1PCT designated stage Publication Date: 2026-05-21PROFILEPRINT PTE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PROFILEPRINT PTE LTD
Filing Date
2025-10-27
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Conventional defect detection in granular food materials, such as coffee beans or rice, requires significant manual labor and is inefficient for detecting non-visual defects, especially for large batches, leading to exponential effort requirements.

Method used

A system and method utilizing machine learning models, including supervised and unsupervised classifiers, to analyze spectra collected from granular food samples using a spectroscope, combining selected spectra into aggregates, and providing defect predictions based on these aggregates.

Benefits of technology

Automates defect detection, reducing manual labor and efficiently identifying both visual and non-visual defects in granular food materials, improving throughput and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050697_21052026_PF_FP_ABST
    Figure SG2025050697_21052026_PF_FP_ABST
Patent Text Reader

Abstract

A system and a method of defect detection. The method includes: receiving a plurality of first spectra associated with a first food sample disposed in a sensing space, the plurality of first spectra being measured by moving the first food sample relative to a spectroscope, wherein each of the plurality of first spectra corresponds to a portion of the first food sample; obtaining a plurality of first aggregates by combining selected ones of the plurality of first spectra, each of the plurality of first aggregates comprising multiple ones of the plurality of first spectra, wherein each of the plurality of first aggregates corresponds to a respective one of a plurality of first spatial zones of the sensing space; providing the plurality of first aggregates to a machine learning model; and obtaining a defect prediction characteristic of the first food sample based on an output of the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM FOR AND METHOD OF DEFECT DETECTIONCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of priority to the Singapore application no.10202403569T filed 15 November, 2024, the contents of which are hereby incorporated by reference in their entirety for all purposes.TECHNICAL FIELD

[0002] This application relates generally to the field of defect detection, and more particularly, to a system for defect detection and a method of defect detection.BACKGROUND

[0003] Conventionally, defect detection for food material requires a high amount of laborious work, such as manual screening. This is especially so for detecting defects in granular food materials, such as coffee beans or rice. While screening the full batch of granular food materials presents better reliability in defect prediction, the effort needed for screening increases exponentially.SUMMARY

[0004] According to an aspect, disclosed herein a system. The system comprises: memory storing instructions; and a processor coupled to the memory and configured to process the stored instructions to implement: a module configured to perform a method of defect detection. The method comprises: receiving a plurality of first spectra associated with a first food sample disposed in a sensing space, the plurality of first spectra being measured by moving the first food sample relative to a spectroscope, wherein each of the plurality of first spectra correspondsto a portion of the first food sample; obtaining a plurality of first aggregates by combining selected ones of the plurality of first spectra, each of the plurality of first aggregates comprising multiple ones of the plurality of first spectra, wherein each of the plurality of first aggregates corresponds to a respective one of a plurality of first spatial zones of the sensing space; providing the plurality of first aggregates to a machine learning model; and obtaining a defect prediction characteristic of the first food sample based on an output of the machine learning model.

[0005] According to another aspect, the method comprises: receiving a plurality of first aggregates obtained by combining selected ones of a plurality of first spectra, each of the plurality of first aggregates comprising multiple ones of the plurality of first spectra, wherein the plurality of first spectra is associated with a first food sample disposed in a sensing space, the plurality of first spectra being measured by moving the first food sample relative to a spectroscope, wherein each of the plurality of first spectra corresponds to a portion of the first food sample, wherein each of the plurality of first aggregates corresponds to a respective one of a plurality of first spatial zones of the sensing space; inputting the plurality of first aggregates to a machine learning model; and providing a defect prediction characteristic of the first food sample based on an output of the machine learning model.BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various embodiments of the present disclosure are described below with reference to the following drawings:FIG. l is a schematic diagram of a system for defect detection according to embodiments of the present disclosure.FIG. 2 is a schematic top view of the system of FIG. 1.FIG. 3 is a schematic top view of another system according to embodiments of the present disclosure.FIG. 4 is a schematic diagram of another system for defect detection according to embodiments of the present disclosure.FIG. 5 is a schematic diagram showing various modules of a system according to embodiments of the present disclosure.FIG. 6 is a data flow chart of the system of FIG. 5.FIG. 7 is a schematic diagram illustrating a plurality of first spectra according to embodiments of the present disclosure.FIG. 8 is a schematic diagram illustrating a plurality of first aggregates according to embodiments of the present disclosure.FIGs. 9 and 10 illustrate various configurations of a plurality of first aggregates according to embodiments of the present disclosure.FIG. 11 is a schematic diagram illustrating an alignment method using a plurality of first spectra according to embodiments of the present disclosure.FIGs. 12A and 12B are schematic diagrams illustrating an alignment method using a plurality of first aggregates according to embodiments of the present disclosure.FIG. 13 is a schematic diagram showing various modules of a system according to embodiments of the present disclosure.FIG. 14 is a schematic diagram illustrating a local site and a remote site according to embodiments of the present disclosure.FIG. 15 is a flowchart of a method of defect detection according to embodiments of the present disclosure.FIG. 16 is a schematic diagram illustrating a local site and a remote site according to other embodiments of the present disclosure.FIG. 17 is a flowchart of a method of defect detection according to embodiments of the present disclosure.FIG. 18 is a flow chart illustrating a method of defect detection using an unsupervised machine learning model according to embodiments of the present disclosure.FIG. 19 is a flow chart illustrating a method of defect detection using a supervised machine learning model according to embodiments of the present disclosure.FIG. 20 is a flow chart illustrating a method of defect detection using a machine learning model for lot-level predictions according to embodiments of the present disclosure.FIG. 21 is a flow chart illustrating a method of defect detection using a Multiple Instance Learning (MIL) problem-based model according to embodiments of the present disclosure. FIGs. 22 and 23 shows the results from a non-visual defect prediction model built on 58 samples.FIG. 24 is a schematic diagram of a processor system.DETAILED DESCRIPTION

[0007] The following detailed description is made with reference to the accompanying drawings, showing details and embodiments of the present disclosure for the purposes of illustration. Features that are described in the context of an embodiment may correspondingly be applicable to the same or similar features in the other embodiments, even if not explicitly described in these other embodiments. Additions and / or combinations and / or alternatives as described for a feature in the context of an embodiment may correspondingly be applicable to the same or similar feature in the other embodiments.

[0008] In the context of various embodiments, the articles “a”, “an” and “the” as used with regard to a feature or element include a reference to one or more of the features or elements.

[0009] In the context of various embodiments, the term “about” or “approximately” as applied to a numeric value encompasses the exact value and a reasonable variance as generally understood in the relevant technical field, e g., within 10% of the specified value.

[0010] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.

[0011] As used herein, terms “concurrently”, “simultaneously”, “at the same time”, or the like, may refer to events or actions that coincide or overlap within a period of time, regardless of whether the events start at the same time instant, and regardless of whether the events end at the same time instant.

[0012] As used herein, the term “spectra” or “spectrum” may be used interchangeably with the terms “Near Infrared (NIR) spectra or spectrum”, “NIR reflectance spectra or spectrum”, “reflected spectra or spectrum”, “diffuse reflectance spectra or spectrum” etc., and may generally refer to a spectra reflected from a sample in a sensing space which is measured or sensed by a spectroscope, or a spectra measured using diffuse reflectance spectroscopy as it may be understood. In avoidance of doubt, the reflective spectra may comprise NIR light (with corresponding spectra) reflected directly by the food sample and received by the spectroscope. In addition, the reflective spectra may also comprise NIR light scattered and / or reflected indirectly (e g. via secondary reflections) from the food sample and received by the spectroscope.

[0013] As used herein, the term “aggregate” may generally refer to a set of or multiple spectra corresponding to a spatial zone or spatial portion in a sensing space. As such, the set of spectra may be measured or collected from part of or the whole of the sample in the respective spatial zone / spatial portion in the sensing space. As a non-limiting example, one aggregate may comprise a combination of multiple spectra measured from three coffee beans disposed in orpositioned in one quadrant in the sampling sample. As such, the one aggregate may correspond to the one quadrant.

[0014] In addressing the various limitations presented by current methods, the present disclosure presents a proposed system and a method of defect detection incorporating the use of machine learning model(s) for defect prediction. The proposed system may include one or more machine learning model, such as: a supervised classifier, an unsupervised classifier, a multiple-instance learning (MIL) model, a learning from label proportions (LLP) model, a convolutional neural network-based model or a machine learning model based on a convolutional neural network architecture. In some embodiments, the machine learning model may comprise an ensemble learning-based model. The defect prediction may be characteristic of the food sample(s), and may include one or more of: a defect identification prediction; a defect rate prediction; a prediction of a probability of defect, and a defect class prediction. The defect prediction may be a sample-level prediction or a lot-level (multiple samples) prediction.

[0015] The proposed system and method of defect detection may be used on granular food materials, such as dried crop produce from the Fabaceae and Poaceae families, alleviating the laborious effort required for conventional defect screening processes. The granular food materials may include but is not limited to raw coffee beans, soybeans, mung beans, peas, rice, wheat, barley, cocoa beans, peanuts, cashew nuts, hazelnuts, bambara groundnuts, etc. As an exemplary non-limiting diametrical range, each grain of the granular food material may have a size or a dimension between 3mm to 18mm. For example, the granular material may have a maximum dimension of 18mm. Another granul r material may have a minimum dimension of 3 mm.

[0016] In various embodiments, the proposed system and method of defect detection may be used to detect non-visually observable defects corresponding to defects not observable by the naked eye of a screening user, such as internal fungal / mould growth, overfermentation, aged.Such non-visual defects may affect a grain wholly or partially. In the example of raw coffee beans, such non-visual defects may include descriptive terms such as riado (rioy), rio light (rio suave), rio strong (rio iodoforme), rio, phenolic (phenol defect), potato defect, moldy, aged, sour (fermentada), earthy, dirty, stink, quakers, etc., which are descriptors of non-visual defects.

[0017] Additionally or alternatively, the proposed system and method of defect detection may be used to detect visually observable defects such as defects observable by the naked eye of a screening user, such as insect damage through biting or boring, physical damage such as cracking broken edges, fungus or mould infestation, blackened, discolouration, etc. Such visual defects may also affect a grain wholly or partially.

[0018] Referring to FIG. 1, in one aspect, the proposed system 100 for defect detection is presented according to various embodiments of the present disclosure. The system 100 may comprise a processing system 900 configured for performing a method of defect detection for a first food sample 80. The first food sample 80 may be a plurality of granular food sample. The system 100 may comprise a spectroscope 200 in signal communication with the processing system 900. The spectroscope 200 may be configured for measuring a plurality of spectra reflected (directly or indirectly) from the first food sample 80 disposed or contained in a sensing space 400.

[0019] In addition, the system 100 may optionally comprise a light source 300 for illuminating the first food sample 80 during defect detection. The light source 300 may be integrated with the spectroscope 200. The spectroscope 200 and optionally the light source 300 may be configured for diffuse reflectance spectroscopy. In an example, an operable wavelength range used by the spectroscope 200 may be between 300nm (nanometres) to 2500nm. Preferably, the operable wavelength range may be between 400nm to HOOnm. Similarly, an illumination wavelength range provided by the light source 300 may be between 300nm to 3500nm. Preferably, the illumination wavelength range may be between 400nm to 1 lOOnm.

[0020] In various embodiments, the spectroscope 200 and / or the light source 300 may be oriented towards the first food sample 80 disposed or contained in the sensing space 400. In an exemplary embodiment, the first food sample 80 may be held in a specimen dish 410 or a crucible defining the sensing space. The proposed system 100 may provide a defect prediction characteristic of the first food sample 80. In a non-limiting example, during defect detection, the specimen dish 410 may hold or contain 5g to 20g of granular food material.

[0021] In some implementations, the system 100 and the processing system 900 may be located in a computing system (such as a cloud server) remote from the spectroscope 200 and / or the light source 300. In such scenarios, the processing system 900 may access the spectroscope 200 via a network connection 205. In other embodiments, the processing system 900 may be local to the spectroscope 200 and / or the light source 300, and may be configured as a mobile device or a personal computer.

[0022] Still referring to FIG. 1, in various embodiments, the spectroscope 200 (or a sensing module 210 of the spectroscope 200) may be movable or displaceable relative to the sensing space 400. As such, the spectroscope 200 and / or the sensing module 210 may be movable relative to the first food sample 80 during defect detection. In various embodiments, the spectroscope 200 may be moved relative to the first food sample 80 in one or more rotations about a continuous closed path 92.

[0023] Referring to FIG. 2, according to some embodiments of the disclosure, the spectroscope 200 may be held stationary with the sensing space 400 (or the specimen dish 410) moving in one or more rotations about the continuous closed path 92. The specimen dish 410 may move relative to the sensing module 210 in a circular manner. For example, the specimen dish 410 moves relative to a static incident beam of the sensing module 210 in a rotating / circular motion. As such, the incident beam traces a circular path around the specimen dish 410, i.e. the incident beam may move relative to a static specimen dish along a circular path. The incidentbeam may trace a circle or an annulus (ring) when moving relative to the specimen dish 410 to collect the spectra.

[0024] In an example, the specimen dish 410 may rotate at a speed of 4 to 12 revolutions per minute. In an example, the spectroscope 200 may measure or collect a minimum of 2000 distinct spectra within one to three rotations about the continuous closed path 92.

[0025] In various embodiments, each rotation / revolution about the continuous closed path 92 may have a different starting point and ending point, i.e. the revolutions are offset radially relative to each other. Similarly, different measuring sessions may have a different starting point and ending point.

[0026] In other embodiments, the sensing space 400 (or the specimen dish 410) may be held stationary with the spectroscope 200 moving in a circular continuous closed path 92.

[0027] In various embodiments, referring to FIG. 3, both the spectroscope 200 and the sensing space 400 (or the specimen dish 410) may be collectively and concurrently displaced to move the spectroscope 200 relative to the specimen dish 410 in a continuous non-circular closed path 94. For example, both the incident beam of the sensing module 210 and the specimen dish 410 may be in motion to move the incident beam relative to the specimen dish in a non-circular manner.

[0028] During operation, the illumination provided by the light source 300 may vary (or drift). This may be due to prolonged use or the age of the light source 300, such that the illumination provided is less uniform and may affect the measurement of spectra. In various embodiments, a relative speed of moving the first food sample 80 relative to the spectroscope 200 may be varied. For example, the speed of displacement (or rotation) between the specimen dish 410 relative to the sensing module 210 (or incident beam) may be varied. This helps with “correcting” the non-uniform illumination provided by the light source 300, thus providing a more robust measurement.

[0029] Further referring to FIG. 4, in various embodiments, the system 100 may be provided with a mixer 420 (or a stirrer) rotatable about a longitudinal axis 425. The mixer 420 may be positioned or disposed in the sensing space 400 for moving the first food sample 80 relative to the spectroscope 200. The provision of the mixer 420 allows relative movement or displacement between the spectroscope 200 (or a sensing module 210 of the spectroscope 200) relative to the first food sample 80 in the sensing space 400. Tn addition to the embodiments as described previously in which the first food sample 80 moves in two dimensions (radially and tangentially), the provision of the mixer 420 enables relative movement axially along the longitudinal axis 425 between the first food sample 80 and the spectroscope 200. It may be said that the mixer 420 may move the first food sample 80 relative to the spectroscope 200 in three orthogonal directions This enables more granular food samples 80 to be screened or measured at one time, improving defect detection throughput of the system 100.

[0030] In an exemplary embodiment, 50g to 500g of the granular food material may be placed in a cylindrical receptacle, of 5-15cm in diameter and 5-20cm in height. The receptacle may be provided with a mixer (or rotating device) mounted within the receptacle. The mixer enables mixing of the granular food material whilst the incident beam collects or measures the spectra. In an example, the mixer may be in the form of an auger screw, tapered or otherwise. The mixer may be in the fonn of a central shaft, with fins mounted to it that may be vertical or angled 0-90° to the shaft. The mixer may rotate at a speed of 0-100rpm. The mixer may increase the chance of the granular food material in the receptacle being measured by / exposed to the incident beam, as the mixer moves the granular material in an up-down and angular motion.

[0031] FIGs. 5 and 6 illustrate the various functional modules of the processing system 900 for defect detection. The processing system 900 may comprise a measurement module 510 for measuring a plurality of first spectra 512 associated with the first food sample 80 disposed in the sensing space 400. Further referring to FIG. 7, the plurality of first spectra 512 may bemeasured by moving the first food sample 80 relative to the spectroscope 200 as described earlier. The plurality of first spectra 512 may be distinct spectra. As the plurality of first spectra 512 are measured during movement, the plurality of first spectra 512 may correspond to different portions of the first food sample 80, for example, different coffee beans or parts of a coffee bean. Hence, each of the plurality of first spectra 512 may correspond to a portion of the first food sample 80. Tn some implementations, the plurality of first spectra 512 may comprise at least 2000 distinct spectra. In various embodiments, the plurality of first spectra 512 may be measured during a single rotation about the continuous closed path 92 / 94.

[0032] Referring to FIGs. 5 to 8, in various embodiments, the processing system 900 may further comprise an aggregation module 520 for obtaining a plurality of first aggregates 522 by combining selected ones of the plurality of first spectra 512. It may also be said that the plurality of first spectra 512 are collapsed or combined into a plurality of first aggregates 522. It may be appreciated that for some embodiments, only selected ones of the plurality of first spectra 512 are used to form the plurality of first aggregates 522. The selected ones of the plurality of first spectra 512 may be distinct spectra. However, for other embodiments, all of the plurality of first spectra 512 are used to form the plurality of first aggregates 522.

[0033] Referring to FIG. 7, each of the plurality of first aggregates 522 may comprise multiple ones of the plurality of first spectra 512 In an example, each of the plurality of first aggregates 522 may include 20 to 80 first spectra. In addition, referring to FIG. 8, each of the plurality of first aggregates 522 may correspond to a respective one of a plurality of first spatial zones 430 of the sensing space 400. In an example, each of the plurality of first aggregates 522 may be understood to correspond to a slice or sector of a circular space or an annulus space. The plurality of first aggregates 522 may be spatially divided, partitioned or sectioned.

[0034] In various embodiments, the aggregation module 520 may use a clustering model for obtaining the plurality of first aggregates 522 by combining selected ones of the plurality offirst spectra 512. In some examples, the clustering model may be an unsupervised clustering algorithms to determine each of the plurality of first aggregates 522 selected from the plurality of first spectra 512. In other words, the selection of spectra for aggregation may be performed using clustering methods, which may include, but are not limited to, algorithms such as spectral clustering, k-nearest neighbours (k-NN), HDBSCAN and DBSCAN.

[0035] In various embodiments, the plurality of first aggregates 522 may correspond to a plurality of non-overlapping spatial zone 430 of the sensing space 400. In various embodiments, each of first spatial zones 430 may be regarded as a datapoint for further processes. In some embodiments, aggregation is done by averaging of points across the whole wavelength range considered.

[0036] In an exemplary embodiment as shown in FIG. 8, the sensing space 400 may be divided into 16 non-overlapping first spatial zones 430. One of the first aggregate 522A may correspond to one of the first spatial zone 430A. In other words, first aggregate 522A may include spectra 512 measured or obtained from the first spatial zone 430A. Similarly, another of the first aggregate 522B may correspond to another of the first spatial zone 430A. In various embodiments, as shown in FIG. 8, the plurality of first aggregates 522 may correspond to a plurality of uniformly divided spatial zones 430 of the sensing space 400. It may be appreciated that the exemplary embodiment as shown in FIG. 8 is non-limiting in nature.

[0037] Referring again to FIGs. 5 and 6, in various embodiments, the processing system 900 may further comprise a machine learning model 530 or a machine learning module. The plurality of first aggregates 522 may be provided to the machine learning model 530 as an input. In addition, the processing system 900 may further comprise a prediction module 540 for obtaining a defect prediction 542 characteristic of the first food sample 80 based on an output 532 of the machine learning model 530. The defect prediction 542 may be characteristic of the food sample(s). The defect prediction 542 may correspond only to non-visually observabledefect(s). In an example, the defect prediction 542 may correspond to at least one non-visually observable defect. The defect prediction 542 may include one or more of: a defect identification prediction; a defect rate prediction; a prediction of a probability of defect; and a defect class prediction. For example, the output 532 of the machine learning model 530 may correspond to a “defective” prediction outcome or a “defect-free” prediction outcome. In another example, the output 532 of the machine learning model 530 may be a value indicative of a defect rate prediction. In yet another example, the output 532 of the machine learning model 530 may be a value indicative of a defect class or defect type prediction. In non-limiting examples, the defect prediction 542 may be obtained by further processing of one or more outputs 532 of the machine learning model 530. In some instances, the output 532 may also be referred to as “Model Output Features” or “Intermediate Model Output”.

[0038] In various embodiments, the processing system 900 may further comprise a preprocessing module 550. The pre-processing module 550 may receive the plurality of first aggregates 522 from the aggregation module 520 and perform at least one pre-processing step on the plurality of first spectra 522, prior to providing plurality of first spectra 522 to the aggregation module 520. In other words, at least one pre-processing step on the plurality of first spectra 522 may be performed prior to obtaining the plurality of first aggregates 522. In various embodiments, the at least one pre-processing step comprises one or more of: a clustering step; a truncation step; a normalization step; a noise reduction step; and a dimensional reduction step. In an example, the spectra may or may not be further truncated to an appropriate section, for example 700nm-l lOOnm. The plurality of first spectra 522 may be pre-processed through one or more preprocessing techniques, such as: Standard Normal Variate, Multiplicative Scatter Correction, Savitsky Golay filter, 1st derivative, 2nd derivative, etc. In some embodiments, the spectra are further processed through Principal Components Analysis, UMAP or other dimensionality reduction techniques.

[0039] In various embodiments, as shown in FIG. 9, the plurality of first aggregates 522 may correspond to a plurality of non-uniformly divided spatial zones 430 of the sensing space 400. For example, spatial zone 430A may include 5 spatial units, spatial zone 430B may include 3 spatial units, spatial zone 430C may include 4 spatial units, spatial zone 430D may include 2 spatial units, spatial zone 430E and 430F may each include 1 spatial unit. Each of the spatial zones 430A to 430F may correspond to a respective first aggregate 522.

[0040] In other embodiments, as shown in FIG. 10, the plurality of first aggregates 522 may correspond to a pair of adjacent first spatial zones 430A and 430B which are spaced apart in the sensing space 400. In various non-limit examples, the plurality of first aggregates 522 may be spatially divided, partitioned or sectioned, such that each slice is of a substantially equal spatial size, e.g. area, volume, dimension. In various non-limit examples, the plurality of first aggregates 522 may be non-uniformly divided / partitioned / sectioned such that each slice is of a substantially different spatial size, e.g. area, volume, dimension. In various non-limit examples, the plurality of first aggregates 522 may spatially overlap with one other.

[0041] In various embodiments, the plurality of first spectra 512 may undergo an alignment method prior to the determination of the plurality of first aggregates 522. This enables a more informed choice of hyperparameters for unsupervised clustering algorithms to determine each of the plurality of first aggregates 522.

[0042] In various embodiments, the plurality of first spectra 512 may be measured by moving the first food sample 80 relative to the spectroscope 200 in a plurality of rotations about the continuous closed path 92 / 94. Referring to FIG. 11, the alignment method may comprise dividing the plurality of first spectra 512 into a plurality of first spectrum sections 513. Each of the plurality of first spectrum sections 513 may comprise multiple ones of the plurality of first spectra 512. The alignment method may further comprise: based on a similarity measure between ones of the plurality of first spectrum sections 513, determining from the plurality offirst spectrum sections: a starting spectrum 5121 and an ending spectrum 5122. The similarity measure may be a distance metric between pairs of the plurality of first spectrum sections 513 at the same positional index. The distance metrics may include, but are not limited to, Cosine Distance, Euclidean Distance, Manhattan distance, etc.

[0043] The starting spectrum 5121 to the ending spectrum 5122 may correspond to a single one of the plurality of rotations about the continuous closed path 92 / 94 Tn other words, the starting spectrum 5121 may correspond to a first spectrum measured from the single rotation, and the ending spectrum 5122 may correspond to a last spectrum measured from the single rotation.

[0044] In various embodiments, based on the starting spectrum 5121 and the ending spectrum 5122, selected ones (or all) of the plurality of first spectra 512 from the starting spectrum 5121 to the ending spectrum 5122 are combined into the plurality of first aggregates 522 and provided as the plurality of first aggregates 522 to the machine learning model 530.

[0045] In various embodiments, the plurality of first aggregates 522 may undergo an alignment method for the affirmation of spatial overlap between the plurality of first aggregates 522. This enables a more informed choice of hyperparameters for unsupervised clustering algorithms to determine each of the plurality of first aggregates 522. Similarly, the plurality of first spectra 512 may be measured by moving the first food sample 80 relative to the spectroscope 200 in a plurality of rotations about the continuous closed path 92 / 94.

[0046] Referring to FIG. 12A, in various embodiments, the alignment method may comprise: dividing the plurality of first aggregates 522 into a plurality of first aggregate sections 523. Each of the plurality of first aggregate sections 523 may comprise multiple ones of the plurality of first aggregates 522. In some embodiments, each of the first aggregate sections 523 may comprise multiples ones of the plurality of first spectra 512. The length of each of the plurality of first aggregate sections 523 may be varied. In other embodiments, the number of firstaggregates 522 (or length) of each of the plurality of first aggregate sections 523 may be varied for the determination of the starting aggregate 5231 and the ending aggregate 5232.

[0047] In alternative embodiments, each of the first aggregate sections 523 may comprise partial ones of the plurality of first aggregates 522.

[0048] The alignment method may further comprise: based on a similarity measure between ones of (such as a pair of) the plurality of first aggregate sections 523, determining from the plurality of first aggregates 522: a starting aggregate 5231 and an ending aggregate 5232. The similarity measure may be a distance metric between pairs of the plurality of first aggregates sections 523 at the same positional index. The distance metrics may include, but are not limited to, Cosine Distance, Euclidean Distance, Manhattan distance, etc.The starting aggregate to ending aggregate may correspond to a single one of the plurality of rotations about the continuous closed path 92 / 94.

[0049] In various embodiments, based on the starting aggregate 5231 and the ending aggregate 5232, selected ones (or all) of the plurality of first aggregates 522 from the starting aggregate 5231 and the ending aggregate 5232 are provided to the machine learning model 530.

[0050] Further referring to FIG. 12B, for illustration purposes, the alignment of aggregate sections 523 is shown with the tracing of aggregates represented as a repeating linear sequence of 5 aggregates. As examples, the partitioned aggregate sections 523 of different lengths: 3, 4, 5 are shown to illustrate the method of alignment with the variation of the length of first aggregate sections 523. In the example, the different lengths (3, 4 or 5) of the partitioned aggregate sections are initial guesses. Taking the example of length 3, the partitioned aggregate sections will be [ABC] [DEA] [BCD] [E], As such, comparison between the first index of each of the partitioned aggregate section of length 3, i.e. [A—] [D— ] [B— ][E-], results in a higher distance value. Therefore, length 3 is not selected as the suitable length.

[0051] In an exemplary embodiments, assuming that the tracing covers at least two revolutions where the plurality of first aggregates 522 spatially overlap, the plurality of first aggregates 522 are partitioned into a plurality of first aggregate sections 523 of equal length, following the sequence in which the aggregates were captured. The similarity of two plurality of first aggregate sections 523 are evaluated by taking the mean of a distance metric between pairs of aggregates at the same positional index within each section, such as Cosine Distance, Euclidean Distance, Manhattan distance, etc. If there are insufficient first aggregates 522 to generate a first aggregate section of desired length (i.e. the last section 5233), the distance metric score for that last section 5233 may be evaluated against truncated versions of previous first aggregate section 523.

[0052] In an exemplary embodiment, the alignment method (or search process) may be repeated for sections of various lengths (or the number of first aggregates 522 in each first aggregate section 523) up to a maximum value equal to half of the total number of aggregates or 40, whichever is higher. The length returning lower distances may be suitable as a most appropriate length corresponding to one complete revolution. The length and total number of revolutions determined may be used to obtain hyperparameters for unsupervised clustering algorithms.

[0053] In various embodiments, the spectroscope 200 may measure at a first sampling rate for one of the plurality of rotations, and measure at a second sampling rate for another one of the plurality of rotations, wherein the first sampling rate is different from the second sampling rate.

[0054] In exemplary embodiments, the plurality of first aggregates 522 may be spatially divided, partitioned or sectioned differently for each revolution of spectra collection. As an example, the number of the plurality of first spectra 512 collected in the first revolution may bedivided into 12 sectors, and the number of the plurality of first spectra 512 collected in the second revolution may be divided into 18 sectors.

[0055] Referring to FIG. 13, in various embodiments, the system 100 may comprise a processing system 900 configured for performing a method of defect detection for a plurality of food samples, such as a first food sample 80 and a second food sample 82. In other words, the method of defect detection may be performed on a “lot-level” This is in addition to the method of defect detection previously described, which may be performed on a single first food sample 80, i.e. on a “sample-level”. For example, multiple food samples may be measured or sampled to provide multiple spectra. The first food sample 80 and the second food sample 82 may be a distinct batch of granular food samples. The first food sample 80 and the second food sample may be food samples of the same type, such as coffee beans. The machine learning model(s) may thus be configured to generate defect prediction at the lot-level instead of the sample-level. In various embodiments, lot-level predictions may be aggregated from samplelevel predictions, through methods such as: averaging, or AveragePooling or MaxPooling for CNNs. The machine learning model may also generate lot-level predictions by directly predicting on aggregated characteristics of the lot based on the spectra of the individual samples.

[0056] In various embodiments, the processing system 900 may comprise a measurement module 510 for measuring a plurality of first spectra 512 associated with the first food sample 80 and a plurality of second spectra 514 associated with a second food sample 82. The first food sample 80 and the second food sample 82 may be disposed in the sensing space 400 at different time instances. Each of the plurality of first spectra 512 and the plurality of second spectra 514 may be measured by moving the first food sample 80 / the second food sample 84 relative to a spectroscope. The plurality of first spectra 512 may correspond to a portion of the first foodsample 80. The plurality of second spectra corresponds to a portion of the second food sample 82.

[0057] The processing system 900 may further comprise an aggregation module 520 for obtaining a plurality of first aggregates 522 by combining selected ones of the plurality of first spectra 512, and obtaining a plurality of second aggregates 524 by combining selected ones of the plurality of second spectra 514.

[0058] Each of the plurality of first aggregates 522 may comprise multiple ones of the plurality of first spectra 512. Each of the plurality of second aggregates 524 may comprise multiple ones of the plurality of second spectra 514. As shown in FIG. 8, each of the plurality of first aggregates 522 may correspond to a respective one of a plurality of first spatial zones 430 of the sensing space 400. Similarly, each of the plurality of second aggregates 524 may correspond to a respective one of a plurality of second spatial zones of the sensing space 400.

[0059] In the example shown in FIG. 8, a circle is divided into sectors (or first spatial zones 430) by radial lines, with coffee beans spread across. Each sector represents an aggregate of collapsed spectra, with the total aggregates ranging from 20 to 80.

[0060] The processing system 900 may further comprise a machine learning model 540 or a machine learning module. In other embodiments, the machine learning module may comprise a plurality of machine learning models 530. The plurality of first aggregates 522 and the plurality of second aggregates 524 may be provided to the machine learning model 530 as input. Alternatively, the plurality of first aggregates 522 and the plurality of second aggregates 524 may be provided to the plurality of machine learning models 530 as input.

[0061] The processing system 900 may further comprise a prediction module 540 for obtaining a collective defect prediction 544 characteristic of the first food sample 80 and the second food sample 82 based on an output 532 of the machine learning model 530. In other words, the collective defect prediction 544 may be characterized by a lot-level defect predictionof multiple food samples. In some examples, the collective defect prediction 544 may comprise at least one of: a defect identification prediction; a defect rate prediction; a prediction of a probability of defect; and a defect class prediction.

[0062] Referring to FIG. 14, in various embodiments, the proposed system 100 may be split into a local site 810 corresponding to the measurement(s) of the food sample (s), and a remote site 910 corresponding to the processing of the measurement(s). For example, the local site 810 may correspond to a client site in which the food sample(s) are located, and the remote site 910 may correspond to a service provider site in which the computational processes are performed. In an example, the measurement module may be located at the local site 810. Further, the aggregation module 520, the machine learning model 530 and the prediction module 540 may be located at the remote site 910.

[0063] In various embodiments, referring to FIG. 15, the proposed system 100 may be configured to perform a method 710 of defect detection. The method 710 may comprise: in 712, receiving a plurality of first spectra associated with a first food sample disposed in a sensing space, the plurality of first spectra being measured by moving the first food sample relative to a spectroscope, wherein each of the plurality of first spectra corresponds to a portion of the first food sample; in 714, obtaining a plurality of first aggregates by combining selected ones of the plurality of first spectra, each of the plurality of first aggregates comprising multiple ones of the plurality of first spectra, wherein each of the plurality of first aggregates corresponds to a respective one of a plurality of first spatial zones of the sensing space; in 716, providing the plurality of first aggregates to a machine learning model; and in 718, obtaining a defect prediction characteristic of the first food sample based on an output of the machine learning model.

[0064] In various embodiments, the plurality of first spectra is measured by moving the first food sample relative to the spectroscope in a plurality of rotations about a continuous closedpath. In various embodiments, the method further comprises: dividing the plurality of first spectra into a plurality of first spectrum sections, wherein each of the plurality of first spectrum sections comprises multiple ones of the plurality of first spectra.

[0065] In various embodiments, the method further comprises: based on a similarity measure between ones of the plurality of first spectrum sections, determining from the plurality of first spectrum sections: a starting spectrum and an ending spectrum, the starting spectrum to ending spectrum corresponding to a single one of the plurality of rotations about the continuous closed path. In various embodiments, selected ones of the plurality of first spectra from the starting spectrum to the ending spectrum are combined into the plurality of first aggregates and provided as the plurality of first aggregates to the machine learning model.

[0066] In various embodiments, the method further comprises: dividing the plurality of first aggregates into a plurality of first aggregate sections, wherein each of the plurality of first aggregate sections comprises multiple ones of the plurality of first spectra. In various embodiments, the method further comprises: based on a similarity measure between ones of the plurality of first aggregate sections, determining from the plurality of first aggregates: a starting aggregate and an ending aggregate, the starting aggregate to ending aggregate corresponding to a single one of the plurality of rotations about the continuous closed path. In various embodiments, selected ones of the plurality of first aggregates from the starting aggregate to the ending aggregate are provided to the machine learning model.

[0067] In various embodiments, the method further comprises: using a clustering model, obtaining the plurality of first aggregates by combining selected ones of the plurality of first spectra. In various embodiments, the method further comprises: varying a relative speed of moving the first food sample relative to the spectroscope. In various embodiments, the method further comprises: performing at least one pre-processing step on the plurality of first spectra prior to obtaining the plurality of first aggregates, wherein the at least one pre-processing stepcomprises one or more of: a clustering step; a truncation step; a normalization step; a noise reduction step; and a dimensional reduction step.

[0068] Referring to FIG 16, according to other embodiments of the proposed system 100, the measurement module and the aggregation module 520 may be located at the local site 810. Further, the machine learning model 530 and the prediction module 540 may be located at the remote site 910. As such, the measurement and aggregation processes may be performed on the local site 810 (or client site), and the prediction process may be performed on the remote site 820 (or the service provider site).

[0069] In various embodiments, referring to FIG. 17, the proposed system 100 may be configured to perform a method 720 of defect detection. The method 720 may comprise: in 722, receiving a plurality of first aggregates obtained by combining selected ones of a plurality of first spectra, each of the plurality of first aggregates comprising multiple ones of the plurality of first spectra, wherein the plurality of first spectra is associated with a first food sample disposed in a sensing space, the plurality of first spectra being measured by moving the first food sample relative to a spectroscope, wherein each of the plurality of first spectra corresponds to a portion of the first food sample, wherein each of the plurality of first aggregates corresponds to a respective one of a plurality of first spatial zones of the sensing space; in 724, inputting the plurality of first aggregates to a machine learning model; and in 726, providing a defect prediction characteristic of the first food sample based on an output of the machine learning model.

[0070] Exemplary Machine Learning Models for the proposed system and method

[0071] Unsupervised Machine Learning Model

[0072] FIG. 18 illustrates a flow chart of a method of defect detection using an unsupervised machine learning model for the system 100 according to various embodiments of the disclosure. The unsupervised machine learning model may be built by the following steps: i) Spectra are1acquired from a set of samples generally regarded as defect free and preprocessed as pr described above. Based on an assumption that the majority of aggregated spectra correspond to “non-defect” spectra, the “non-defecf ’ spectra are used to train an anomaly detection model or machine learning model. The minority aggregated spectra are regarded as “defect” spectra and used to aid the choice of algorithms.

[0073] Majority and / or minority of aggregated spectra may be determined by way of clustering or clustering methods. For example, given x number of clusters, the cluster with the largest representation is regarded as “non-defect”. The anomaly detection model may be any of the following: one-class support vector machines, gaussian mixed models, histogram-based outlier score, isolation forest, local outlier factor.

[0074] Tn some embodiments, the above algorithms may also be ensembled by voting or stacking as it is possible for clean samples to contain some defective grains, and hence may correspond to a defective spectrum. A scoring function may be used to determine if the sample in question is predicted as containing defects or otherwise. This aligns the model with a subjective evaluation by humans in which the samples may be labelled. A sample with more sectors comprising defective spectra would score higher. For sample with a number of defective spectra beyond a specified threshold, the sample is considered a ‘defective’ sample.

[0075] Supervised Machine Learning Model

[0076] FIG. 18 illustrates a flow chart of a method of defect detection using a supervised machine learning model for the system 100 according to various embodiments of the disclosure. The supervised machine learning model may be built by the following steps: i) Spectra are acquired from a set of samples generally regarded as defect-free and preprocessed as per described above, ii) Spectra are additionally acquired from a set of samples known to contain specific defects in appreciable prevalence, and pre-processed as described above, iii) For non-defect samples, the aggregated spectra are clustered and the spectra in the majority class are labelled as non-defect or “clean”, iv) For defective samples, the aggregated spectra are clusteredand the spectra in the minority class(es) are labelled according to the claimed defect type as labelled.

[0077] In various embodiments, a supervised classifier may be trained with the above labelled spectra corresponding to non-defect spectra and various types of labelled defect spectra. The supervised classifier algorithms may be any of the following: support vector machines, logistic regression, k-nearest neighbours, linear discriminant analysis, quadratic discriminant analysis, Naive Bayes classifier, decision tree.

[0078] In some embodiments, the above algorithms may be ensembled. Ensembled techniques may include but is not limited to: random forest, extra trees, XGBoost Classifier, Adaboost Classifier, Catboost classifier. A sample with a notable number of sectors comprising defective spectra classified as a specific defect would be classified as belonging to that defect type.

[0079] In various embodiments, for both the unsupervised machine learning model and the supervised machine learning model, the machine learning model may be used to predict on 10 separate samplings of a particular lot sample (lot). The model predicts a number X out of 10, and this is indicative of the prevalence of defects in that lot sample (lot). For example, a sample that results in 8 out of 10 samples predicted as “defective” is regarded to be more defective than another sample with 3 out of 10 samples predicted as “defective”.

[0080] Machine Learning Model for Lot-level prediction

[0081] FIG. 20 illustrates a flow chart of a method of defect detection using a machine learning model for lot-level predictions. The machine learning model may be built by the following steps: i) Spectra are acquired from a set of samples generally regarded as defect-free and pre-processed as described above, ii) Spectra are additionally acquired from a set of samples known to contain specific defects in appreciable prevalence, and preprocessed as per described above, iii) A machine learning model is built to generate predictions at the lot level,where each lot may comprise variable numbers (or equal numbers) of samples. Predictions pertaining to the aggregated characteristics of the lot are predicted based on the spectra from the individual specimens. The representation of the spectra at the lot level is dynamic in two key dimensions: the clustering compression dimension and the number of repeated samplings.

[0082] Clustering compression involves clustering the spectra into n clusters, representing n spectra per specimen, thus allowing for dimensionality reduction while retaining key spectral features. Each lot may comprise a dynamic, non-fixed number of samples / specimens, resulting in a flexible model structure. To handle the position-invariance of the spectra across samples, the machine learning model may adapt neural networks capable of learning representations that are invariant to the order and positioning of the inputs. These networks may include convolutional neural networks (CNNs) for spatial feature extraction, recurrent neural networks (RNN) for handling non-Euclidean data relationships, and transformers for sequence-based learning, ensuring robustness in capturing spectral variations.

[0083] The final model optimises a loss function that may include metrics such as crossentropy loss for classification tasks, mean squared error for regression tasks, or other relevant loss functions, depending on the nature of the target prediction. Tn various embodiments, the model may predict: the probability of having a defect within a lot; the proportion of defects within a lot; and the proportion and type of defects within a lot.

[0084] Multiple Instance Learning (MIL) problem-based Model

[0085] FIG. 21 is a flow chart of a method of defect detection using a Multiple Instance Learning (MIL) problem-based model for the system 100 according to various embodiments of the disclosure. The defect detection method may be formulated as a Multiple Instance Learning (MIL) problem, with the following steps: i) Spectra are acquired from lots, including both defect-free and known defect-containing samples, and preprocessed as described above, ii) Each lot is treated as a “bag” containing multiple “instances” (spectra from different samples or specimens), allowing MIL models to be applied where only bag-level labels (i.e., defect ornon-defect) are known rather than instance-level labels, iii) Various MIL algorithms, such as Multiple Instance Support Vector Machine (MI-SVM), may be used to classify each bag, leveraging the assumption that a defect label applies if at least one instance (sample spectra or specimen spectra) in the bag exhibits defect characteristics.

[0086] In various embodiments, as an extension of the MIL approach, a convolutional neural network (CNN) may be trained to classify individual instances within each sample. For lotlevel predictions, instance predictions are aggregated using techniques such as max pooling or mean pooling across instances in each bag, thus allowing for a single defect score or classification to be representative at the lot level.

[0087] In various embodiments, the aggregated lot-level predictions may be used as the final output, with appropriate loss functions applied during training to optimize for accurate baglevel defect classification. In various embodiments, an ensemble of MIL and CNN-based models is used to improve robustness, with voting or stacking methods applied to combine model outputs.

[0088] Learning from Label Proportions (LLP) problem-based Model

[0089] According to various embodiments of the disclosure, the machine learning model may be a Label Proportions (LLP) problem-based model The LLP problem-based model may allow for the prediction of the types and proportions of defects in a particular sample or a particular lot. The LLP problem-based model may be built, with the following steps: i) Spectra are acquired from lots, including both defect-free and known defect-containing samples, and preprocessed as described above, ii) Each lot is treated as a “bag” containing multiple “instances” (spectra from different samples or specimens), allowing LLP algorithms to be applied where only the proportions of classes in a bag are known rather than relying on instancelevel labels, iii) Various LLP algorithms, such as Deep LLP, may be used to predict the proportion of each defect class in each bag. iv) The backbone of this model may be a convolutional neural network (CNN), a recurrent neural network (RNN) or a transformermodel. Predictions may be performed at the lot level, or at the sample levels for more granular predictions (i.e., defining samples or spectra as nested bags within lots, allowing LLP to be performed at a more granular level).

[0090] In various embodiments, for lot-level predictions, granular predictions are aggregated using techniques such as max pooling or mean pooling across instances in each bag. The aggregated lot-level predictions may be used as the final output, optimized with a relevant proportion loss function such as Kullback-Leibler Divergence or cross-entropy loss.

[0091] In various embodiments, as an extension of this LLP approach, an LLP-GAN methodology may be employed, with the discriminator and generator architectures being the same, or in other cases, different. In various embodiments, as an extension of the LLP-GAN methodology, a CNN may be used as the backbone of both the discriminator and generator architectures. In various embodiments, as an extension of the LLP approach, a contrastive loss is added to the objective to encourage better representation learning in the model.

[0092] In various embodiments, the model may be used to predict on 10 separate samples of a particular lot. The model predicts a number X out of 10 for each defect class, and this is indicative of the prevalence of that class of defects in that lot. The prediction is hence presented as an array of integers that sum to 10, with each index corresponding to one of the defect classes. For example, a sample may result in 3 samplings out of 10 predicted as riado (rioy), 2 samplings out of 10 predicted as mouldy and 5 samplings out of 10 as clean cup (or defect-free).

[0093] Exemplary training sample set

[0094] The following section describes various exemplary training sample sets used in the proposed system and methods. Multiple lots of raw coffee beans were used, wherein each lot comprises multiple samples or specimens, with each sample corresponding to a single sampling of the lot and represented by an acquired and preprocessed spectra.

[0095] The training sample set comprises between 200 and 500 lots. The lots originate from multiple geographic sources including, but not limited to, Brazil, China, Colombia, and Honduras The beans within each lot exhibit size range between approximately 3 mm and 18 mm.

[0096] Each lot in the training sample set is labelled according to its defect content. Labels may be assigned in proportional form, such as 2 / 10 Mofado, 2 / 10 Rio, and 6 / 10 Clean Cup, or may be simplified into a binary classification of defective versus non-defective. In the binary formulation, the training sample set may be curated such that approximately half of the lots are defective and half are non-defective, with defective lots exhibiting defect prevalence between 10% and 100%. Defect classes represented in the training sample set include, but are not limited to, Mofado, Riado, Rio Light, Rio Strong, Fermented, and Others.

[0097] The training sample set may be used for the different machine learning model as described above. For Multiple Instance Learning (MIL), each lot is treated as a bag of specimens and inherits a defect label if at least one specimen exhibits defect characteristics. For Learning from Label Proportions (LLP), each lot is associated with a proportional label specifying the distribution of defects among its specimens. For conventional supervised classification models and for convolutional neural network (CNN) extensions, the same training sample set provides the specimen-level spectra and the lot-level or proportional labels as appropriate.

[0098] Verification of the trained models may be carried out by partitioning the training sample set into training and testing subsets using standard train-test split procedures, ensuring that unseen lots and their associated specimens are reserved for evaluation. Such train-test splits could be split randomly, optionally with stratification, or as a time split, where the most recent samples are used as the test subset. Such split ratios could range from 50:50 to 95:5 for the training and testing subsets respectively.

[0099] Experiments

[0100] Experiments were performed on using an exemplary implementation of the proposed defect detection system 100 and method. FIGs. 22 and 23 shows the results from a non-visual defect prediction model built on 58 samples.

[0101] Table 1 below shows the results from another test set of 34 samples. Table 1 shows the breakdown by number of specimens out of 10 in each sample (lot) predicted as “Defect”.Table 1. Results from test set of 34 samples.

[0102] Table 2 below shows an exemplary output of the LLP -based method with Mean Average Error, maximum error and standard deviation of errors of the number of specimens out of 10 in sample (lot). “Clean Cup MAE” suggests that for a sample (lot), the actual number of specimens out of 10 that are manually evaluated as “Clean Cup” could differ on average 2.1224 specimens from the predicted number of specimens.Table 2. LLP -based method

[0103] On another set of blind samples, the coffee was scanned and simulated through 200 repetitions. For the medium defect sample, the model predicts between 4-10 cups, with the mode at 6 cups, whereas for the high defect sample, the model predicted 8-10 cups with the mode at 9 cups out of 10.

[0104] The proposed system and method may be implemented by a processor system 900 as illustrated in the schematic block diagram of FIG. 24. Components of the processing system 900 may be provided within one or more computing device to carry out the functions of the modules or any other modules. One skilled in the art will recognize that the exact configuration or arrangement illustrated in FIG. 24 is provided by way of example only, e.g., each processing system provided may be different and the exact configuration of processing system 900 may vary.

[0105] In embodiments of the present disclosure, the processing system 900 may include a controller 901 and user interface 902. User interface 902 is configured to enable manual interactions between a user and the computing module as required. For this purpose, the processing system 900 includes the input / output components required for the user to enter instructions to provide updates to each of the modules. A person skilled in the art will recognize that components of user interface 902 may vary from embodiment to embodiment but may typically include one or more input devices 935 such as but not limited to a touchscreen, akeyboard, a joystick, a mouse, a microphone, etc. The user interface 902 can also include a media player 940, which can be in the form of one or more playback devices, including but not limited to a display, a speaker, earphones, headsets, etc.

[0106] The controller 901 is configured to be in data communication with the user interface 902 via bus 915. The controller 901 includes memory 920 and processor 905 mounted on a circuit board to process instructions and data, e.g., to perform the method of the present disclosure. The controller 901 includes an operating system 906, an input / output (I / O) interface 930 for communicating with user interface 902, and a communications interface, e.g., a network card 950. The network card 950 may, for example, be configured to send data from the controller 901 via a wired or wireless network to other processing devices or to receive data via the wired or wireless network. Wireless networks that may be utilized by the network card 950 include, but are not limited to, Wireless-Fidelity (Wi-Fi), Bluetooth, Near Field Communication (NFC), cellular networks, satellite networks, telecommunication networks, Wide Area Networks (WAN), and etc.

[0107] Memory 920 and operating system 906 are in data communication with central processing unit (CPU) 905 via bus 910 The memory 920 may include both volatile and nonvolatile memory. The memory 920 may include more than one of each type of memory, e.g., Random Access Memory (RAM) 923, Read Only Memory (ROM) 925, and a mass storage device 945. The mass storage device 945 may include one or more solid-state drives (SSDs). One skilled in the art will recognize that the memory described above includes non-transitory computer-readable media and shall be taken to include all computer-readable media except for a transitory, propagating signal. Typically, instructions are stored as program code in the memory but can also be hardwired. Memory 920 may include a kernel and / or programming modules such as a software application that may be stored in either volatile or non-volatile memory.

[0108] Herein, the term “processor” is used to refer generically to any device or component that can process computer-readable instructions, including for example, a microprocessor, microcontroller, programmable logic device, or other computational device. That is, processor 905 may be provided by any suitable logic circuitry for receiving inputs, processing them in accordance with instructions stored in memory, and generating outputs (for example to the memory components or media player 940). In the present disclosure, processor 905 may be a single core or multi-core processor with memory addressable space. In one example, processor 905 may be multi-core, comprising — for example — an 8 core CPU. In another example, it could be a cluster of CPU cores operating in parallel to accelerate computations.

[0109] Further, one skilled in the art will recognize that certain functional units in this description have been labelled as modules throughout the specification. The person skilled in the art will also recognize that a module may be implemented as circuits, logic chips or any sort of discrete component. Still further, one skilled in the art will also recognize that a module may be implemented in software which may then be executed by a variety of processor architectures. In embodiments of the disclosure, a module may also comprise computer instructions or executable code that may instruct a computer processor to carry out a sequence of events based on instructions received. In further embodiments, the module may comprise a combination of different types of modules or sub-modules. The choice of the implementation of the modules may be determined by a person skilled in the art and does not limit the scope of the claimed subject matter in any way.

[0110] All examples described herein, whether of apparatus, methods, materials, or products, are presented for the purpose of illustration and to aid understanding, and are not intended to be limiting or exhaustive. Modifications may be made by one of ordinary skill in the art without departing from the scope of the invention as claimed.

Claims

CLAIMS1. A system, comprising:memory storing instructions; anda processor coupled to the memory and configured to process the stored instructions to implement:a module configured to perform a method of defect detection, the method including:receiving a plurality of first spectra associated with a first food sample disposed in a sensing space, the plurality of first spectra being measured by moving the first food sample relative to a spectroscope, wherein each of the plurality of first spectra corresponds to a portion of the first food sample; obtaining a plurality of first aggregates by combining selected ones of the plurality of first spectra, each of the plurality of first aggregates comprising multiple ones of the plurality of first spectra, wherein each of the plurality of first aggregates corresponds to a respective one of a plurality of first spatial zones of the sensing space;providing the plurality of first aggregates to a machine learning model; and obtaining a defect prediction characteristic of the first food sample based on an output of the machine learning model.

2. The system as recited in claim 1, wherein the plurality of first spectra is measured by moving the first food sample relative to the spectroscope in at least one rotation about a continuous closed path.

3. The system as recited in any one of the above claims, wherein the plurality of first spectra is measured by moving the first food sample relative to the spectroscope in a pluralityof rotations about a continuous closed path, wherein the method further comprises: dividing the plurality of first spectra into a plurality of first spectrum sections, wherein each of the plurality of first spectrum sections comprises multiple ones of the plurality of first spectra.

4. The system as recited in claim 3, wherein the method further comprises: based on a similarity measure between ones of the plurality of first spectrum sections, determining from the plurality of first spectrum sections: a starting spectrum and an ending spectrum, the starting spectrum to ending spectrum corresponding to a single one of the plurality of rotations about the continuous closed path.

5. The system as recited in claim 4, wherein selected ones of the plurality of first spectra from the starting spectrum to the ending spectrum are combined into the plurality of first aggregates and provided as the plurality of first aggregates to the machine learning model.

6. The system as recited in any one of claims 1 and 2, wherein the plurality of first spectra is measured by moving the first food sample relative to the spectroscope in a plurality of rotations about a continuous closed path, wherein the method further comprises: dividing the plurality of first aggregates into a plurality of first aggregate sections, wherein each of the plurality of first aggregate sections comprises multiple ones of the plurality of first spectra.

7. The system as recited in claim 6, wherein the method further comprising: based on a similarity measure between ones of the plurality of first aggregate sections, determining from the plurality of first aggregates: a starting aggregate and an ending aggregate, the starting aggregate to ending aggregate corresponding to a single one of the plurality of rotations about the continuous closed path.

8. The system as recited in claim 7, wherein selected ones of the plurality of first aggregates from the starting aggregate to the ending aggregate are provided to the machine learning model9. The system as recited in any one of the claims 3 to 8, wherein the spectroscope measures at a first sampling rate for one of the plurality of rotations, and measures at a second sampling rate for another one of the plurality of rotations, wherein the first sampling rate is different from the second sampling rate.

10. The system as recited in any one of the above claims, wherein the plurality of first aggregates corresponds to a plurality of non-overlapping spatial zone of the sensing space.

11. The system as recited in any one of claims 1 to 10, wherein the plurality of first aggregates corresponds to a plurality of uniformly divided spatial zones of the sensing space.

12. The system as recited in any one of claims 1 to 11, wherein the plurality of first aggregates corresponds to a plurality of non-uniformly divided spatial zones of the sensing space.

13. The system as recited in any one of the above claims, wherein a respective pair of adjacent first spatial zones are spaced apart in the sensing space.

14. The system as recited in any one of the above claims, wherein the defect prediction of the first food sample comprises at least one of: a defect identification prediction; a defect rate prediction; a prediction of a probability of defect; and a defect class prediction.

15. The system as recited in any one of the above claims, wherein the method further comprising: using a clustering model, obtaining the plurality of first aggregates by combining selected ones of the plurality of first spectra.

16. The system as recited in any one of the above claims, further comprising a light source for illuminating the first food sample.

17. The system as recited in any one of the above claims, wherein the method further comprises: varying a relative speed of moving the first food sample relative to the spectroscope.

18. The system as recited in any one of the above claims, further comprising a mixer disposed in the sensing space for moving the first food sample relative to the spectroscope in three orthogonal directions.

19. The system as recited in any one of the above claims, wherein the method further comprising: performing at least one pre-processing step on the plurality of first spectra prior to obtaining the plurality of first aggregates, wherein the at least one pre-processing step comprises one or more of: a clustering step; a truncation step; a normalization step; a noise reduction step; and a dimensional reduction step.

20. The system as recited in any one of the above claims, wherein first food sample comprises a plurality of granular food sample.

21. The system as recited in any one of the above claims, wherein the method further comprising:receiving a plurality of second spectra associated with a second food sample disposed in the sensing space, the plurality of second spectra being measured by moving the second food sample relative to a spectroscope, wherein each of the plurality of second spectra corresponds to a portion of the second food sample;obtaining a plurality of second aggregates by combining selected ones of the plurality of second spectra, each of the plurality of second aggregates comprising multiple ones of the plurality of second spectra, wherein each of the plurality of second aggregates corresponds to a respective one of a plurality of second spatial zones of the sensing space;providing the plurality of first aggregates and the plurality of second aggregates to the machine learning model; andobtaining a collective defect prediction characteristic of the first food sample and the second food sample based on an output of the machine learning model.

22. The system as recited in claim 22, wherein the collective defect prediction of the first food sample and the second food sample comprises at least one of: a defect identification prediction; a defect rate prediction; a prediction of a probability of defect; and a defect class prediction.

23. The system as recited in any one of the above claims, wherein the machine learning model comprises at least one of: a supervised classifier, an unsupervised classifier, a multipleinstance learning (MIL) model, a learning from label proportions (LLP) model.

24. The system as recited in any one of the above claims, wherein the machine learning model comprises an ensemble learning-based model.

25. The system as recited in any one of the above claims, wherein the defect prediction corresponds to at least one non-visually observable defect.

26. A system, comprising:memory storing instructions; anda processor coupled to the memory and configured to process the stored instructions to implement:a module configured to perform a method of defect detection, the method including:receiving a plurality of first aggregates obtained by combining selected ones of a plurality of first spectra, each of the plurality of first aggregates comprising multiple ones of the plurality of first spectra, wherein the plurality of first spectra is associated with a first food sample disposed in a sensing space, the plurality of first spectra being measured by moving the first food sample relative to a spectroscope, wherein each of the plurality of first spectra corresponds to a portion of the first food sample, wherein each of the plurality of first aggregates corresponds to a respective one of a plurality of first spatial zones of the sensing space; inputting the plurality of first aggregates to a machine learning model; and providing a defect prediction characteristic of the first food sample based on an output of the machine learning model.