Automatic detection of the chemical composition of moving objects

The system addresses the inaccuracies in hyperspectral image segmentation by using predefined profiles and composite bands to enhance relevant information, achieving efficient and precise object detection.

JP2026122944APending Publication Date: 2026-07-29X DEVELOPMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
X DEVELOPMENT LLC
Filing Date
2026-03-17
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing image segmentation techniques are unreliable and inaccurate due to real-world complexities such as partial occlusion, perspective issues, and varying illumination, especially in hyperspectral images, leading to false detections and increased computational cost.

Method used

A computer system uses hyperspectral images with predefined profiles that specify different combinations of wavelength bands for accurate and efficient detection and segmentation of objects, employing optical systems with specific camera angles and light sources, and generating composite bands to enhance relevant information while filtering noise.

Benefits of technology

The system achieves high-accuracy and efficient image segmentation by selectively using subsets of wavelength bands tailored to object types, reducing computational expense and improving detection precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026122944000001_ABST
    Figure 2026122944000001_ABST
Patent Text Reader

Abstract

The goal is to more reliably and accurately detect the characteristics of what is depicted in an image and the elements it represents. [Solution] Image data is acquired showing the degree to which one or more objects reflect, scatter, or absorb light in each of multiple wavelength bands, and the image data is collected while a conveyor belt moves the object(s). The image data is preprocessed by performing analysis across frequency and / or analysis across spatial dimension representation. A set of feature values ​​is generated using the image preprocessed image data. A machine learning model generates an output using the feature values. The output is used to generate a prediction of the identity of a chemical substance in one or more objects or the level of one or more chemical substances in one or more objects. The output is a prediction of the identity of a chemical substance in one or more objects, or data showing the level of one or more chemical substances in at least one of the one or more objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit and priority of U.S. Patent Application No. 17 / 811,766, filed on July 11, 2022, and U.S. Patent Application No. 17 / 383,293, filed on July 22, 2021, which are hereby incorporated by reference in their entirety for all purposes.

[0002] This specification generally relates to collecting and automatically analyzing hyperspectral images of moving objects. More specifically, hyperspectral images of moving objects are collected, features are extracted, and a machine learning model is used to predict each of one or more components of an object depicted using the features.

Background Art

[0003] Computer vision can be used to automatically process images to predict, for example, what is depicted within the image. Various situations can introduce complexity into this automated processing. For example, in some cases, one object (or animal, plant, etc.) is in front of another object (or animal, plant, etc.), obscuring the other object partially. As another example, perspective problems can occur where the depiction of an object has a shape, texture, color, etc. that is different from one typical characteristic of the object. As yet another example, the illumination of the environment, or the brightness of a light source as seen by a camera, can be different on a given day compared to another day (or different for another environment or camera), which can complicate the effort to use intensity and / or frequency data from the captured image to accurately characterize what is depicted.

[0004] Image segmentation is a digital image processing technique that divides an image into meaningful parts such that pixels belonging to a particular part share similar characteristics. This enables the analysis of digital images by defining the shapes and boundaries of objects within the image. Image segmentation is widely used in several fields, including autonomous vehicles, medical imaging, and satellite imaging. However, image segmentation can be unreliable and / or inaccurate when real-world complexities of image processing arise (e.g., those mentioned above).

[0005] It would be advantageous if image processing systems and technologies were improved to more reliably and accurately detect the characteristics of what is depicted in an image and the elements depicted. [Overview of the project]

[0006] According to one innovative aspect of the subject matter described herein, a computer system can perform image segmentation with greater accuracy and efficiency than previous methods using hyperspectral images. Hyperspectral images provide image data about a subject across multiple optical bands with different wavelengths (e.g., “wavelength bands,” “spectral bands,” or simply “bands”) and contain significantly more information than conventional color or grayscale images. This information is provided in many forms and often includes more bands than a typical RGB image, information about spectral bands narrower than conventional RGB Bayer filter bands, and information about bands outside the visible range (e.g., infrared, ultraviolet, etc.).

[0007] However, not all wavelength bands in a hyperspectral image are relevant to the boundaries of each type of segmentation. As a result, depending on the type of object being imaged and its characteristics (e.g., material, composition, structure, texture, etc.), image data for different hyperspectral wavelength bands may indicate regional boundaries. Similarly, for some object types and regional types, information in some wavelength bands may add noise or actually obscure the desired boundaries, resulting in reduced segmentation accuracy and increased computational cost of segmentation analysis.

[0008] Furthermore, real-world imaging introduces complexities, such as those described above, which can lead to false detections and / or characterizations. Some of these complexities may be selectively more pronounced, or even more pronounced, in certain wavelength bands compared to others.

[0009] The techniques described below illustrate how computer systems can generate and use profiles that specify different combinations of wavelength bands to provide accurate and efficient detection and / or segmentation of different object and region types, and / or accurate and efficient characterization of depicted objects. Using these profiles, the system can selectively use image data within an image (e.g., hyperspectral and / or visible light image), resulting in the use of different combinations of image bands to identify the locations of different types of regions or boundaries within the image. For example, for a particular object type, a profile may indicate that for objects of that object type, a first type of region should be segmented using image data from bands 1, 2, and 3, and a second type of region should be segmented using image data from bands 3, 4, and 5. When processing an image of a particular object type, segmentation parameters specified in the profile are used, including subsets of bands for each region type, e.g., image data for bands 1, 2, and 3 to identify a first type of region, and image data for bands 3, 4, and 5 to identify a second type of region.

[0010] Hyperspectral images may be generated using an optical system comprising one or more light sources and one or more cameras (and / or one or more photosensors). The light source(s) may be configured to emit (e.g.) infrared light, near-infrared light, shortwave light, coherent light, etc. In some cases, the light source(s) may include (e.g.) optical fibers that can separate the thermal output of the light source from its spectral output. In some cases, at least one of the cameras(s) may be positioned such that the optical axis of the camera lens or image sensor is at 75–105 degrees, 80–90 degrees, 85–95 degrees, or 87.5–92.5 degrees, 30–60 degrees, 35–55 degrees, 40–50 degrees, 42.5–47.5 degrees, or less than 15 degrees relative to the surface (e.g., a conveyor belt) supporting the object(s) being imaged. In some cases, the optical system includes multiple cameras, and the angle between the optical axes of the first camera and the surface supporting the object(s) is different from the angle between the optical axes of the second camera and the surface. The difference may be (for example) at least 5 degrees, at least 10 degrees, at least 15 degrees, at least 20 degrees, at least 30 degrees, less than 30 degrees, less than 20 degrees, less than 15 degrees, and / or less than 10 degrees. In some cases, the first camera filters different types of light for the second camera. For example, the first camera may be an infrared camera and the second camera may be a visible light camera.

[0011] As an exemplary implementation, the characteristics and quality of fruit can be automatically evaluated by a computer vision system using fruit image segmentation. Beyond simply segmenting fruit from the background, the system can be used to segment different parts of the fruit from one another. In the case of a strawberry, the external components include leaves (e.g., calyx, sepals, fruit stalk), seeds (e.g., achenes), and pulp (e.g., receptacle). The pulp may have areas in different states, such as ripe, unripe, damaged, moldy, or rotten. To facilitate rapid and efficient machine vision analysis of individual strawberries for quality control or other purposes, the system can generate strawberry object type profiles that specify the type of area of ​​interest (e.g., leaves, seeds, and pulp) and a subset of the hyperspectral image bandwidth used to segment or identify areas of each area type. These subsets of bandwidth can be determined through data-driven analysis of training examples, which include hyperspectral images and ground truth segmentation showing the area types for the examples. A profile can specify other parameters for each region type, such as functions to apply to image data of different bandwidths and thresholds to use. Using defined profiles, the system can accurately and efficiently process strawberry hyperspectral images and segment each region type. For each region type, the system can define boundaries for instances of that region type using a subset of bandwidth and other parameters specified in the profile. As a result, each region type can be accurately segmented using the subset of bandwidth that best represents the region boundary, and processing becomes more efficient by limiting the number of bandwidths used for segmentation of each region type.

[0012] As another exemplary implementation, image segmentation of waste materials can be used to better identify and characterize recyclable materials. For example, a system can be used to accurately segment regions of image data representing different types of plastics (e.g., polyethylene (PE), polyethylene terephthalate (PET), polyvinyl chloride (PVC), polypropylene (PP), etc.) to automatically detect the material of an object and identify where different types of objects are located. Furthermore, segmentation techniques can be used to identify and characterize instances of additives and contaminants in materials. For example, in addition to identifying regions containing one or more main materials (e.g., PE vs. PET), or instead, segmentation techniques can also identify objects or parts of objects that contain different additives (e.g., phthalates, bromides, chlorates, UV-resistant coatings) or contaminants (e.g., oil, food residues, etc.). To better characterize different types of regions, the system can generate and store profiles for different types of objects and materials that specify the type of region in question (e.g., different types of materials, different additives present, different contaminants), as well as for subsets of hyperspectral image bandwidths used to segment or identify regions of each region type. These subsets of bandwidths can be determined through data-driven analysis of training examples, which may include ground truth segmentation indicating region types for hyperspectral images and examples. The profiles may specify other parameters for each region type, such as functions to apply to image data of different bandwidths and thresholds to use. Once the profiles are defined, the system can process hyperspectral images and segment each region type accurately and efficiently. For each region type, the system can define boundaries for regions composed of different materials, regions where different contaminants are detected, regions where different types of contamination exist, and so on.As a result, each region type can be precisely segmented using a subset of bandwidth that best represents the region boundary, and the process becomes more efficient by limiting the number of bandwidths used for segmentation of each region type.

[0013] As will be further explained below, the system can also define a composite band that modifies the bands before performing segmentation. The composite band can be based on one or more image bands in a hyperspectral image, but one or more functions or transformations may be applied to it. For example, the composite band may be a combination or aggregation of two or more bands to which a function is applied (e.g., addition, subtraction, multiplication, division, etc.). One example is to calculate a normalized index based on two bands, such as dividing the difference between the two bands by the sum of the two bands, as the composite band. For a hyperspectral image where the image for each band has dimensions of 500 pixels × 500 pixels, the result that produces the normalized indices for bands 1 and 2 may be a 500 pixels × 500 pixels 2D image, where each resulting pixel is calculated by combining two pixels Pband 1 and Pband 2 at the same position in the source image according to the formula (Pband 1 - Pband 2) / (Pband 1 + Pband 2). Another example is to accept all bands across a set of pixels and apply a function that maps each set of pixels to a single number. Next, this function is convolved along the height and width of the image to generate a new composite bandwidth across the entire image. The last example is to apply a function that projects all bandwidths to a single number to each individual pixel. Similarly, this function can be applied to each pixel in the image to generate a new composite bandwidth at each point. Of course, this is just one way of combining image data of different bandwidths, and many different functions can be used.

[0014] The composite band, along with other parameters, can be used by the system to amplify or enhance the type of information indicating region boundaries while filtering or reducing the influence of image information that does not indicate region boundaries. This provides an enhanced image to which a segmentation algorithm can be applied. By predefining the band and function for each target region, the segmentation process can be much faster and less computationally expensive than other techniques, such as processing each hyperspectral image using a neural network. Generally, the composite band can combine information about region boundaries distributed across image data of various different bands, allowing the system to extract hyperspectral image components that best signal region boundaries from various bands and combine them into one or more composite images that enable high-accuracy, high-confidence segmentation. Another advantage of this method is that it enables an empirical, data-driven approach to customizing segmentation for different object and region types, while requiring far less training data and training computation than is typically needed to train neural networks and similar models.

[0015] To generate a profile, the system can perform a selection process to identify subsets of wavelength bands in a hyperspectral image that enable more accurate segmentation of different region types. This process may include multiple phases or iterations applied to training examples. A first phase may involve evaluating the image data for each band and selecting a subset of individual bands that most clearly demonstrate the difference between target regions (e.g., the bands that show the highest or most consistent difference between a particular region to be segmented and one or more other region types represented in the training data). A predetermined number of bands, or a subset of bands that meet certain criteria, may be selected for further evaluation in a second phase. The second phase may include generating composite bands based on the application of different functions to the individually selected bands from the first phase. For example, if bands 1 and 2 were selected in phase 1, the system may generate several different candidate bands based on different ways of combining those two bands (e.g., band 1 minus band 2, band 1 plus band 2, normalized index of bands 1 and 2, etc.). The system then evaluates how clearly and consistently the normalized bands distinguish the target region from other regions and can select a subset of these composite bands (e.g., a predetermined number with the highest scores, or those with region type identification scores above a minimum threshold). The selection process can optionally continue into further phases to evaluate different combinations of composite bands and select from them, with each additional phase selecting a new combination that provides higher accuracy and / or consistency in identifying the target region type.

[0016] In applications of optical sorting and classification of plastics, the system's discriminative power can be significantly improved by generating combined or composite bands of image data and determining which bands should be used to detect different materials. For example, analysis may be performed to determine which bands best identify different base plastic types and to distinguish them from other common materials. Similarly, analysis can be used to select bands that best identify a base type of plastic without additives (e.g., pure PE) from plastics of the same type containing one or more additives (e.g., phthalates, bromides, chlorates, etc.) and to distinguish uncontaminated areas from areas with different types of surface contaminants. Band selection may depend on the type of base plastic that sets the baseline amount of reflectance and variation in a particular spectral region. Thus, different combinations of bands may be selected to identify areas of different additives or contaminants present. In some implementations, band selection may be indicated by the set of materials to be identified and the specific types of additives and contaminants of interest.

[0017] In some cases, preprocessing is performed before a portion of an image (e.g., a segment, region, pixel group, or pixel) is processed (e.g., for object identification, object characterization, component detection, etc.). Preprocessing may include identifying one or more features. Subsequent processing (e.g., predicting whether a depicted object contains one or more components) may then be configured to process a relatively small set of features instead of many (e.g., hundreds to thousands) intensities in one or more spectra. For example, normalization and / or subtraction processes may be performed using a spectrum and one or more other baseline spectra. As another example, a spectrum may be preprocessed by computing the derivative (or second derivative) of the spectrum, and the frequencies corresponding to the intersection of the baseline or zero line in the derivative (or second derivative) can be defined as features. As yet another example, a spectrum (or the derivative or second derivative of a spectrum) may be convolved using a kernel (e.g., which may have a Gaussian shape, a single sinusoidal cycle shape, a shape corresponding to one or more chemicals, etc.), and the convolved spectrum (itself or its properties) may be defined as features. As yet another example, baselines may be removed from the spectrum. Another example involves using smoothing techniques (e.g., local least squares or Savitzky-Golay) to identify one or more features from a spectrum. For example, the Savitzky-Golay method can fit a p-th degree polynomial (if the polynomial is differentiable) to n points surrounding each data point, thereby smoothing the data and providing an easy estimate of the derivative. Another example is using local least squares to fit any functional form to the data.

[0018] In one general embodiment, a method performed by one or more computers includes: one or more computers acquiring image data of a hyperspectral image, wherein the image data includes image data for each of a plurality of wavelength bands; one or more computers accessing stored segmentation profile data for a particular object type, which indicates a predetermined subset of wavelength bands designated for segmenting images of objects of a particular object type into different region types; one or more computers segmenting the image data into a plurality of regions using the predetermined subset of wavelength bands designated in the stored segmentation profile data for segmenting into different region types; and one or more computers providing output data indicating the plurality of regions and the respective region types of each of the plurality of regions.

[0019] In some implementations, a computer system uses a machine learning model to predict the chemical composition of an object based on an image of the object. The machine learning model can be trained on an image of the object and the results of conventional chemical analysis. For example, the training image may be of the object before a destructive testing process, and the training label or training target may be the result of the destructive testing process, such as the amount or concentration of one or more chemicals in the object. Through training, the machine learning model can learn to predict the chemical properties of an object based on the characteristics of the object's undamaged external appearance, enabling the prediction of chemical composition without damaging the object.

[0020] One application of this technology is predicting the sugar content in fruit based on images of the fruit. The system can acquire hyperspectral images of the fruit and feed the data from the hyperspectral images to a machine learning model trained to predict sugar content. The model may be trained on a specific type of fruit to achieve high accuracy. For example, one model may be specialized to predict the sugar content of strawberries, another model may be trained to predict the sugar content of cherries, and so on. The model may be configured to provide regression outputs such as predictions of numerical values ​​indicating the amount of the chemical substance of interest. For example, the model may predict a value in Brix degrees or another unit representing sugar concentration. These technologies can provide a rapid and non-destructive determination of sugar content and other chemical components of fruit or other types of samples.

[0021] Another application of this technology is predicting the chemical composition of other materials, such as plastics. This technology can be used to assess the composition of waste materials, facilitating material sorting and recycling. The system can segment hyperspectral image regions corresponding to different plastic types. For example, a facility could include a conveyor that moves materials, such as plastic items (alone or mixed with other items), taking hyperspectral cameras into consideration. The captured hyperspectral image is then segmented using different segmentation profiles for different base plastics (e.g., PE, PEG, PET, etc.). The segmentation profiles can specify different sets of spectral bands, including potentially different composite bands, that modify or combine the captured spectral band data to segment regions of different base plastic types and to distinguish regions with contaminants and plastics with additives from regions where the base plastic exists. Using the segmented image data, the system can generate inputs for machine learning models used to predict chemical composition and concentration, using segmented regions that isolate data describing specific materials (e.g., regions of base plastic, regions of plastic with specific additives, regions where contamination exists, etc.). For example, by segmenting a hyperspectral image to identify specific regions of PET-based plastics that have oil contamination from food residues, the system can use the spectral data of those segmented regions to generate input to a machine learning model trained to characterize the contamination, for example, predicting the chemical composition and concentration of the contaminant. Different models can be created to predict different chemical properties or properties of different regions. For example, some models can be trained to predict the concentrations of different contaminants, while others can be trained to predict the concentrations of different additives, and so on.When multiple types of regions are segmented, hyperspectral image data for each region type can be processed using a machine learning model trained for chemical analysis of that particular region type.

[0022] In some implementations, imaging or scanning data generated through techniques other than hyperspectral imaging can be used for segmentation and to predict the quantity and concentration of chemical substances. For example, X-ray fluorescence and laser-induced breakdown spectroscopy can be used in addition to or as an alternative to hyperspectral imaging. More generally, any suitable form of spectroscopy can be used, such as absorption spectroscopy, emission spectroscopy, elastic scattering and reflection spectroscopy, impedance spectroscopy, Raman spectroscopy and other analyses of inelastic scattering phenomena, coherent or resonance spectroscopy, or nuclear spectroscopy. These additional analytical techniques can be used to generate data that is essentially another information band that the computer system can consider. When deciding which information bands to use to segment different types of information and which bands to use to predict different chemical properties, the computer system can select which spectroscopic techniques and parts of the spectroscopic results are most useful regarding the presence and quantity of different chemical substances. Similarly, the information obtained from these spectroscopic techniques can be used to train models and provide the information as input to the trained models to generate predictions.

[0023] Generally, one innovative aspect of the subject matter described herein can be carried out in a manner that includes: an action by one or more computers to acquire image data of an image containing a representation of an object, wherein the image data includes image data for each of a plurality of wavelength bands, each including one or more wavelength bands outside the visible wavelength range; an action by one or more computers to provide feature data derived from the image data to a machine learning model trained to predict levels of chemical components within an object in accordance with feature data derived from image data representing the exterior of the object; an action by one or more computers to receive an output of a machine learning model generated in accordance with the feature data, wherein the output includes predictions of levels of chemical components within an object represented in the image; and an action by one or more computers to provide output data indicating the predicted levels of chemical components within an object.

[0024] In some implementations, the output data is displayed on a screen by one or more computers. In some implementations, the output data is provided to client devices via a communication network. The output data may be provided upon request from a client device associated with the image data.

[0025] A machine learning model may be trained to provide values ​​indicating the predicted concentration or amount of additives or contaminants in plastics. For example, a machine learning model may be trained to indicate the predicted level of additives present in a particular type of plastic, in response to receiving feature data derived from images of a particular type of plastic. In some implementations, the prediction is based solely on image data (and / or features of the image data), without receiving any results of other chemical analysis of the plastic or any modification of the object (e.g., destructive chemical treatment of the plastic). In some implementations, the output is a regression output, such as a numerical value indicating the predicted concentration of a chemical substance. In other implementations, the output of the machine learning model may be a classification of whether a chemical substance is present or not, or a classification corresponding to different ranges or levels of concentration. Examples of additives that can be evaluated include phthalates, bromides, chlorates, and coatings. Similarly, the types and concentrations of different types of contaminants, such as oils, greases, and food residues, can also be predicted using an appropriate model. In some implementations, the model may be trained to predict the properties of the base plastic resin, such as the type of resin(s) present and the level of resin purity.

[0026] In some implementations, the system may be used to generate an output for a single object, such as a single plastic item. In other implementations, the system may be used to generate an output characterizing a collection of multiple items, such as a group of plastic items on a conveyor belt in a recycling facility. In some implementations, different chemical output predictions can be determined for different parts of an object, indicating different levels of contamination in different parts of the object.

[0027] The machine learning model can be trained using training data that includes (i) exemplary image data of plastic items and (ii) indicators of the chemical composition of the plastic items represented in the exemplary image data. The machine learning model can be trained using training data that includes (i) exemplary image data for a specific type of base plastic (e.g., PE, PP, PET, etc.) and (ii) indicators of the level of one or more additives or contaminants that have been measured for the plastic items represented in the exemplary image data. The machine learning model is trained to remotely and non-destructively predict the chemical composition of the plastic represented within an image, and the machine learning model is trained based on (i) exemplary image data that provides a representation of an exemplary plastic item prior to a physical or chemical test (which may include a destructive test) of the plastic item and (ii) measurement results obtained through testing of the exemplary plastic item represented within the exemplary image data.

[0028] The machine learning model can be trained to provide a value indicating the predicted concentration or amount of sugar. For example, the machine learning model can be trained to indicate a predicted level of sugar content in a specific type of fruit in response to receiving feature data derived from an image of the specific type of fruit. In some implementations, the prediction is based only on the image data, without receiving any results of other chemical analyses regarding the object (e.g., the fruit), and without modifying the object (e.g., without cutting or obtaining information about the interior of the fruit).

[0029] The output of the machine learning model includes an indicator of the predicted level of sugar content in an entire fruit of a particular type (e.g., not cut or segmented), and the output data indicates the level of sugar content in an entire fruit of a particular type. The model can be generated and trained for a particular type of fruit and also to potentially predict the concentration or amount of a particular chemical substance or set of chemical substances. Different models can use different features such as features for different sets of spectral bands to predict the characteristics of different types of objects or to predict different chemical characteristics such as the concentration or amount of different chemical substances.

[0030] The machine learning model can be trained to indicate a predicted level of sugar content in the juice of a fruit in response to receiving feature data derived from image data about an image of the exterior of the fruit. The machine learning model can be configured to indicate a predicted level of sugar content in the juice as at least one value of degrees Brix, degree Plato, specific gravity of sugar, mass fraction of sugar, or concentration of sugar. In other implementations, the model can predict a classification indicating the level or range of the amount of a chemical substance in a sample.

[0031] In some implementations, the particular fruit being evaluated is a particular strawberry, e.g., a particular single strawberry. The machine learning model is trained to indicate a predicted level of sugar content in the juice from the strawberry in response to receiving feature data derived from image data about an image of the exterior of the strawberry. The output of the machine learning model includes an indicator of the predicted level of sugar content of the juice from the particular strawberry, and the output data indicates the predicted level of sugar content of the juice from the particular strawberry.

[0032] A machine learning model can be trained using training data that includes (i) exemplary image data of fruit and (ii) an index of the sugar content of the fruit represented in the exemplary image data. A machine learning model has been trained using training data that includes (i) exemplary image data of strawberries and (ii) an index of the sugar content level measured for the strawberries represented in the exemplary image data. A machine learning model has been trained to non-destructively predict the sugar content of fruit represented in images, and the machine learning model has been trained on (i) exemplary image data that provides a representation of the exemplary fruit before a destructive test of the exemplary fruit and (ii) results obtained through a destructive test of the exemplary fruit represented in the exemplary image data.

[0033] Various types of machine learning models can be used. For example, the machine learning model can be a decision tree or a neural network. In some cases, the machine learning model can be a gradient-boosted regression tree.

[0034] In some cases, image segmentation can be performed on image data to identify a predetermined type of region, and feature data is derived from the image data of the identified region of the predetermined type. The feature data provided to the model can exclude information about regions that are not of the predetermined type. For example, image data of regions of an image that show the object being analyzed but still do not meet certain predetermined criteria (e.g., representing parts of an object that are undesirable for analysis of a particular characteristic) can be excluded from the input to the model. Thus, features can represent information about a predetermined type of region while omitting information about regions that are not of the particular type. A predetermined type of region used for analysis might be a region showing the flesh of a strawberry, and the feature data input to the machine learning model would be derived from image data representing the strawberry flesh. Image data of other regions, such as regions corresponding to the calyx and achene of the strawberry, can be excluded from the analysis.

[0035] The image data used for analysis can be hyperspectral image data, and the data is acquired over one or more wavelength bands outside the visible wavelength range. For example, the image data may include reflectance (or absorptive or absorbance) data in one or more wavelength bands of infrared light. Feature data for input to the model can also be based on image data captured over one or more wavelength bands of infrared light.

[0036] In some embodiments, a computer implementation method is provided, comprising: acquiring image data indicating the extent to which one or more objects reflect, scatter, or absorb light in each of a plurality of wavelength bands, the image data being collected while a conveyor belt is moving one or more objects; preprocessing the image data to generate preprocessed image data, wherein the preprocessing includes performing an analysis over frequency and / or an analysis over spatial dimension representation; generating a set of feature values ​​derived from the preprocessed image data; generating predictions of chemical identity in one or more objects or levels of one or more chemicals in one or more objects based on the output generated by a machine learning model in response to a set of feature values ​​provided as input to a machine learning model; and providing data indicating predictions of chemical identity in one or more objects or levels of one or more chemicals in at least one of the one or more objects.

[0037] Other implementations of this embodiment and other embodiments include corresponding systems, devices, and computer programs configured to perform the actions of the method and encoded on computer storage devices. One or more computer systems can be configured in this way by software, firmware, hardware, or a combination thereof installed on the system that causes the system to perform actions during operation. One or more computer programs can be configured in this way by having instructions that cause the device to perform actions when executed by a data processing device.

[0038] Details of one or more embodiments of the subject matter described herein are shown in the accompanying drawings and the following description. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. [Brief explanation of the drawing]

[0039] [Figure 1] This block diagram shows an example of a system implemented to perform image segmentation using hyperspectral images. [Figure 2] This figure illustrates an example of performing enhanced image segmentation using profiles to select different subsets of wavelength bands in a hyperspectral image in order to segment different types of regions of an object. [Figure 3] This figure shows an example of automatically generating and selecting different bandwidths of image data for performing image segmentation. [Figure 4] This flowchart illustrates the process of automatically generating and selecting the wavelength bands used when segmenting an image of an object. [Figure 5A] This figure shows an example of a system that predicts the chemical composition of an object based on an image of the object. [Figure 5B] This figure shows an example of a system that predicts the chemical composition of an object based on an image of the object. [Figure 6A] This is a block diagram showing image segmentation of hyperspectral images of a specific type of fruit. [Figure 6B] This block diagram shows the generation of a set of feature values ​​from a segmented image to obtain a feature vector to provide as input to a machine learning model. [Figure 7] This flowchart illustrates the process of performing chemical analysis using image segmentation and machine learning. [Figure 8A] This shows exemplary hyperspectral data for a dark-referenced hyperspectral image. [Figure 8B] This shows exemplary hyperspectral data for a dark-referenced hyperspectral image. [Figure 8C] This shows exemplary hyperspectral data for a dark-referenced hyperspectral image. [Figure 9A] Corresponding exemplary hyperspectral data for a light-referenced hyperspectral image is shown. [Figure 9B] Corresponding exemplary hyperspectral data for a light-referenced hyperspectral image is shown. [Figure 9C] Corresponding exemplary hyperspectral data for a light-referenced hyperspectral image is shown. [Figure 10] An example of the spectrum of polyethylene plastic is shown. [Modes for carrying out the invention]

[0040] Figure 1 is a block diagram of an exemplary system 100 implemented to perform band selection and image segmentation of hyperspectral images. System 100 includes a camera system 110 for capturing images of an object (e.g., hyperspectral images). Each of one or more of the captured images may have image data for each of a plurality of bands, each band representing a measurement of reflected light for a particular band of wavelength. Figure 1 further illustrates an exemplary data flow shown in steps (A) to (E). Steps (A) to (E) may occur in the order shown, or may occur in a different order.

[0041] System 130 can be used to perform image segmentation for many different applications by selecting bandwidths and / or image features. For example, the system can be used to select bandwidths for identifying and evaluating different types of fruits, vegetables, meats, and other foods. As another example, the system can be used to select bandwidths and / or features for identifying and evaluating waste materials, such as detecting recyclable material types, as well as detecting the presence of additives or contamination.

[0042] In applications of optical sorting and classification of plastics, the system's discriminative power can be significantly improved by utilizing strategic pre-processing techniques, strategic processing techniques, and / or generating combined or composite bands of image data to determine which bands should be used to detect different materials. For example, pre-processing may include performing normalization on a reference image, generating derivatives, detecting frequencies corresponding to baseline (or zero) crossings, performing filtering, generating convolutions, and / or performing other calculations to produce outputs (e.g., frequencies with detected baseline crossings, convolution sum statistics, etc.). As another example, analysis can be performed to determine which bands and / or features best identify different base plastic types and distinguish them from other common materials. Similarly, analysis can be used to select bands that best identify base plastic types without additives (e.g., pure PE) from plastics of the same type containing one or more additives (e.g., phthalates, bromides, chlorates, etc.), and to distinguish uncontaminated areas from areas with different types of surface contaminants. The selection of bandwidths may depend on the type of base plastic, which sets the baseline amount of reflectance (or absorptive or absorbance) and its variation in a particular spectral region. Therefore, different combinations of bandwidths may be selected to identify regions of different additives or contaminants present. In some implementations, bandwidth selection may be indicated by the set of materials to be identified and the specific types of additives and contaminants of interest.

[0043] In the example in Figure 1, the camera system 110 captures a hyperspectral image 115 of an object 101, which is a strawberry in the figure. Each hyperspectral image 115 contains image data for N bands. Generally, a hyperspectral image can be thought of as having three dimensions x, y, and z, where x and y represent the spatial dimensions of a 2D image for a single band, and z represents an index or step through the number of wavelength bands. Thus, a hyperspectral image contains multiple 2D images, each represented by the x and y spatial dimensions, and each image represents the captured light intensity (e.g., reflectance) of the same scene for different spectral bands of light.

[0044] The images collected and analyzed by the camera system 110 may include non-hyperspectral images, and it should be understood that any disclosure herein referring to hyperspectral images may be adapted to use non-hyperspectral images. For example, the camera system 110 may include lenses with low absorption in the visible, near-infrared, short-wave infrared, or mid-wave infrared ranges. Thus, the captured images may show signals in the visible, near-infrared, short-wave infrared, or mid-wave infrared ranges, respectively. The camera system 110 may include any of a variety of illumination sources, such as light-emitting diodes (which may be synchronized to the camera sensor exposure), incandescent light sources, lasers, and / or blackbody illumination sources. The light-emitting diodes included in the camera system 110 may have peak emission that matches the peak resonance of a given target chemical substance (e.g., a given type of plastic), and / or bandwidth that matches the absorption of multiple molecules. In some cases, multiple LEDs may be included in the camera system 110 so that they can cover a given particular spectral region. LED light sources provide coherent light, and controlling emission and signal-to-noise ratios may be more feasible than with other light sources. Furthermore, they are generally more reliable and consume less power compared to other light sources.

[0045] In some cases, the camera system 110 is configured such that the optical axis of the camera lens or image sensor is at an angle of 75–105 degrees, 80–90 degrees, 85–95 degrees, 87.5–92.5 degrees, 30–60 degrees, 35–55 degrees, 40–50 degrees, 42.5–47.5 degrees, or less than 15 degrees with respect to the surface (e.g., conveyor belt) supporting the object(s) being imaged. In some cases, the optical system includes multiple cameras, and the angle between the optical axes of the first camera with respect to the surface supporting the object(s) is different from the angle between the optical axes of the second camera with respect to the surface. The difference may be (e.g.) at least 5 degrees, at least 10 degrees, at least 15 degrees, at least 20 degrees, at least 30 degrees, less than 30 degrees, less than 20 degrees, less than 15 degrees, and / or less than 10 degrees. The difference may facilitate the detection of signals from objects having different shapes or being positioned at different angles with respect to the underlying surface (e.g., having different inclinations). In some cases, the first camera filters different types of light for the second camera. For example, the first camera may be an infrared camera and the second camera may be a visible light camera. In some cases, the camera system 110 includes a light source that is specularly reflective to the camera and a second light source that is diffusely reflective to the camera (for example, to facilitate the detection of objects having different specular and diffuse reflectances).

[0046] The camera system 110 may include an optical guide (e.g., fiber optic, hollow, solid, or liquid-filled type) for transmitting light from the illumination source to the imaging location, which can reduce the heat emitted at the imaging location. The camera system 110 may include a light source or optical system of a type such that light from the light source(s) is focused into a line or to match the projection size of the entrance slit to the spectrometer. The camera system 110 may be configured such that the illumination source and imaging device (camera) are arranged under specular reflection conditions (to produce a bright-field image), non-specular (or diffuse) conditions (to produce a dark-field image), or a mixture of conditions. Most hyperspectral images have image data for each of several or tens of wavelength bands, depending on the imaging technique. In many applications, processing images with a large number of bandwidths is computationally expensive (resulting in delays and high power consumption when obtaining results), high-dimensional spaces can prove impractical to explore, or they have inappropriate distance metrics ("curse of dimensionality"), or bandwidths are highly correlated ("correlation regressor problem"), making it desirable to reduce the number of bandwidths in hyperspectral images to a manageable amount. Many different dimensionality reduction techniques have been proposed in the past, such as principal component analysis (PCA) and pooling. However, these techniques are still computationally expensive, require specialized training, and often do not provide the desired accuracy in applications such as image segmentation. Furthermore, many techniques attempt to use almost all or all bandwidths for segmentation decisions, even though different wavelength bands often have dramatically different information values ​​for segmenting different types of boundaries (e.g., boundaries of different types of regions with different properties such as material, composition, structure, and texture). This has traditionally resulted in the inefficiency of processing image data over more wavelength bands than are required for segmentation analysis. Furthermore, data in bandwidths with low relevance to the segmentation boundary obscures key signals within data containing noise and only slightly relevant data, thus limiting accuracy.

[0047] In particular, the importance of different wavelength bands to segmentation decisions varies significantly depending on the type of region. Of the 20 different wavelength bands, one type of region (e.g., having a specific material or composition) may strongly interact with only a portion of the total band being imaged, while a second type of region (e.g., having a different material or composition) may strongly interact with different subsets of the total band being imaged. Many conventional systems lack the ability to determine, store, and use region-dependent variations in which a subset of bands produces the best segmentation results, which often results in inefficient handling of band data that is only slightly relevant or irrelevant to the segmentation of at least some of the target regions. As described below, the technique described herein allows segmentation parameters for each object type and region type to be determined and stored based on an analysis of training examples, and then used to better identify and distinguish each type of region for a given type of object. This can be done for many different object types, and allows the system to select profiles for different objects or scenes and segment the various region types that may exist for different objects or scenes using appropriate sets of bands and parameters.

[0048] In the example in Figure 1, the camera system 110 includes or is associated with a computer or other device that can communicate via a network 120 with a server system 130 that processes hyperspectral image data and returns segmented images or other data derived from segmented images. In other implementations, the functions of the computer system 130 (e.g., generating profiles, processing hyperspectral image data, performing segmentation, etc.) may be performed locally at the location of the camera system 110. For example, system 100 can be implemented as a standalone unit housing the camera system 110 and the computer system 130.

[0049] Network 120 may include a local area network (LAN), a wide area network (WAN), the internet, or a combination thereof. Network 120 may also include any type of wired and / or wireless network, satellite network, cable network, Wi-Fi network, mobile communication network (e.g., 3G, 4G, etc.), or any combination thereof. Network 120 may utilize communication protocols, including packet-based and / or datagram-based protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Network 120 may further include several devices that facilitate network communication and / or form the hardware infrastructure for the network, such as switches, routers, gateways, access points, firewalls, base stations, repeaters, or a combination thereof.

[0050] In some implementations, the computer system 130 provides a bandwidth selection and image segmentation module that analyzes images and provides a selected bandwidth configuration and segmented images as output. In some implementations, the computer system 130 may be implemented by a single remote server or by a group of multiple different servers distributed locally or globally. In such implementations, the functions performed by the computer system 130 can be performed by multiple distributed computer systems, and the machine learning model is provided as a software service via the network 120.

[0051] In short, Figure 1 shows an example of how the computer system 130 generates segmentation profiles of object types and / or region types through the analysis of various training examples. The computer system 130 then receives additional hyperspectral images and uses the object type profiles of the objects in the images to efficiently generate accurate segmentation results. Figures 1, 2, 5, and 6A-6B show strawberries as the type of object to be detected and evaluated, but the same technique described can be used to process other types of objects.

[0052] During stage (A), as part of the setup process, the computer system 130 generates profiles of the types of objects whose images will be segmented. For example, to enable the system to segment an image of a strawberry, a profile 153 of the strawberry object type can be created. Profile 153 can specify a subset of bandwidth to use when segmenting the strawberry, or more potentially, different bandwidths to use for segmenting different types of regions of the strawberry.

[0053] To generate object type profiles, the computer system 130 processes various training examples 151, each containing hyperspectral images of instances of the object type to be profiled. Each hyperspectral image may be preprocessed according to one or more techniques. Preprocessing may include (for example) modifying the spectrum corresponding to each pixel in the hyperspectral image. Modification may include (for example) calculating the first derivative of the spectrum, calculating the second derivative of the spectrum, convolving the spectrum (or the derivative or second derivative of the spectrum) with each of one or more kernels (e.g., a Gaussian kernel, a single-cycle sinusoidal kernel, or a kernel having a signature of a given type of material), removing a baseline from the spectrum, and / or normalizing the spectrum based on one or more reference spectra (e.g., corresponding to light and dark conditions where the light source in the camera system 110 is completely on or completely off and there are no objects in the field of view). For example, baseline removal and / or normalization may be performed to produce results showing that the intensity in the frequency bands in the spectrum decreases to a maximum value corresponding to light conditions and a minimum value corresponding to dark conditions (for each of the set of frequency bands).

[0054] In some implementations, the bandwidth evaluation module 150 performs a bandwidth selection process in which each bandwidth of the processed hyperspectral image is analyzed to generate a selected bandwidth configuration that enables high accuracy while performing hyperspectral image segmentation.

[0055] The band evaluation module 150 can perform an iterative process of band selection 152 for object type and / or region type. During the first iteration, individual bands of a hyperspectral image undergo selection process 152a. Process 152a selects a subset of bands from multiple bands of a hyperspectral training image. For example, during the first iteration, module 150 evaluates bands 1-N of various hyperspectral image training examples 151 and assigns each band a score indicating how well it distinguishes between a particular type of target region (e.g., strawberry pulp) and other regions (e.g., leaves, seeds, background, etc.). In this example, the first iteration of process 152 selects bands 1 and 3 from bands 1-N. As another example, module 150 may identify bands associated with (e.g.) peaks in a spectrum, threshold crossovers in the derivative or second derivative of a spectrum, local minima in a spectrum, etc.

[0056] In some implementations, after selecting a subset of individual bandwidths, a composite bandwidth or modified bandwidth is generated. A composite bandwidth may be generated by processing image data on one or more of the bandwidths selected in the first iteration. For example, each bandwidth in the subset of bandwidths can undergo one or more operations (e.g., image processing operations, mathematical operations, etc.) which may include operations that combine data from two or more different bandwidths (e.g., the bandwidths selected in the first iteration). Each of various predetermined functions can be applied to the image data for different combinations of the selected bandwidths (e.g., for each pair of bandwidths or each permutation within the selected subset of bandwidths). This can create a new set of composite bandwidths, each representing a different modification or combination of bandwidths to the bandwidths selected in the first iteration. For example, during selection by selection process 152a, module 150 performs operations on bandwidths 1 and 3 to create three new composite bandwidths, including (1) bandwidth 1 + bandwidth 3, (2) bandwidth 1 / bandwidth 3, and (3) bandwidth 1 - bandwidth 3. A composite band can also be derived from convolution or projection applied to image data, which are functions that map groups of pixels and individual pixels to a single number. A new composite band can be created by applying one or more convolutions and / or one and multiple projections across the entire image.

[0057] Next, the composite bands thus created are evaluated, for example, scored, to determine the level at which they distinguish between the target region type (e.g., strawberry pulp) and other region types. Then, for a second iteration, the computer system 130 selects from the composite bands in the selection process 152b. In this example, the composite band created as band 1-band 3 is selected by process 152b. The iterative process of generating new modified or composite bands and then selecting the most effective one from among them can continue until the desired level of accuracy is reached.

[0058] In this example, the information for segmenting the strawberry flesh is distilled or aggregated into a single 2D image. However, this is not mandatory, and in some implementations, profile 153 may indicate that multiple separate bands (e.g., original bands or combined / modified bands) should be generated and used for segmentation. For example, the system may specify that segmentation should use image data for three bands: band 1 + band 3, band 1 / band 3, and band 1-band 3.

[0059] The bandwidth evaluation module 150 performs a selection process for each of the multiple region types of the object type from which the profile 153 is being generated. This generates a selected subset of bandwidths to be used for each region type. If the selected bandwidths are composite bandwidths, the component input bandwidths and the functions applied to generate the composite bandwidths are stored in the profile. As a result, the profile 153 for an object type may include the selected bandwidth configurations to be used for each region type of the object, enabling high accuracy for image segmentation for each region type. For example, in the case of a profile 153 for segmenting a strawberry, the region types may be leaves, seeds, and pulp. As another example, in a profile for segmenting elements of a dining room, the multiple region types may include chairs as the first region, a table as the second region, and walls as the third region.

[0060] In some implementations, the region type of profile 153 represents regions of different materials, and as a result, segmentation can easily distinguish between regions of different materials shown in the image. The bandwidth evaluation module 140 generates bandwidth configurations for each material type to enable high accuracy while performing image segmentation. For example, to evaluate furniture, multiple material types include wood, plastic, and leather. More generally, the selection can determine image bandwidth parameters for any of various properties, including material, composition, texture, density, and structure.

[0061] In some implementations, the bandwidth evaluation module 150 performs a selection and operation process 152 for each of several condition types of an object. For example, the flesh of a strawberry may be thought to have different types of regions, such as ripe, unripe, damaged, and powdery mildew. The bandwidth evaluation module 150 may generate bandwidth configurations for each condition type and store them in a profile 153, enabling high-precision distinction between regions of different conditions while performing image segmentation.

[0062] The process for generating the profile described in stage (A) can be performed for many different object types to create a library of segmentation profiles 153, which can be stored and retrieved by system 130 to accurately segment each of the different object types. For each object type, multiple different region types may be specified, each region type having a corresponding wavelength band, operator, algorithm, and other parameters specified for use in segmenting the image region of that region type.

[0063] During stage (B), the camera system 110 captures a hyperspectral image of the object. For example, the camera system 110 takes a hyperspectral image 115 of the strawberry 101, which includes image data for each of N different wavelength bands. In some implementations, the hyperspectral image 115 may be sent as one of many images in a series of images of different objects, such as objects on a conveyor belt for manufacturing, packaging, or quality assurance.

[0064] During stage (C), the hyperspectral image 115 is transmitted from the camera system 110 to the computer system 130, for example, using the network 120. The hyperspectral image 115 may be transmitted in connection with a request to process the image, such as to generate a segmented image, to inspect the characteristics or quality of an object represented in the image, or for other purposes.

[0065] During step (D), upon receiving the hyperspectral image 115, the computer system 130 performs processing (and potentially preprocessing as well) to identify and generate image data for different types of regions to be segmented. As referred to herein, preprocessing may include (e.g.) modifying the spectrum corresponding to each pixel in the hyperspectral image. Modification may include (e.g.) calculating the first derivative of the spectrum, calculating the second derivative of the spectrum, convolving the spectrum (or the derivative or second derivative of the spectrum) with each of one or more kernels (e.g., a Gaussian kernel or a kernel having a signature of a given type of material), removing the baseline from the spectrum, and / or normalizing the spectrum based on one or more reference spectra (e.g., corresponding to light / dark conditions where there are no objects in the field of view). Removing the baseline from the spectrum may reduce or remove signal components due to (e.g.) CO2 in the environment. Normalizing the spectrum using reference spectra may result in more consistent spectral results despite variability in the intensity of the light source (e.g., over time or relative to another light source).

[0066] The processing may include a pre-segmentation step of identifying object types represented in the hyperspectral image 115, a step of retrieving object type profiles 153 (e.g., from a database and from profiles of multiple different object types), a step of pre-processing image data of different bands (e.g., a step of generating a composite or combined image, a step of applying thresholds, functions, or filters, etc.), a step of reducing the number of images (e.g., a step of projecting or combining image data from multiple bands into fewer images or a single image), and / or a step of preparing the hyperspectral image 115 for segmentation processing using parameters in the profile 153. In some cases, the processing may include detecting peaks (e.g., in the spectrum or a pre-processed version thereof), baseline crossings (e.g., in the spectrum or its first or second derivative), or zero crossings (e.g., in the first or second derivative of the spectrum). The frequency band(s) in which each peak, baseline crossing, or zero crossing was detected can be identified and used as features to characterize the spectrum.

[0067] The computer system 130 identifies the object type of the object 101 represented in the hyperspectral image 115, and then selects and retrieves a segmentation profile 153 corresponding to that image type. The data provided in relation to the hyperspectral image 115 can indicate the type of object 101 represented in the image. For example, a request to process the hyperspectral image 115 may include an instruction that the object being evaluated is of the "strawberry" object type. In another example, system 100 may be configured to repeatedly process hyperspectral images showing the same type of object, and as a result, computer system 130 is already configured to interpret or process incoming hyperspectral images 115 as images of strawberries. This may be the case in a manufacturing facility or packaging workflow where items of the same type are processed sequentially. In yet another example, computer system 130 may use an object recognition model to detect the type of object represented in the hyperspectral image 115, and then automatically select a profile corresponding to the identified object type.

[0068] Once an appropriate profile 153 is selected for the object type of the object depicted in the hyperspectral image 115, the computer system 130 processes the hyperspectral image 115 by applying the information in the selected profile 153. For example, the profile 153 may specify different composite or combined images to be generated from image data of different bands within the hyperspectral image 115. The computer system 130 may generate these images and may apply any other algorithms or operations specified by the profile. As a result, module 140 prepares one or more images to which segmentation processing has been applied. In some cases, this may result in a single 2D image, or different 2D images for each of several different region types to be segmented, or multiple 2D images for each of several different region types. In practice, module 140 can use the profile 153 to act as a preprocessing step, filtering out image data of bands unrelated to a given region type and processing the image data into a suitable format for segmentation.

[0069] During stage (E), the segmentation module 160 performs segmentation based on the processed image data from the image processing module 140. Segmentation can determine the boundaries of different objects and different types of regions of those objects. One way to see the segmentation process is that the module 160 can classify different regions of the image it receives (representing the corresponding hyperspectral image 115) into classes or categories, for example, assigning pixels as one of various types, background, or not part of an object, such as leaves, seeds, or pulp. For example, image data of a selected bandwidth configuration generated by the image processing module 140 can be subjected to any of the following segmentation algorithms: threshold segmentation, clustering segmentation, compression-based segmentation, histogram-based segmentation, edge detection, region extension techniques, partial differential equation-based methods (e.g., curve propagation, parametric methods, level-set methods, fast marching methods, etc.), graph partitioning segmentation, watershed segmentation, model-based segmentation, multiscale segmentation, and multispectral segmentation. The segmentation algorithm may use parameters specified by the profile 153 for each region type (e.g., threshold, weight, criterion, different model or model training state, etc.), and as a result, different region types can be identified using different segmentation parameters or different segmentation algorithms. The results of the segmentation can be represented as an image. One example is an image providing a 2D pixel grid, where pixels are given a value of "1" when they correspond to a particular region type (e.g., leaf) and a value of "0" otherwise.

[0070] In this example, profile 153 specifies three region types: strawberry leaves, seeds, and pulp. Profile 153 specifies these three regions to be segmented, as well as the bandwidth and parameters to be used to identify where these three regions exist. The segmentation module 160 generates output 160 containing three images, one for each of the three different region types. Thus, each image corresponds to a different region type and specifies the region of the 2D field of view of the hyperspectral image 115 occupied by an instance of a particular region type. In other words, the segmented images may include an image mask, or otherwise specify the boundaries of the region identified as containing a certain region type. In some cases, regions of different region types may all be specified in a single image using different values ​​that classify different pixels as corresponding to different regions (e.g., 0 for things that are not part of the background or object, 1 for leaves, 2 for seeds, 3 for strawberry pulp, etc.).

[0071] During stage (F), system 130 stores the segmentation results 160 and uses them to generate and provide outputs. Even if the segmentation boundaries are generated using image data for a subset of bands in the hyperspectral image 115, the determined boundaries can be used to process the image data for each of the bands in the hyperspectral image 115. For example, the hyperspectral image may have image data for 20 different bands, and the segmentation process may use image data only for bands 1 and 2. The resulting determined region boundaries can then be applied to segment or select defined regions in the image data for all 20 different bands. Since the images for different bands of the hyperspectral image 115 share the same view and perspective of object 101, segmentation based on image data for one band can be directly applied (e.g., overlay, project, or otherwise mapped) to the image data for the other bands. In this way, segmentation can be consistently applied across all images in the hyperspectral image 115.

[0072] The segmentation results 160 can be stored in a database or other data storage, associated with the sample identifier of object 101 and the captured hyperspectral image 115. For quality control and manufacturing applications, system 130 can use the association as part of detecting and recording defects, tracking the quality and characteristics of specific objects as they move through a facility, and assisting in sampled analysis of lots or batches of objects. Some common functions that system 130 can perform using segmented hyperspectral image data include characterizing object 101 or specific parts thereof, such as assigning scores for composition, quality, size, shape, texture, or other characteristics. Based on these scores, or potentially as a direct output of image analysis without intermediate scores, computer system 130 can classify objects based on the segmented hyperspectral image data. For example, system 130 can classify objects into categories such as different quality grades and direct the objects to different areas using a conveyor system based on the assigned categories. Similarly, system 130 can detect defective objects and remove them from the manufacturing or packaging pipeline.

[0073] The segmentation results 160, the results of applying segmentation to the hyperspectral image 115, and / or other information generated using them can be provided. In some implementations, one or more images showing the segmented boundaries are sent to the camera system 110 or another computing device for display or further processing. For example, boundaries of different region types determined through segmentation can be overlaid on the region boundaries and specified in annotation data indicating the region type for the hyperspectral image 115 or composite image or standard color (e.g., RGB) image of the object 101.

[0074] In some implementations, the computer system 130 performs further processing on the segmented image, such as generating input feature values ​​from the pixel values ​​of specific segmented regions and providing the input feature values ​​to a machine learning model. For example, the machine learning model may be trained to classify objects such as strawberries based on characteristics such as size, shape, color, consistency of appearance, and absence of defects. The computer system 130 may use the region boundaries determined through segmentation to separate image data in various spectral bands from the hyperspectral image 115 that correspond to individual strawberries and / or specific types of regions of strawberries. Thus, the computer system 130 can provide an input image as input to the trained machine learning model that excludes background elements and other objects and instead provides only regions showing parts of a strawberry, or only regions showing specific parts of a strawberry (e.g., omitting the flesh, seeds, and leaves of the strawberry).

[0075] In some cases, the input provided to a machine learning model can be derived from segmented images without providing the image data itself to the model. Examples include the ratio of the number of pixels classified as one region type to the number of pixels classified as another region type (e.g., the ratio of the number of seed region pixels to the number of pulp region pixels), the average intensity of pixels segmented as a certain region type (e.g., strawberry pulp) (potentially for each of the various spectral bands), and the distribution of the intensity of pixels segmented as a certain region type. In some cases, the spectrum from each pixel in a region is preprocessed according to the techniques disclosed herein, and the preprocessed spectrum is analyzed to generate features. Exemplary features include the percentage of the spectrum in which baseline (or 0) crossings are observed in a particular frequency band, the mean integral of the convolution of the spectrum (or its derivative or second derivative) with a given kernel, the median number of peaks detected in the spectrum, the median number of 0 crossings detected in the first derivative of the spectrum, and the average ratio between the intensity in the spectrum in a first frequency band and the intensity in the spectrum in a second frequency band.

[0076] The machine learning model 170 can be trained to perform a variety of different functions, such as classifying the state of an object or estimating its properties. In the case of strawberries, the machine learning model can use segmented input image data to determine classifications (e.g., good condition, unripe, damaged, powdery mildew, etc.). The machine learning model can also be trained to provide scores or classifications for hyperspectral image-based predictions of specific properties such as chemical composition, intensity, defect type or density, texture, and color. The machine learning model 170 may be a neural network, classifier, decision tree, random forest model, support vector machine, or other type of model. The results of the machine learning model processing segmented hyperspectral image data or input features derived from hyperspectral image data may be stored in a database for later use and provided to any of various devices, for example, a client device for display to a user, a conveyor system for orienting object 101 to one of several locations, a tracking system, etc. For example, the results of a machine learning model that classifies objects 101 can be used to generate sorting equipment 180 (for example, causing the sorting equipment 180 to physically move or group objects according to the characteristics indicated by the machine learning model results), packaging equipment that specifies how and where to package objects 101, a robotic arm or other automated manipulator for moving or adjusting objects 101, or commands to be sent to operate objects 101.

[0077] The technique shown in Figure 1 can also be applied to evaluate other types of materials, such as plastics and other recyclable materials. For example, these techniques can be used to improve the efficiency and accuracy of characterizing the chemical or material identity of waste materials, making it possible to sort items by material type, presence of additives, presence of contaminants, and other properties determined through computer vision. This analysis can be used to improve both mechanical and chemical recycling processes.

[0078] Mechanical recycling is a viable strategy for recycling plastics, involving crushing, melting, and re-extruding plastic waste. Recycling facilities are often designed to handle streams of highly purified and sorted materials to maintain high levels of material performance in recycled products. However, impurities in the raw materials, including complex formulations with additives, as well as the physical decomposition of the material, reduce the effectiveness of recycling, even immediately after several cycles of mechanical recycling. For example, among plastic materials, polylactic acid (PLA) is a common waste plastic that is often not detected in polyethylene terephthalate (PET) sorting and mechanical recycling operations. Another example is chlorinated compounds such as polyvinyl chloride (PVC), which are unacceptable in both mechanical and chemical recycling operations because they generate corrosive compounds during the recycling process, which limit the value of the hydrocarbon output.

[0079] Mechanical recycling is limited in its applicability to mixed, complex, and contaminated waste flows, partly due to its use of mechanical separation and reforming processes that are insensitive to chemical contaminants and may not be able to alter the chemical structure of waste materials. System 130 can improve the effectiveness of mechanical recycling through improved identification and classification of plastics and other materials, resulting in more accurate sorting of materials and, therefore, higher purity and more valuable recycled raw materials. Furthermore, System 130 can use imaging data to detect the presence and type of additives and contaminants, enabling materials in which these compounds are present to be treated or removed differently.

[0080] Chemical recycling can overcome the limitations of mechanical recycling by breaking down the chemical bonds of waste materials into smaller molecules. For example, in the case of polymer materials, chemical recycling can provide a means of recovering oligomers, monomers, or even basic molecules from plastic waste raw materials. In the case of polymers, the chemical recycling process may include operations to depolymerize and dissociate the chemical composition of complex plastic products so that their by-products can be upcycled into raw materials for new materials. Elements of chemical recycling can enable the material to be repeatedly dissociated into primary raw materials. In this way, chemical recycling can be integrated into an "end-to-end" platform to facilitate the reuse of molecular components of recyclable materials, rather than being limited to a limited number of physical processes by chemical structure and material integrity, as in the case of mechanical recycling. For example, products of chemical recycling may include basic monomers (ethylene, acrylic acid, butyric acid, vinyl, etc.), raw material gases (carbon monoxide, methane, ethane, etc.), or elemental materials (sulfur, carbon, etc.). Instead of being limited to a single group of recycled products, products that can be synthesized from intermediate chemicals that can be produced from the waste by chemical reactions can be identified based on the molecular structure of the input waste material. In this way, the end-to-end platform can manage waste flows by generating chemical reaction schemes that convert waste materials into one or more target products. For example, the end-to-end platform could direct waste raw materials to a chemical recycling facility for the chemical conversion of waste materials into target products.

[0081] The capabilities of system 130 can also improve the effectiveness of chemical recycling. For example, system 130 can capture hyperspectral images of the waste flow on a conveyor belt and detect which materials are present using the technique shown in Figure 1. For example, camera system 110 may collect line scans along the width of the conveyor belt and then analyze them (e.g., by feeding features derived from pre-treated spectra into a machine learning model) to predict which chemicals are present in the objects passing on the conveyor belt. System 130 can also estimate the amount of different materials present, for example, based on the size and shape of different types of regions identified, as well as the proportion of different materials present. System 130 can also detect which additives and contaminants are present. From this information regarding the composition of the waste flow, system 130 can modify or update chemical processing parameters to change the target product quantity, endpoint, or chemical structure. Some of these parameters may include changes in processing conditions (e.g., residence time, reaction temperature, reaction pressure, or mixing rate and pattern) and the type and concentration of chemical agents used (e.g., input molecules, output molecules, catalysts, reagents, solvents). System 130 can store tables, formulas, models, or other data specifying processing parameters for different input types (e.g., different mixtures or conditions of input materials), and can use the stored data to determine instructions to the processing machine to implement the required processing conditions. In this way, System 130 can use analysis of hyperspectral imaging data of the waste flow to adjust the waste processing parameters so that the chemical recycling processing parameters match the material properties of the waste flow. Monitoring can be performed continuously so that if the mixture of materials in the waste flow changes, System 130 can appropriately change the processing parameters for the incoming mixture of materials.

[0082] As an example of applying the technology shown in Figure 1 to recycling, the camera system 110 can be configured to capture hyperspectral images of waste materials as objects 101 to be imaged. In some implementations, the camera system 110 is configured to capture hyperspectral images of waste materials on a conveyor. The waste materials may include many objects of different types and compositions imaged within a single image. The results of processing the images are used to sort the waste materials and generate instructions for equipment to process the waste materials mechanically or chemically, if optional.

[0083] During stage (A), as a setup process, the computer system 130 generates profiles of one or more types of materials of interest. Different profiles may be generated for different plastic types (e.g., PE, PET, PVC, etc.). The profiles may include spectra and / or spectral features. Information to facilitate the segmentation and detection of additives and contaminants may be included in the base material profile or other profiles. For example, in this application, profile 153 shown in Figure 1 may represent a profile for PET and may show a subset of spectral bands used when segmenting regions that coincide with PET. Profile 153 may also show different subsets of bands used to segment different types or variations of PET, or to segment regions having different types of additives or contaminants, respectively. Profile 153 may also, or alternatively, identify features such as bands corresponding to zero crossings in the derivatives of the normalized spectrum, functions of the convolution between the pre-processed spectrum (or the derivative or second derivative of the spectrum) and the kernel, and projections of the spectrum into a new vector space.

[0084] Using the techniques described above, the computer system 130 generates a profile of a material by processing various training examples 151, which include hyperspectral images of instances of the material to be profiled. The training examples may include examples showing the target material (e.g., PET) identified in the presence of various other different materials, including other waste materials such as other types of plastics. The training examples 151 may include at least several examples of images of the material having regions where additives or contaminants are present, so that the system 130 can learn which bands and / or features distinguish clean or pure regions of the material from regions where various additives or contaminants are present.

[0085] The Bandwidth Evaluation Module 150 performs a Bandwidth Selection process in which each of the Bandwidths and / or Features of the processed hyperspectral image is analyzed to generate a selected configuration that enables high accuracy in segmenting the desired material from other types of materials, in particular other types of plastics and other waste materials that may be imaged together with the material of interest. As described above, the Bandwidth Evaluation Module 150 can perform an iterative process of Bandwidth Selection 152 for each of several different material types and / or region types. For region segmentation of different base plastic types, the system 130 can use clustering techniques, component analysis (e.g., principal component analysis or independent component analysis), or support vector machines (SVM) to determine the variation or difference between pixel groups obtained from identifying materials using different Bandwidths. In some cases, the pixel intensities of each type of plastic can be clustered together, and the difference between clusters (e.g., between the mean values ​​of clusters) can be determined. In general, Bandwidth and / or Feature Selection Analysis can attempt to maximize the difference or margin between pixel groups of different materials.

[0086] For example, to distinguish between PE, PET, PVC, and PP, the band evaluation module 150 can evaluate the difference in pixel intensity (or the intensity of the derivative or second derivative of the untreated or pretreated spectrum) in different spectral bands and identify the band that provides the largest and most consistent amount of difference (e.g., margin) between the reflectance (or absorption) intensities for the different plastics. For illustrative purposes, the band evaluation module 150 may determine that a first band has similar average reflectance for each of the four plastic types described above, but a second band shows a larger amount of difference in reflectance for at least some of the plastic types. The analysis may be performed pairwise to identify which band is most effective in distinguishing which pair of materials. In various iterations, the band selection module 150 can select the band with the greatest discriminative power (e.g., the highest margin between pixel intensity groupings), and different combinations of these can be made to generate composite bands, which are similarly evaluated until the maximum number of iterations is reached or until the margins meet a minimum threshold of discriminative power. Similarly, this process can be used to determine the bands and synthetic bands that best distinguish plastics with additives or contaminants from pure base plastics.

[0087] As another example, to distinguish between PE, PET, PVC, and PP, the evaluation module 150 can identify one or more kernels that can be used to convolve the spectra (or pre-processed versions thereof). The convolution function can then be defined as features such that each spectrum is associated with a value for each of the one or more kernels. In some cases, this value may simply be a scalar projection of the spectrum onto another vector, or many such values ​​may be generated by a projection based on a vector space. In some cases, each of the one or more kernels corresponds to a chemical substance (e.g., including the spectrum of a chemical substance or a pre-processed version of the spectrum). In some cases, each of the one or more kernels corresponds to a higher level of shape or attribute.

[0088] As a result of band selection, the system 130 may determine, for example, that a first subset of bands or a first set of features provides identification of a desired plastic type (e.g., PET) from other plastic or waste materials (e.g., better identification than that provided by one or more other bands or one or more other features), and / or that a second different subset of bands or a second set of features is effective for segmenting specific additives or contaminants, such as oil as food residue (e.g., better segmentation than that provided by at least one other band or at least one other feature). This may include generating and evaluating a composite band that combines data from multiple different spectral bands of the original hyperspectral image. This results in a repository of profiles for different materials, each representing the best set of parameters and bands identified by the system to distinguish the target material from other materials it is likely to be near.

[0089] Continuing the application of the technology in Figure 1 to recycling applications, during stage (B), the camera system 110 captures hyperspectral images of the waste flow. For example, the camera system 110 takes hyperspectral images 115 of the waste flow on a conveyor in the process of sorting or other processing. During stage (C), the hyperspectral images 115 are transmitted from the camera system 110 to the computer system 130, for example, using the network 120. The hyperspectral images 115 may be transmitted in connection with a request to process the image, such as to generate segmented images, to identify materials represented in the image 115, to determine the amount of one or more specific materials represented by the image, to evaluate the level or type of contamination of the sample, or for other purposes.

[0090] During step (D), the computer system 130 retrieves profiles of different types of materials to be identified in the hyperspectral image data 115. Based on the information in the profiles, the computer system 130 may select one band, generate one or more composite bands, and / or identify one or more features of the image data to identify each different type of region to be detected. For example, the profile may specify different sets of bands for different plastic types and for different additives and contaminants. For example, one set of bands may be used to segment clean PET regions, another set of bands may be used to segment oil-contaminated PET regions, a third set of bands may be used to segment PET regions with specific additives, and so on. On the other hand, one or more bands, composite bands, or features may be used to predict whether the depicted object contains any of a set of chemicals (and / or parts of an object containing chemicals).

[0091] During stage (E), the segmentation module 160 performs segmentation based on the processed image data from the image processing module 140. In the case of plastic recycling, the segmented regions may be regions of different plastic types (e.g., PET, PE, PVC, etc.), as well as regions where additives or contamination are present (e.g., PE with oil contamination, PET with UV-resistant additives, etc.). The system 130 can perform further processing to interpret the segmented results, such as counting the number of different items or regions of each type, or determining the area covered by each region type (e.g., as an indicator of the amount and proportion of different materials).

[0092] During stage (F), system 130 stores the segmentation results 160 and other data characterizing the imaged region. System 130 can store metadata that marks specific objects or regions within the imaged region (e.g., waste material on a conveyor) along with the material type determined through the segmentation analysis. It can also store the boundaries of different regions, as well as the area or proportion of different region types. System 130 can then use this information to generate instructions for processing the waste material. For example, system 130 can provide information to a mechanical sorting machine to direct different plastic pieces to different bins, conveyors, or other devices according to the type of plastic detected. As another example, system 130 can identify plastic items that have one or more additives and separate them from plastics that do not have additives. Similarly, system 130 can identify items with at least a minimum amount of contamination (e.g., at least a minimum area where contaminants are present) and remove these items to avoid contamination of the rest of the recycled material. More generally, the system 130 can characterize various properties of waste material, such as assigning a score to any of the various properties of the detected material, using segmented hyperspectral image data. Based on these scores, or as a direct output of image analysis without potentially intermediate scores, the computer system 130 can classify parts of the waste material or classify a set of waste material as a whole.

[0093] As described above, the segmentation results 160, the results of applying segmentation to the hyperspectral image 115, and / or other information generated using them may be provided to other devices for display or further processing. The segmented image may also be used by computer system 130 or another system to generate input for one or more machine learning models. For example, computer system 130 may generate input feature values ​​from the pixel values ​​of a segmented region and provide the input feature values ​​to a machine learning model. For example, a machine learning model may be trained to classify whether an item or set of items is recyclable or not based on the type of plastic and the amount and type of additives and / or contamination present. Computer system 130 may use the region boundaries determined through segmentation to separate data from the hyperspectral image 115 for regions of different material types, different additives, or different contaminants, respectively. In some cases, information from the hyperspectral image data for a segmented region may indicate further material properties that may be relevant to the classification decision, such as the density, thickness, and quality of the material. Therefore, the computer system 130 can provide one or more input images as input to the trained machine learning model, excluding background elements (e.g., a conveyor belt) and other irrelevant objects, and instead providing only the regions relevant to the model's classification decision. The bandwidth of the information provided may differ from that used for segmentation. The spectral bandwidth that best distinguishes one material from another may be entirely different from the spectral bandwidth that exhibits the properties of that material, or the spectral bandwidth that distinguishes different states or changes in that material.Additionally or alternatively, one or more features disclosed herein (e.g., whether zero crossings of the first or second derivative of a spectrum are detected in a given frequency band, a function of the convolution of the spectrum (or its derivative or second derivative) and the kernel, a scalar projection of the spectrum onto a particular vector, the number of detected peaks in the spectrum, the number of detected crossings of a given baseline in the spectral derivative, etc.) may be used to characterize a material and / or determine subsequent actions for an object containing the material. In some cases, characterization is performed for each set of pixels in a segmented object depiction, and an aggregate of the characterizations (e.g., including mean, mode, median, percentage above threshold, range, standard deviation, etc.) is used to characterize the object and / or identify subsequent actions for the object. In some cases, the spectra of a set of pixels in a segmented object depiction are first aggregated (e.g., to calculate median, mean, mode, etc.), and then processed to generate characterizations for the object and / or identify subsequent actions for the object.

[0094] Segmented hyperspectral images (and / or features corresponding to hyperspectral image data) can be processed by trained machine learning models or other computational methods such as procedural or rule-based models to search for patterns in the signal that relate to material signatures, additive or contaminant signatures, or other information indicating chemical type, composition, morphology, structure, or purity. For materials incorporating multiple different additives, contaminants, or impurities into the main material, such as different forms of recycled PET objects containing various plasticizers often received by recycling facilities, data for multiple region types can be provided. Data for multiple bands of interest can also be provided, including image data for subsets of spectral bands that exclude spectral bands that are less useful because they have similar properties across many forms of recycled raw materials. As an example, data for specific bands of interest can be provided to a classifier implementing an SVM trained to classify materials, with different types of segmented regions marked or otherwise indicated.

[0095] In some cases, the input provided to a machine learning model can be derived from segmented images without providing the image data itself to the model. Examples include the ratio of the number of pixels classified as one region type to the number of pixels classified as another region type (e.g., the amount of pixels indicating plastic to the amount of pixels indicating non-plastic, the ratio of pixels representing PET to the amount of pixels representing PE, the amount of pixels representing clean PET to the amount of pixels representing contaminated PET), the average intensity of pixels segmented as a certain region type (potentially for each of various spectral bands), and the distribution of the intensity of pixels segmented as a certain region type. Any output of the machine learning model can be used by system 130 to control machines for sorting and processing waste materials. For example, this can be done by labeling items with metadata indicating their classification, or by generating instructions and sending them to sorting devices to operate specific items in a specified manner.

[0096] In some embodiments, the waste materials to be imaged and analyzed may include, but are not limited to, polymers, plastics, plastic-containing composites, non-plastics, lignocellulosic materials, metals, glass, and / or rare earth materials. Polymer materials and plastic materials may include materials formed by one or more polymerization processes, and may include highly crosslinked polymers and linear polymers. In some cases, the waste materials may include additives or contaminants. For example, plastic materials may include, for example, plasticizers, flame retardants, impact modifiers, rheology modifiers, or other additives contained in the waste material 111 to impart desired properties or promote formation properties. In some cases, the waste materials may incorporate constituent chemicals or elements that may not be compatible with a wide range of chemical recycling processes, and therefore the characterization data 113 may include information specific to such chemicals. For example, the decomposition of halogen or sulfur-containing polymers may produce corrosive byproducts, which may hinder or impair the chemical recycling of waste materials containing such elements. An example of a waste material containing halogen components is polyvinyl chloride (PVC). For example, the decomposition of PVC can produce chlorine-containing compounds that can act as corrosive byproducts.

[0097] Figure 2 is an exemplary diagram showing region type segmentation based on the band configuration specified by profile 153. Figure 2 provides additional detail following the example in Figure 1, where a strawberry is the object 101 being imaged, and profile 153 has already been defined to specify a subset of bands used to segment different region types within the strawberry. Here, the profile specifies three different composite or combined images to be created for use in segmentation, each labeled as components A, B, and C, each derived from image data of two or more bands of the hyperspectral image 115. Naturally, the profile is not required to specify an image that combines multiple bands; instead, in some cases, it may simply specify the original bands selected from the hyperspectral image 115 from which the image data should be passed to the segmentation module 160. As indicated by profile 153, the image processing module 140 generates various composite images 220a-220c, and these combined images are used by the segmentation module to generate masks 230a-230c that show the segmentation results (e.g., regions or boundaries of regions determined to correspond to different region types). The mask can then be applied to part or all of the image for different bands of the hyperspectral image.

[0098] System 100 utilizes the fact that regions with different compositions, structures, or other properties can have very different responses to light of different wavelengths. In other words, two different types of regions (e.g., seeds versus leaves) can each strongly reflect light in different bands. For example, the first band may be largely absorbed by the first region type (e.g., seeds) but reflected much more strongly by the second region type (e.g., leaves), making it a good band to use when segmenting the second region type (e.g., leaves). In this example, the different reflectivity characteristics captured in the image data showing the intensity values ​​captured for the first band of light tend to reduce or remove the first type of region at least partially, leaving a signal that more strongly corresponds to the second region type. Different bands can show the opposite if the first region type has a higher reflectivity than the second region type. However, in many cases, the level of differential reflectivity between two region types is not as clear as illustrated. In particular, different types of regions may not be effectively distinguishable based on only a single spectral band.

[0099] The bandwidth evaluation module 140 receives the hyperspectral image 115 as input and uses the profile 153 to determine the bandwidth configuration to use for each of three different region types. For example, the bandwidth evaluation module 140 generates three composite images 220a to 220c. The pulp and seeds are most prominently shown in composite image A, generated by (band 1 + band 3) / (band 1 - band 3), the strawberry seeds are most prominently shown in composite image B (band 1 + band 3), and the strawberry leaves are most prominently shown with the bandwidth configuration (band 1 - band 3) / (band 1 / band 5). In each of these, a reference to a bandwidth refers to image data for that bandwidth, and consequently, for example, "band 1 + band 3" represents the sum of image data for both bandwidth 1 and 3 (for example, summing the pixel intensity value for each pixel in the bandwidth 1 image with the corresponding pixel intensity value for the image for bandwidth 3). (It will be understood that each composite image may instead represent the value of a feature or a weighted sum of two or more features. Profile 153 can specify transformations and aggregations of image data for different bandwidths that best highlight different region types, and more specifically, it can highlight the differences between different region types to make boundaries clearer for segmentation.)

[0100] The segmentation module 160 also receives information from the profile 153, such as instructions on which bands (or features) or combinations thereof should be used to segment different region types, and what other parameters should be used for different region types (e.g., thresholds, which algorithms or models to use). The segmentation module 160 can perform the segmentation operation and determine masks 230a to 230c for each target region type. Each mask for a region type identifies the region corresponding to that region type (e.g., specifies pixels that are classified as depicting that region type). As seen in mask 230a, the segmentation process can remove seeds, leaves, and background, leaving regions where the strawberry flesh is clearly identified and shown. Next, masks 230a to 230c can each be applied to any or all of the images in the hyperspectral image 115 to generate segmented images, e.g., images 240a to 240c, which show the variation in intensity values ​​for each band, but restrict the data to image data corresponding to the desired region type.

[0101] As described above, the segmentation results can be used to evaluate object 101, such as determining whether shape, size, proportion, or other characteristics meet predetermined criteria. The segmented regions can also be provided for analysis, including by machine learning models, to determine other characteristics. For example, by isolating regions corresponding to a particular region type, the system can limit the analytical processes that should be performed using regions of that region type. For instance, an analysis of the chemical composition of a strawberry (e.g., Brix degree or sugar content in other units) can exclude pixels corresponding to seeds, leaves, or background that would distort the results if considered, based on a set of pixels identified as corresponding to the strawberry pulp. For each image data of one or more spectral bands of the hyperspectral image 115, information about pixels corresponding to the strawberry pulp region type is provided to a machine learning model that makes estimates about the strawberry's sugar content or other characteristics. Similarly, segmented data can be used to evaluate other characteristics such as ripeness, overall quality, and expected shelf life.

[0102] The spectral bands used for segmentation may differ from those used for subsequent analysis. For example, segmentation to distinguish between pulp and seeds may use image data in bands 1 and 2. The segmentation results may then be applied to image data in band 3, which shows chemical properties such as sugar content, and band 4, which shows water content. In general, this method allows for segmenting each region type using image data in the band(s) that most accurately distinguish the region type. Analysis of any property of a subject (e.g., presence or concentration of different chemicals, surface features, structural properties, texture, etc.) can benefit from its segmentation, regardless of the set of bands that best provides image data of the property to be evaluated.

[0103] Figure 3 is an illustrative process diagram showing an example of iteratively selecting different band configurations to specify a segmentation profile. As previously mentioned, a hyperspectral image consists of multiple 2D images, each representing the reflectance (or absorptive or absorbance) measured for a different wavelength band. In some cases, a hyperspectral image may include, or may be, a hypercube, which includes line scans for each of the sets of frequencies across each of the sets of positions in the 2D image. For simplicity, the example in Figure 3 uses a hyperspectral image 301 with image data for only three bands, band 1, band 2, and band 3, but many implementations use more bands. Furthermore, the example in Figure 3 shows an analysis for selecting the band configuration to use for a single region type of a single object type. The same process can be performed for each of multiple region types and for each of various different object types.

[0104] In some implementations, during the first iteration of the selection process, the system can evaluate image data for each individual band within the source hyperspectral image 301 to determine how well a desired region type can be segmented from that band. For example, each image 302-304 of the hyperspectral image 301 has segmentation applied to generate segmentation results 305-307. The scoring module compares the segmentation results based on the images of each band with the segmentation ground truth for the region and generates scores 312-314 representing the performance of the segmentation for a particular band. For example, band 1 image 302 of the hyperspectral image 301 is provided as input to the segmentation module 160. The segmentation module 160 performs segmentation based on the processing of band 1 image 302 and generates segmentation result 305. The scoring module 308 compares segmentation result 305 with ground truth segmentation 310 and generates a score 312 indicating 95% accuracy. The same process can be performed to evaluate each of the other bands in the hyperspectral image 301.

[0105] Next, the system compares scores 312–314 (e.g., segmentation accuracy scores) for each bandwidth and selects the one that exhibits the highest accuracy. For example, a predetermined number of bandwidths can be selected (e.g., n bandwidths with the highest scores), or a threshold can be applied (e.g., selecting bandwidths with accuracy above 80%). The selected bandwidths are then used in a second iteration of the selection process to evaluate potential combinations of bandwidths.

[0106] For the second selection iteration, the system generates new combinations of bands to test. This may involve combining different pairs of bands selected in the first iteration using different functions (e.g., sum, difference, product, quotient, maximum, minimum values ​​of images of different band pairs). This results in a composite or combined image that combines the selected bands with a specific function as a candidate in segmentation. For example, during iteration 1, bands 1 and 3 of hyperspectral image 301 were selected based on these bands having the highest precision scores 312 and 332. In the second iteration, the image data of these two bands are combined in different ways to generate a new composite or aggregated image. The selected bands 1 and 3 are combined to form three new images: (1) band 1 + band 3, (2) band 1 - band 3, and (3) band 1 / band 3. The images obtained from these three new band combinations constitute the image collection 351.

[0107] The second iteration performs the same steps as in iteration 1 on the images in image collection 351, for example, performing segmentation on each image, comparing the segmentation result with ground truth segmentation 310, generating a score for segmentation accuracy, comparing the scores, and finally selecting a subset of images 352-354 from image collection 351 that provides the best accuracy. For example, during iteration 2, each image in image collection 351 undergoes the same segmentation and selection process as described for iteration 1. For example, image 352 (formed by adding image data of band 1 and image data of band 3) is provided as input to segmentation module 160. Segmentation module 160 performs segmentation and generates a segmentation result 355 for this band combination (band 1 + band 3). Scoring module 308 compares the segmentation result 355 with ground truth 310 and generates a score 322 indicating an accuracy of 96%. For each of the other images generated using different operators used to combine data from bandwidths 1 and 3, the segmentation result and score are determined. The system then selects the bandwidth configuration(s) that provide the best segmentation accuracy.

[0108] The process can be continued for additional iterations as needed, for example, until the maximum number of iterations is reached, until the minimum level of accuracy is reached, until the candidate band combination reaches the maximum number of operations or bands, or until another condition is met, provided that the highest accuracy achieved in each iteration increases by at least a threshold amount. The highest-accuracy bands across various iterations are selected and added to the profiles of the region type and object type being evaluated. This may include specifying a subset of the original bands of the hyperspectral image 301 and / or a specific composite band, for example, a set of bands combined with a specific operator or function.

[0109] Figure 4 is a flowchart illustrating an example of process 400 for band selection and hyperspectral image segmentation. Process 400 is an iterative process, and for each region of interest, the process is repeated multiple times until the termination criteria are met. During each iteration, process 400 performs hyperspectral image segmentation, selects a set of bands based on the segmentation, combines bands within the set to generate new bands, and generates a new hyperspectral image with the new bands. In short, process 400 involves accessing hyperspectral image data containing multiple wavelength bands. Based on the hyperspectral image data, it generates image data for each of several different combinations of wavelength bands. It performs segmentation on each of the generated image datasets to obtain segmentation results for each of the several different combinations of wavelength bands. It determines a precision metric for each segmentation result for each of the several different combinations of wavelength bands. It selects one of the wavelength band combinations based on the precision metric. It provides an output showing the selected combination of wavelength bands.

[0110] More specifically, hyperspectral image data is acquired that includes multiple images having different wavelength bands (410). As previously mentioned, a hyperspectral image has three dimensions x, y, and z, where x and y represent spatial dimensions and z represents the number of spectral / wavelength bands. In one interpretation, a hyperspectral image includes multiple two-dimensional images, each two-dimensional image represented by the spatial dimensions x and y, and each two-dimensional image has a different spectral / wavelength band represented by z. Most hyperspectral images have hundreds or possibly thousands of bands, depending on the imaging technique. For example, camera system 110 takes a hyperspectral image 115 of a strawberry 101. The hyperspectral image 115 includes N images, each having a different wavelength band.

[0111] For each of the multiple region types, process 400 performs hyperspectral image segmentation and generates segmentation results (420). As previously mentioned, a hyperspectral image includes multiple images having different wavelength bands. Each image of the hyperspectral image having a particular wavelength band undergoes segmentation. For example, during iteration 1, image 301 includes three images 302, 303, and 304 of different wavelength bands provided as input to the segmentation module 160. Segmentation of each image at a particular wavelength generates segmentation results. For example, segmentation of image 302 generates segmentation result 305. Similarly, segmentation of images 303 and 304 generates segmentation results 306 and 307, respectively.

[0112] The segmentation results are compared to ground truth segmentation to generate performance scores (430). For example, scoring module 308 compares the segmentation result 305 of image 302 to ground truth segmentation to generate segmentation accuracy for image 302 having bandwidth 1. Similarly, scoring module 308 compares segmentation results 306 and 307 to segmentation ground truth 310 to generate accuracy 322 and 332.

[0113] The segmentation accuracy of the hyperspectral image in any given band is compared to a desired criterion (440). For example, if the user desires a 99% segmentation accuracy, but the highest segmentation accuracy in a band for a particular iteration is not 99% or higher, process 400 performs a selection process from the bands with the highest accuracy. However, if the highest segmentation accuracy in any of the bands meets the desired criterion, process 400 provides that band as output.

[0114] If the segmentation accuracy of the wavelength band does not meet the desired criteria, process 400 selects several different bands from among those having a specific performance score and uses these different bands to generate a new wavelength band (450). For example, the segmentation of three images 302, 303, and 304 with different bands in iteration 1 produces accuracy scores of 95%, 70%, and 98%, respectively, thereby failing to meet the desired criterion of 99%. Based on their high accuracy, process 450 selects bands 1 and 3 to generate new bands 352, 353, and 354. These three new bands form a new hyperspectral image 351 in iteration 2.

[0115] Figure 5A is a block diagram of an exemplary system 100 configured to predict the chemical composition of an object by processing one or more images of the object using a machine learning model. System 500 includes a computer system 130 that uses a machine learning model 170 to generate a prediction of the chemical composition of sample 101 based on a hyperspectral image 115 of sample 101. System 100 includes a camera system 110 for capturing images of the object being analyzed. The camera system 110 may be configured to acquire a hyperspectral image having, for example, pixel intensity values ​​for each of several different spectral bands. In some implementations, the camera system 110 may acquire data using other imaging or scanning techniques, including X-ray fluorescence and laser-induced breakdown spectroscopy. Results from spectroscopy or other scanning techniques may be used in addition to, or instead of, the hyperspectral imaging results for purposes such as segmentation, model training, and prediction of the amount and concentration of chemicals.

[0116] In the example in Figure 5A and several other examples below, system 100 uses a model trained to predict the sugar content of a specific type of fruit, such as a strawberry. The same technique can be used to train the model and predict the content of other chemicals in other types of objects. Generally, system 130 can use machine learning models to process images of a sample and infer sample properties that can only be tested directly through destructive analysis of the sample. This not only allows for testing without damaging the sample but also enables much faster testing and higher sampling rates for testing in many applications.

[0117] As will be further explained below with respect to Figure 5B, this technique can be used to predict the properties of plastics or other materials for recycling and waste management. For example, this technique can be used to train a model that uses information from hyperspectral image data to detect the presence of additives (e.g., phthalates, bromides, chlorates, surface coatings, etc.) and / or contaminants (e.g., oil or food residues on plastic items) in plastics and predict their concentrations. In some implementations, the same technique can be used to train a model to predict the type of base resin used (e.g., PE, PET, PVC, etc.) and other properties of the object. The information generated by the model can be used to characterize the flow of waste items or waste materials.

[0118] One way in which system 100 can provide high accuracy when predicting chemical components is by using hyperspectral images, which provide a much higher level of information than conventional RGB images. A hyperspectral image can be thought of as having three dimensions, with x and y representing the spatial dimensions of the pixel grid and a third z dimension representing the different hyperspectral bands or wavelengths from which the data is captured. In other words, a hyperspectral image can be thought of as containing a plurality of two-dimensional images, each representing the measured reflectance (or absorptiveness or absorbance) of different spectral bands or wavelength bands represented by z. For example, image 115 contains a plurality of images representing the spatial dimensions (e.g., image 1, image 2...image N), each of which has a different wavelength band (e.g., band 1, band 2...band N). The images may contain, or may contain, hypercubes.

[0119] Hyperspectral images often contain information about many different spectral bands (e.g., 5, 10, 15, 20, etc.). These bands may also be narrower than the RGB bands and may include more than three bands covering the spectral regions covered by RGB data. Hyperspectral images also often include data about non-visible spectral bands, e.g., bands in the ultraviolet and / or infrared ranges, potentially including one or more bands for each of the short-wavelength infrared (SWIR), mid-wavelength infrared (MWIR), and long-wavelength infrared (LWIR) ranges. Different chemical compounds have different light reflectance and absorption properties for different spectral bands. In other words, different chemicals interact strongly with different wavelength bands depending on their chemical structure. As a result, the ability to use hyperspectral images to assess the reflectance of a sample across many relatively narrow spectral bands can help system 130 identify distinct interactions that are characteristic of the chemicals under study. It can also help distinguish chemicals that have similar properties with respect to some spectral bands but may have different properties with respect to interactions with other spectral bands.

[0120] The set of spectral bands used in sample analysis may depend on the nature of the prediction task. For example, the system may identify a subset of spectral bands relevant to the analysis of a target chemical and use that subset for training and inference. As an example, to assess the sugar content of fruit, images of spectral bands(s) corresponding to the OH bond vibration frequencies in sugar molecules can predict the sugar content. However, a single band for this frequency may not provide the most reliable prediction, and a single band may not adequately distinguish the interaction between light and sugar from the interactions of other compounds that may be present (e.g., water, other carbohydrates). Consequently, when developing a machine learning model, system 500 can perform various steps to perform data-driven selection of spectral band combinations to be used in chemical component prediction. The process of testing which bands should be used in the model may include creating many different models trained to receive input feature values ​​(e.g., mean intensity features) for each of the different sets of spectral bands, and then testing the models to determine which set of spectral bands works best.

[0121] In some cases, if the properties of the chemical substance to be predicted are known, one or more spectral bands may be manually specified to include bands known to interact highly with the chemical substance. For example, to predict sugar concentration, since sugar molecules have bonds that interact with various spectral bands in the range of 100 nm to 1000 nm (and similarly or alternatively others), one or more bands covering this range may be included. Various bands can be used that correspond to different interaction peaks of the chemical compound to be predicted. These may include bands containing primary interactions, as well as secondary, tertiary, or other harmonics for stretching or vibration of chemical bonds. In some cases, the bands used correspond to different absorption peaks of the chemical substance over a certain range. Thus, the spectral bands used in the model for training and inference can be selected based on the chemical structure of the chemical substance to be predicted and / or the known interactions of the chemical substance with different wavelengths, and therefore the feature values ​​of those spectral bands directly show the effect of the chemical substance on the increase or decrease in reflectance (or absorptive or absorbance) in those bands. As mentioned above, other bands can also be selected to show the effect of other chemical substances that are expected to be present with the target chemical substance but are not actually predicted. For example, if both chemical substances A and B have similar interactions in a first band, but only chemical substance B strongly interacts with a second band (or more generally, the relative level or type of interaction differs for the two chemicals in different bands), then even if only the concentration of chemical substance A is predicted, features from both bands can be included, and therefore the model can learn to attribute or allocate interactions in the first band between different chemical substances.

[0122] In some cases, information in a spectral band directly indicates the presence of a chemical substance, for example, due to the reflectance (or absorptive rate or absorbance) of the chemical substance, or through a difference in reflectance due to the presence of the chemical substance. In other cases, feature data may provide indirect information about the presence of a chemical substance through reflectance data that indicates a chemical substance that is not predicted but, nevertheless, typically occurs together with the target chemical substance. For example, for a given application, two chemical substances A and B may typically be present together in a certain ratio, e.g., 60 / 40. Considering constraints on the application or chemical properties, it may not be possible to directly measure the reflectance in a spectral band that clearly indicates the concentration of chemical substance A that is desired to be predicted. Nevertheless, if there is a spectral band that can indicate the concentration level of chemical substance B, that band can be used in the model. Through machine learning training, the model can learn to predict the concentration of chemical substance A using image data that indicates the concentration of chemical substance B, based on the relationship between two chemical substances in the application. A data-driven or experimentally determined selection of which bands are most correlated and predicted (e.g., most effective when used in predictive modeling) can reveal these relationships and indicate which bands provide the most accurate and reliable results. This can arise from a frequency band showing a direct interaction between light and the predicted chemical substance, or from a frequency band showing interactions with other chemical substances related to the target chemical substance, and therefore can serve as a substitute or indicator of the concentration level of the target chemical substance.

[0123] The selection of spectral bands to be used may involve generating, evaluating, and incorporating extended bands that can combine information from different spectral bands. For example, many different extended bands may be created by adding images of two or more bands, subtracting an image of one band from another, or performing other operations. The system then determines which of the extended bands has the highest predictive value for inferring one or more properties of the subject. This allows the system to identify extended bands that are useful, for example, for separating the reflectance (or absorptive or absorbance) contribution of a chemical substance from the contributions of other chemicals. For example, both chemicals A and B may strongly interact with the first band, but only chemical B may strongly interact with the second band. For a model to predict the content of chemical A, system 500 can test many permutations of extended bands that add and subtract values ​​for different pairs of spectral bands, respectively. System 100 can determine that the extended band formed by subtracting the image data of the second band from the image data of the first band best indicates the content of chemical substance A (for example, because it is more useful in showing the interaction with chemical substance A than the first band alone). As a result, System 130 can use the identified extended band data as input to the model when training a machine learning model and when using the trained model to make inferences about the level of chemical substance A present in the sample.

[0124] To provide higher accuracy when predicting chemical components, system 130 may isolate specific types of regions of an image of a sample that are useful for chemical component prediction or for which chemical component prediction is to be made. System 130 can use isolated regions of a sample within an image, rather than the entire image or even the entire portion of the sample shown within the image, to generate input to a machine learning model. This can be done using an automated image segmentation process that identifies regions that meet certain criteria. Segmentation can identify different types of regions of a sample shown in an image using data for one or more spectral bands of a hyperspectral image. System 130 can then perform inference processes that treat different types of regions differently, for example, by using only specific types of regions in the image to infer the level of content of a particular chemical, weighting the values ​​of different regions differently, or using different regions to infer the level of content of different chemicals.

[0125] In some implementations, the camera system 110 may capture an image of a single side of the sample 101, or it may capture images of the sample 101 from multiple different angles, poses, or orientations. As described above, the camera system 110 may include a camera capable of capturing images of the fruit in light spectra other than the visible spectrum. For example, the camera system 110 may include a camera that takes images of the strawberry in specific bands of infrared and / or ultraviolet light. In the illustrated example, the camera system 110 captures a hyperspectral image 115 of the strawberry.

[0126] As will be further explained below, the intensity of reflectance in different spectral bands is used when making predictions about chemical composition. To provide high accuracy, the images captured and used in training, as well as the images used for inference processing, can be captured under controlled conditions. These may include capturing under a consistent level of artificial lighting, at a consistent distance from the light source, or capturing images in a housing that blocks ambient light. In this way, variations in reflectance (or absorptiveness or absorbance), indicated by the difference in pixel intensity values ​​across different spectral bands, can serve as a reliable indicator of differences in chemical composition.

[0127] The system includes a communication network 120. The network 120 may include a local area network (LAN), a wide area network (WAN), the internet, or a combination thereof. The network 120 may also include any type of wired and / or wireless network, satellite network, cable network, Wi-Fi network, mobile communication network (e.g., 3G, 4G, etc.), or any combination thereof. The network 120 may utilize communication protocols, including packet-based and / or datagram-based protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. The network 120 may further include several devices that facilitate network communication and / or form the hardware infrastructure for the network, such as switches, routers, gateways, access points, firewalls, base stations, repeaters, or a combination thereof.

[0128] The computer system 130 may include one or more computers that may be located locally or remotely with respect to the camera system 110 and the sample 101. The computer system 130 may be a server system that receives image data from a remote client along with a request to determine the chemical composition of a potentially specific chemical substance. The computer system 130 may process the image data using a trained machine learning model 170 to generate predictions of the amount or concentration of a specific chemical substance, and then provide data indicating the predictions to a device, for example, to be stored in a database, output to a user interface, or to control a sorting system or packaging system.

[0129] One application of the technology illustrated in Figure 5A is to use a computer system 130 to train and use a machine learning model 170 to predict (e.g., infer) the sugar content of fruit. The machine learning model 170 can be trained to infer the amount or concentration of a particular chemical or group or class of chemicals. Similarly, the model 170 can be trained to infer for a particular type of fruit, such as a strawberry. Different models can be trained to predict the content of different chemicals in different sample types (e.g., different types of fruit). In Figure 1, the model 170 is used to infer the sugar content in the juice of a strawberry based on an external hyperspectral image of the entire strawberry. The camera system 110 captures a hyperspectral image 115 of sample 101 and transmits the image data 115 to the computer system 130 using the network 120. Upon receiving the image captured by the camera system 110, the computer system 130 processes the image 115 using the trained machine learning model 170 to predict the sugar content of the fruit in the image as Brix degree values.

[0130] In some implementations, the components of system 500 may be geographically distributed. In such implementations, the computer system 130 may be implemented by a single remote server or by a group of multiple different servers distributed locally or globally. In such implementations, the functions performed by computer system 130 may be performed by multiple distributed computer systems, and the predictive model may be provided as a software service via network 120. In some implementations, system 500 may be implemented locally. For example, system 500 may be implemented as a standalone sensor unit housing camera system 110 and computer system 130.

[0131] Model training may be performed on a remote system, but the trained model 170 may be delivered to different client devices via the network 120, and the client devices may process locally captured images using the locally stored model 170. In such an implementation, a user can point the camera of their user device and take an image of a specific type of object. In such an implementation, the machine learning model 170 may be stored on and used by the user device to generate chemical component inferences.

[0132] Machine learning models can also be provided as software services accessed by user devices. For example, a machine learning model may be implemented on a remote server. When a user takes an image of fruit in a scene using a user device such as a smartphone, the user can transmit the image over network 120 by uploading it to a remote server. Upon receiving the image, the remote server processes the image to predict the sugar content of the fruit in the image. The predicted sugar content is then transmitted to the user device over network 120 for display or notification to the user. Such an implementation can be based on a third-party application running on the user device's operating system, or it can be implemented by accessing a remote server website or portal connected via network 120 or the internet.

[0133] Figure 5A further illustrates an exemplary data flow shown in stages (A) to (F). Stages (A) to (F) may occur in the order shown, or in a different order. In some implementations, one or more of stages (A) to (F) may occur offline, and the computer system 130 may perform calculations when the user device is not connected to the network 120. Stages (A) to (C) describe training the machine learning model 170, and stages (D) to (F) describe using the trained model 170.

[0134] During stage (A), the computer system 130 acquires a set of training images 540. The set of training images 540 includes multiple images of a particular type of fruit. For example, to predict the sugar content of strawberries, the set of training images 540 includes multiple images of different strawberries, including examples of strawberries with different levels of sugar content. Each image in the set of training images 540 is associated with a ground truth label 541 that indicates the actual sugar content of the strawberry, determined through sugar content measurement. The ground truth label 541 of the fruit can be obtained by first capturing an image of the fruit and then performing invasive and / or destructive chemical tests to determine the sugar content of the fruit. Typically, the measurement is performed by destructive testing, such as crushing the strawberry after the image has been captured and measuring the sugar content of the released juice in Brix degrees using, for example, a refractometer or hydrometer.

[0135] In some implementations, the set of training images 540 may comprise two or more images that could correspond to the same particular fruit, each of which may be captured by a different light spectrum, angle, distance, pose, or camera system. For example, the set of training images 540 may include two images, the first showing one side of a strawberry and the second showing the other side of the same strawberry. In such an exemplary scenario, both images would have the same ground truth sugar content because the images refer to the same fruit.

[0136] To enable robust model training, the training images 540 may include images of strawberries taken with different camera systems, different lighting levels, different distances from the camera to the subject, different orientations or poses of the sample, etc.

[0137] During stage (B), the computer system 130 uses the feature extraction module 550 to process the image and generate input features to provide to the machine learning model 170. (It will be understood that, in some cases, the feature extraction module 550 and the machine learning model 170b may operate as a single, identical model. For example, one feature may be the image itself.) Image processing may include several steps. The system 130 segments the training image into different types of regions. (In some cases, it will be understood that segmentation is performed such that the segmentation profile 153 may be a null operation. Additionally or alternatively, image processing such as segmentation may be performed as part of the pipeline within the machine learning framework. Different types may correspond to different parts of the type of object that the model 170 is being trained to analyze. For example, different region types may be determined for the strawberry flesh, strawberry seeds, and strawberry leaves or calyx. Another type of region may refer to background pixels that represent the environment or surroundings of sample 101, rather than sample 101 itself. In this example, the images used for training are segmented, and the regions corresponding to the strawberry flesh are used to determine feature values, while other regions are discarded and not used for sugar content analysis. For example, the outer surface of a strawberry generally includes both the flesh and seeds. In such an example, the feature extraction module 550 may perform image segmentation by processing images in the set of training images 540 to separate the image regions representing the flesh portion of the fruit from the seeds.

[0138] The feature extraction module 550 can also determine values ​​for various different features. These features may be the average pixel intensity in a selected segmented region of the training image for each of the different spectral bands. For example, it can provide the average intensity for band 1, the average intensity for band 2, and so on. The average intensity values ​​for a selected set of spectral bands can be used as input vectors for input to the machine learning model 170 during training.

[0139] As an alternative to segmentation for separating parts of an object used for chemical prediction from parts not used, system 130 can instead use sampling techniques. For example, in a strawberry, regions with seeds and regions without seeds have very different reflectance curves across the hyperspectral band. In particular, seeds are generally more reflective for most, if not all, spectral bands. To take advantage of this property, feature extraction module 150 can take small samples of image data across regions of an image showing a strawberry and determine the reflectance curve or reflectance mean in each of the small sampled regions. Sampling within the region showing a strawberry may be done randomly or pseudo-randomly in a regular or grid-like pattern, or using sampling that covers the entire region of the strawberry. The reflectance results for each sample region are clustered together into clusters, whether as curves, mean vectors across bands, etc. In this application, there are two main clusters, one with a generally higher reflectance representing seed inclusion, and the other with a lower reflectance where the sample avoids seeds. Next, values ​​within clusters with lower reflectivity can be selected as the group to be used to generate features for training, and the same technique can be used to determine features for inference processing. This technique uses sampling to separate the contribution of a specific type of target region (e.g., strawberry pulp) from the contributions of other regions (e.g., showing seeds) that would otherwise cause inaccuracies in predictions.

[0140] In some implementations, the computer system 130 can perform an analysis to determine which spectral bands are most effective in predicting the type of characteristic that the model 170 is being trained to predict. For example, before training in step (B), the computer system 130 can perform an analysis process to examine the predicted values ​​of different spectral bands individually or collectively to determine a combination of bands that provides the greatest predictive accuracy for predicting the chemical components of a particular chemical from a particular type of sample (e.g., a type of object). This often allows for dimensionality reduction and smaller models while improving accuracy. Generally, hyperspectral images have dozens of bands depending on the imaging technique used to capture the image, and not all bands provide useful information for predicting the chemical components of a desired chemical. Based on correlations or predicted values ​​identified through the analysis of exemplary data, the system can select a subset of bands, including potentially expanded (and synthesized) bands, that combine data from different bands together to provide information about the desired type of chemical component to be predicted.

[0141] During stage (C), the computer system 130 trains a machine learning model 170 using feature values ​​determined from training images. In some implementations, model 170 is a decision tree, such as a gradient-boosted regression tree (e.g., an XG-boosted tree). In other implementations, model 170 may be a neural network, a support vector regression model, or another type of model. The machine learning model 170 includes several parameters, the values ​​of which are adjusted during training. Training a machine learning model involves adjusting the trainable parameters of the machine learning model so that the machine learning model 170 can predict the level of content of one or more chemicals by processing a set of input feature values ​​determined from selected segmented regions of the training images, e.g., the average intensity values ​​for each of a given set of spectral bands. When a gradient-boosted decision tree is used as the model type, a gradient-boosted training algorithm may be used. In other implementations, a neural network may be used, and backpropagation and other neural network training techniques may be used.

[0142] Training can be carried out using different training examples until Model 170 can predict the desired characteristic(s), such as the amount or concentration of a compound from the input. When analyzing the sugar content of strawberries, with an appropriate set of training data and sufficient iterations, Model 170 can be trained to predict the sugar content in Brix degrees with an accuracy that matches or exceeds the level provided by a typical destructive test.

[0143] After the model is trained, model 170 can be used to make predictions about objects based on images of those objects. During stage (D), the camera system 110 captures an image 115 of sample 101, which is a strawberry, the type of object for which model 170 was trained to make chemical component predictions. The camera system 110 can capture a hyperspectral image 115 that includes image data of spectral bands outside the visible spectrum, in addition to or instead of spectral bands within the visible spectrum. For example, the camera system 110 captures a hyperspectral image 115 of sample 101. Image 115 includes multiple images (e.g., image 1, image 2...image N) representing the spatial dimension, each image having a different wavelength and / or band (e.g., band 1, band 2...band N). The camera system 130 provides the image data of image 115 to the computer system 130 via the network 120.

[0144] During stage (E), the computer system 130 receives the image 115 via the network 120 and processes the data using the machine learning model 170 to determine the predicted level of the chemical components. In this example, upon receiving the image 115, the computer system 130 processes the image 115 to predict the sugar content of the strawberries shown in the image 115.

[0145] The feature extraction module 550 processes image 115, including segmenting the image to identify portions of image 115 that represent sample 101 (e.g., the strawberry, not the background), more specifically, portions of image 115 that show a particular type of region of the strawberry (e.g., the strawberry flesh, as opposed to leaves, seeds, etc.). Using a selected subset of the segmented regions, e.g., the regions of image 115 identified as the strawberry flesh, module 550 determines values ​​for each of a given set of features. Features can correspond to different spectral bands selected for use in predicting one or more chemical properties of a subject. For example, a feature could be the average intensity for each of the subsets of spectral bands in image 115. As a result, a set of feature values ​​is determined for image 115, each feature value representing the average intensity value for different (which may be extended) spectral bands in a given set of spectral bands, and the average is determined over the segmented regions identified as the strawberry flesh. The given set of spectral bands can be the same set of spectral bands for which information was provided to model 170 during training.

[0146] The feature values ​​generated by the feature extraction module 550 are provided as input to the trained machine learning model 170. The machine learning model 170 processes the input feature values ​​using the training state determined during training, for example, using the parameter values ​​set in step (C) described above. The model 170 produces an output 570 that indicates the level of prediction (e.g., inference) of the content of one or more chemicals that the model 170 has been trained to predict. For example, the model 170 may provide a regression output (e.g., a numerical value) indicating the sugar content of the strawberries shown in image 115. The output 570 can be expressed in any appropriate form or type of unit, such as Brix degree, plateau degree, specific gravity of sugar, mass fraction of sugar, or concentration of sugar. The type of measurement used for prediction may be the same type used for the ground truth label 541 used to train the model 170.

[0147] During stage (F), data indicating the predicted chemical composition 570 is provided to the device. For example, the prediction 570 may be stored in a database 590 or other data storage, associated with a sample identifier, timestamp, and potentially contextual data such as sample type and location. As another example, the prediction 570 may be provided to a user device 585, such as a smartphone, laptop computer, or desktop computer, for presentation to the user in a user interface.

[0148] In some implementations, the predicted chemical composition is evaluated by a computer system 130 or another device. For example, the chemical composition prediction 570 can be compared against one or more thresholds, and the user may be given a warning if the chemical composition is above the maximum level, below the minimum level, within the range, outside the range, or meets other predetermined conditions. The prediction 570 can be used by the computer system 130 or another system to control sorting or packaging equipment 575 configured to move strawberries with different sugar levels to different areas or containers. The prediction 570 can additionally or alternatively be used for sample 101 or the source of sample 101 (e.g., a lot or batch containing sample 101, a specific plant, plant community, or field where sample 101 is located or harvested, etc.). This function can be used to estimate or optimize the timing for finely harvesting strawberries or other fruits to achieve desired levels of chemical composition (e.g., sugar concentration or other properties).

[0149] The technique described above is illustrated by an example of predicting the sugar content of sample 101 in Brix degrees, but the same technique can be used to predict the concentrations of other chemicals. More generally, this technique can be used to predict other properties such as texture, hardness or softness, and time to the predicted ideal harvest time. These and other properties can be predicted by training a model using ground truth values ​​for the type of property to be predicted. These other properties may, in some cases (for example, time to harvest may be based in part on sugar concentration), be related to or correlated with chemical components, or not.

[0150] Figure 5B is a block diagram showing another example of system 100 predicting chemical components based on one or more images using a machine learning model. System 100 can perform image data segmentation and predict chemical components using a machine learning model. Both the segmentation process and the chemical component prediction process can be adjusted by selecting a subset of available bandwidths of information from hyperspectral images and / or by defining selective features to characterize objects and / or materials depicted within the hyperspectral images. As a result, system 100 can segment or identify regions in hyperspectral images corresponding to different objects or materials using information from different bandwidths and / or different types of information. Similarly, the system can use information from different bandwidths as input to estimate different chemical properties or properties of different types of objects. This selective use of information from hyperspectral images can provide higher accuracy in prediction, as well as faster and less computationally intensive models.

[0151] An example in Figure 5B shows predictive properties of plastics or other materials for recycling and waste management. For example, this technique can be used to train a model that uses information from hyperspectral image data to detect the presence of additives (e.g., phthalates, bromides, chlorates, surface coatings, etc.) and contaminants (e.g., oil or food residues on plastic items) in plastics and predict their concentrations. In some implementations, the same technique can be used to train a model to predict the type of base resin used (e.g., PE, PET, PVC, etc.) and predict other properties of the object. The information generated by the model can be used to characterize the flow of waste items or waste materials.

[0152] System 100 can perform optical sorting using hyperspectral imaging to separate recyclable plastics from other materials and sort plastics by resin type, level or type of contamination, and / or presence or concentration of additives. System 100 includes a conveyor 505 that transports a waste stream 101b containing various plastic items. Camera system 110 captures hyperspectral images 115b of items in the waste stream 101b while they are on the conveyor 505. The hyperspectral images may be generated using an optical system including one or more light sources and one or more cameras (and / or one or more optical sensors). The light source(s) may be configured to emit (e.g.) infrared light, near-infrared light, shortwave light, coherent light, etc. In some cases, the light source(s) may include (e.g.) optical fibers that can isolate the light source output from thermal output. In some cases, at least one camera(s) is positioned such that the optical axis of the camera lens or image sensor is at an angle of 75–105 degrees, 80–90 degrees, 85–95 degrees, 87.5–92.5 degrees, 30–60 degrees, 35–55 degrees, 40–50 degrees, 42.5–47.5 degrees, or less than 15 degrees relative to the surface (e.g., a conveyor belt) supporting the object(s) being imaged. In some cases, the optical system includes multiple cameras, and the angle between the optical axes of the first camera and the surface supporting the object(s) is different from the angle between the optical axes of the second camera and the surface. The difference may be (e.g.) at least 5 degrees, at least 10 degrees, at least 15 degrees, at least 20 degrees, at least 30 degrees, less than 30 degrees, less than 20 degrees, less than 15 degrees, and / or less than 10 degrees. In some cases, the first camera filters different types of light relative to the second camera. For example, the first camera may be an infrared camera and the second camera may be a visible light camera. In some cases, the camera may be tilted at a first angle with respect to the surface supporting the object(s) being imaged, and the light source may be tilted at a second angle with respect to the surface. The first angle may be approximately the opposite of the second angle (for example). For example, the camera may be tilted 5 degrees with respect to the normal vector of the belt, and the optical axis of the illumination may be tilted -5 degrees with respect to the belt normal.

[0153] Based on the processing of the hyperspectral image 115b, the system 100 sorts items in the waste stream 101b into different bins 515a to 515c. For example, a sorting machine 510 can operate based on the processing of the hyperspectral image 115b to sort plastics into bins such as clean PET items (bin 515a), dirty or contaminated PET items (bin 515b), and PET items with additives exceeding a predetermined concentration (bin 515c). Naturally, instead of storing items in bins, the sorted output stream may, optionally, be sent on different conveyors to be processed directly by chemical or mechanical recycling. As the waste stream 101b is transported, hyperspectral images of different parts of the flow are taken and processed so that individual items (e.g., bottles, containers, etc.) can each have their analyzed composition, and individual items can be sorted based on the resins, contaminants, and additives determined to be present, as well as estimated amounts and concentrations of potentially target chemicals.

[0154] In some implementations, the camera system 110 may use other imaging or scanning techniques, such as X-ray fluorescence and laser-induced breakdown spectroscopy, in addition to or instead of hyperspectral imaging. The camera system 110 may include scanning equipment positioned to acquire scans of objects in the waste flow 101b on the conveyor 505. The conveyor 505 may be configured to carry objects to a position where scanning is to be performed, and then periodically stop so that scans of different items or groups of items on the conveyor 505 are acquired. Scanning for hyperspectral imaging and spectroscopy techniques can be performed from the same scanning location on the conveyor 505 or from different locations on the conveyor 505.

[0155] In plastics, several chemical properties of interest are uniformly distributed throughout the material. For example, the resin type and additive concentration are typically uniform across the entire plastic item. However, contaminants, labels, blockages, and other elements can obscure or interfere with hyperspectral measurements of parts of a plastic item. Similarly, some elements, such as lids, may be made of different materials or compositions than the corresponding container. Segmentation of hyperspectral image data can be performed to separate plastic items from non-plastic items and to separate individual plastic items from each other. Furthermore, segmentation can be performed to identify contaminated areas of an item from uncontaminated areas. This makes it possible to isolate image data of contaminated areas and use it to provide input to a model that predicts the type and concentration of contaminants. Areas identified as uncontaminated can then be used to provide input to a model that predicts the resin type and additive properties. As a result, by using different segmented areas for different models, system 100 can achieve higher accuracy than evaluating the properties of an object based on the entire area of ​​image data representing the object.

[0156] Figure 5B further illustrates the exemplary data flow shown in steps (A) to (F). Steps (A) to (F) may occur in the order shown, or may occur in a different order. Steps (A) to (C) describe the training of one or more machine learning models 170b to predict chemical properties. Steps (D) to (F) describe the use of one or more trained models 170b, including performing segmentation to determine which regions of the hyperspectral image 115b are relevant to the parameters predicted by the model 170b.

[0157] During stage (A), the computer system 130 acquires a set of training images 540b. The set of training images 540b includes images of the plastic or other recyclable item to be analyzed. For example, to predict the resin type, additives, and contaminants of a plastic item, the set of training images 540b may include multiple hyperspectral images of different types of plastics, as well as exemplary hyperspectral images of plastics with different additives and different types of contaminants. Examples with additives and contaminants may include examples with different concentrations and different combinations, as well as examples for each of several different base resin types. Each hyperspectral image in the set of training images 540b may be associated with a ground truth label 541b that indicates the actual characteristics of the imaged item. For example, the ground truth label 541b may include information about any of the characteristics to be trained to predict by the model 170b, such as the resin type, whether different additives or contaminants are present, and the amount or concentration of each of the different additives or contaminants. In some implementations, the ground truth label 541b may indicate the mass fraction or other measure of the present resin, additives, and contaminants. The information in the ground truth label 541b can be determined through testing of the imaged sample or by examining the material properties from manufacturing data or reference data relating to the type of imaged object. To enable training of a robust model, the training images 540 may include images of a given type of fruit, waste, etc., taken with different camera systems, different lighting levels, different camera distances to the subject, different orientations or poses of the sample, etc.

[0158] During stage (B), the computer system 130 uses the feature extraction module 550 to process the hyperspectral image and generate input features to provide to the machine learning model 170b. Processing the hyperspectral image may include several steps. The system 130 segments the training image into different types of regions. The different types can correspond to different parts of the type of object that the model 170b is being trained to analyze. For example, different region types can be determined for regions of different base resin types, or for contaminated and uncontaminated regions. Another type of region may refer to background pixels that represent the environment or surroundings of sample 101b, rather than sample 101b itself. In this example, the image used for training is segmented, and the regions corresponding to the uncontaminated areas are used to determine feature values ​​for predicting resin content and additive properties, while the contaminated areas are used to generate feature values ​​for predicting the type and concentration of contaminants.

[0159] The feature extraction module 550 can also determine values ​​for various different features. These features may include the average pixel intensity in a selected segmented region of the training image for each of the different spectral bands. For example, it may provide the average intensity for band 1, the average intensity for band 2, and so on. The average intensity values ​​for a selected set of spectral bands may be used as input vectors for input to the machine learning model 170b during training. Other exemplary features disclosed herein include whether zero crossing or baseline crossing is detected in the spectral derivative or second derivative (e.g., the derivative is taken over frequency) in a given frequency band, the number of peaks detected, and the integral of the convolution (or spectral derivative or second derivative) of the spectrum and kernel.

[0160] Different models 170b can be generated and trained to predict different properties. For example, five different classifiers may be trained for five different base resins, each trained to predict the likelihood that an image region represents the corresponding base resin. Similarly, ten different models can be trained so that each predicts the concentration of one of ten different additives. As another example, three different contaminant models can be generated to characterize the contents of three different contaminants, one for each type of contaminant in question. In some cases, the models may be further specialized, such as a first set of models trained to predict the concentrations of different additives in a PET sample, and a second set of models trained to predict the concentrations of different additives in a PE sample. Each of the different chemicals detected and measured may have a different response to various spectral bands of the hyperspectral image. As a result, different models 170b for predicting different chemical properties can be configured and trained to use different subsets of spectral bands of the hyperspectral image as input data. The bands used for each model 170b can be selected to best identify the property in question, and in some cases, to best distinguish it from the presence of other frequently present materials.

[0161] In some implementations, the computer system 130 can perform analyses to predict which spectral bands and / or features are most effective, sufficiently effective, or relatively effective in predicting the type of characteristics that each model 170b is trained to predict. For example, prior to training in step (B), the computer system 130 can perform an analysis process to individually or collectively examine the predicted values ​​of different spectral bands to determine a combination of bands that provides the greatest predictive accuracy for predicting the chemical composition of a particular chemical substance (e.g., a particular resin, additive, or contaminant). This often enables dimensionality reduction and smaller models while improving accuracy. Generally, hyperspectral images have dozens of bands depending on the imaging technique used to capture the image, and not all bands provide useful information for predicting the chemical composition of a desired chemical substance. Based on correlations or predicted values ​​identified through the analysis of exemplary data, the system can select a subset of bands, potentially including extended or composite bands, that combine data from different bands to provide information about a desired type of chemical composition to be predicted.

[0162] The bands or combinations of bands that facilitate high accuracy for each chemical property to be predicted can be identified by System 130 through a process similar to that described herein (e.g., Figure 3) for determining which bands to use to segment different types of regions. For example, different models (e.g., sets of regression trees, support vector machines, etc.) can be trained on data from different individual bands, and the performance of different models can be tested and compared. The model with the highest accuracy is identified, and then a new model is trained in another iteration, the new model using a different pair of bands that yielded the most accurate model in the first set. For example, if there are 20 spectral bands in a hyperspectral image, the first iteration can train 20 models to predict the concentration of the same plastic additive, each using data from a single spectral band among the spectral bands. The bands used by the highest-accuracy subset of models (e.g., the top n models where n is an integer, or a set of models that provide accuracy above a threshold) are then identified. If five bandwidths are selected as most relevant to the plastic additive, the next evaluation iteration can train another 10 models to predict the concentrations of the same plastic additive, each model using a different pair of the five selected bandwidths from the previous iteration. Optionally, the five bandwidths can be combined in different ways to generate extended or composite bandwidths that can be used to additionally or alternatively train and test models. Iterations can continue until at least a minimum level of accuracy is achieved, or until performance no longer increases beyond a threshold amount. This same overall technique can be used to evaluate which other different spectroscopic techniques (e.g., X-ray fluorescence, laser-induced breakdown spectroscopy, etc.) or which parts of the results of these techniques best predict different chemical properties.It will be understood that this type of iterative technique may, instead or in addition, be used to select features (for example, from among multiple types of features and / or features corresponding to multiple frequency bands) to be used to predict which chemicals are present in a depicted object, its chemical properties, etc.

[0163] Other techniques can be used to select subsets of bandwidths used to train different models 170b. For example, the computer system 130 can access data describing reference hyperspectral data showing the response to a pure sample or a high-concentration sample, and the system can identify bandwidths with high and low intensities that indicate the bandwidths most affected and least affected by the presence of plastic additives, respectively. As another example, hyperspectral images for different known concentrations of a chemical can be compared to identify which bandwidths change the most as a change in concentration, and thus which bandwidths are most sensitive to changes in concentration.

[0164] During stage (C), the computer system 130 trains a machine learning model 170b using feature values ​​determined from training images. In some implementations, model 170b is a decision tree, such as a gradient-boosted regression tree (e.g., an XG-boosted tree). In other implementations, each model 170b may be a neural network, a support vector regression model, or another type of model. Each machine learning model 170b includes several parameters, the values ​​of which are tuned during training. Training a machine learning model involves tuning the trainable parameters of the machine learning model so that the machine learning model 170b can predict the level of content of one or more chemicals by processing a set of input feature values, e.g., the mean intensity value of each of a given set of spectral bands selected for that model 170b. The input values ​​can be determined from selected segmented regions of the training images related to the chemicals being characterized (e.g., regions segmented as contaminated, used to generate input for the pollutant characterization model 170b). When a gradient-boosted decision tree is used as the model type, a gradient-boosted training algorithm may be used. Other implementations may utilize neural networks, and backpropagation and other neural network training techniques may be employed.

[0165] Training can be carried out using different training examples until Model 170b can predict the desired property(s), such as the amount or concentration of a compound. When analyzing plastics, with an appropriate set of training data and sufficient iterations, Model 170b can be trained to predict the mass fraction of each of different additives with an accuracy that matches or exceeds the level provided by typical destructive testing. Similarly, Model 170 can be trained to predict the type of resin present, whether different contaminants are present, and the amount or concentration present.

[0166] After model 170b is trained, model 170b can be used to make predictions about objects based on images of objects. During stage (D), the camera system 110 captures a hyperspectral image 115b of sample 101b, which is an image of an area 520 on the conveyor or 505 showing one or more objects to be sorted. The hyperspectral image 115b includes image data of spectral bands outside the visible spectrum, in addition to or instead of spectral bands within the visible spectrum. For example, the camera system 110 captures a hyperspectral image 115b of sample 101b. Image 115b includes multiple images (e.g., image 1, image 2...image N) representing the spatial dimension, each image having a different wavelength and / or band (e.g., band 1, band 2...band N). Where appropriate, the camera system 110 can also acquire scanning results for X-ray fluorescence spectroscopy, laser-induced breakdown spectroscopy, Raman spectroscopy, or other spectroscopic techniques, which can provide additional bands of data about objects on the conveyor 505. The camera system 130 provides the image data of image 115b, as well as any other relevant spectral results, to the computer system 130 via the network 120.

[0167] During stage (E), the computer system 130 receives image 115b via the network 120 (or directly via wired or wireless connection) and processes the data using a machine learning model 170b to determine the chemical properties of one or more objects in region 520. This may include detecting the presence of different chemicals and estimating the levels of chemical components present for each of several different chemicals. If region 520 represents different items (e.g., different bottles in this example), each item may be processed individually to characterize the chemical components of each item and to properly sort each item. In this example, upon receiving image 115b, the computer system 130 processes image 115b to predict, for each plastic object identified from image 115b, the resin type, present additives, and present contaminants, as well as the concentration of each of these compounds.

[0168] The feature extraction module 550 processes image 115b and includes segmenting the image to identify portions of image 115b that represent different objects within sample 101b (e.g., different bottles rather than the background), more specifically, to identify portions of image 115b that show a particular type of region (e.g., PET vs. PE regions, contaminated regions vs. clean regions, etc.). Using a selected subset of the segmented regions, module 550 determines values ​​for each of a given set of features. The features may correspond to different spectral bands selected for use in predicting one or more chemical properties of a subject. In other words, data from different combinations of bands can be used to provide input to different models 170b. Furthermore, data for those bands may be obtained from specific segmented regions relevant to a model, for example, using segmented regions of contamination solely to generate input features for a model trained to predict contaminant concentrations. For example, features may be the average intensity for each of the subsets of spectral bands in image 115b. If data is obtained for other spectroscopic techniques, feature values ​​for these spectroscopic results are also generated and provided as input to model 170b to generate estimates of chemical composition. As a result, a set of feature values ​​is determined for image 115b, where each feature value represents the average intensity value for different spectral bands (which may be extended bands) within a given set of spectral bands, and the average is determined over segmented regions identified as region types corresponding to the model. The given set of spectral bands may be the same set of spectral bands for which each model 170b was informed during training.

[0169] The feature values ​​generated by the feature extraction module 550 are provided as input to the trained machine learning model 170b. As described above, they may be a different set of feature values ​​for each model 170b, determined based on a subset of bands selected as most relevant to or most effective in predicting the properties that model 170b predicts. The machine learning model 170b processes the input feature values ​​using the training state determined during training, for example, using the parameter values ​​set in step (C) above. Each model 170b can produce an output 570b that indicates the level of prediction (e.g., inference or estimation) of the content of one or more chemicals that model 170b has been trained to predict. For example, different models 170b may each provide a regression output (e.g., numerical) indicating the mass fraction of different additives. The type of measurement used for prediction may be the same type used for the ground truth labels 541b used to train model 170b.

[0170] In this example, the segmentation results and model outputs are combined to characterize an object. For example, characterization data 580b for one object indicates that the object is made of PET, that the additive PMDA is present at a mass fraction of 0.68, that the additive PBO is present at a mass fraction of 0.30, and that the object is not contaminated with oil. Each of these elements can be predicted using different models 170b that use the corresponding set of information bandwidths. Several features, such as the resin type and whether or not contamination is present, can be determined based on the segmentation results, and in this example, the object can be segmented as a PET object and it can be shown that there is no contaminated area. In some implementations, based on the base resin type identified through the segmentation or output of model 170b, the system 130 can select which model 170b to use to detect different additives and contaminants. For example, one set of models can be trained to detect additives in the presence of PET, and other sets of models can be trained to detect the same or different additives in other resin types. This allows each model to focus more precisely on spectral features that distinguish each additive from different resin types, which can result in higher accuracy concentration results than general models for multiple resin types.

[0171] During stage (F), data representing the characterization data 580b is provided to the device. For example, the data 580b may be associated with a sample identifier, a timestamp, and potentially contextual data such as sample type and location, and stored in a database 590 or other data storage. As a result, individual objects may be tagged or associated with metadata indicating the chemical composition (e.g., the chemicals present, the amount or concentration of the chemicals, etc.) estimated using the model. As another example, the data 580b may be provided to one or more user devices 585, such as a smartphone, laptop computer, or desktop computer, for presentation to the user in a user interface.

[0172] In this example, the predicted chemical composition is evaluated by the computer system 130 or another device and used to sort objects in the waste stream. For example, the chemical composition can be compared to one or more thresholds corresponding to different bins 515a-515c, conveyors, or other output streams. The prediction can then be used by the computer system 130 or another system to control sorting equipment, such as a sorting machine 510, configured to move objects to different areas or containers. Objects of different categories (e.g., groups with different chemical properties) may be treated differently for mechanical properties or chemical recycling. In some implementations, the properties of the entire waste stream can be evaluated and accumulated to determine the overall mixture of plastics and other compounds along the conveyor 505. Even if items are not sorted, information on the properties of the waste stream can be used to adjust the parameters of the chemical recycling of the waste stream, such as adjusting the concentrations of different inputs, solvents, or catalysts used, or adjusting other recycling parameters. As another example, the user may be given a warning if the chemical composition is above a maximum level, below a minimum level, within a range, outside a range, or meets other predetermined conditions.

[0173] Figure 6A is an exemplary block diagram showing image segmentation of a hyperspectral image that can be performed by the feature extraction module 550. Segmentation can be tailored to a specific type of object, such as a particular type of fruit (e.g., strawberries), which is trained by a machine learning model 170 to make predictions. Image segmentation can identify the boundaries of a sample (e.g., fruit) in the hyperspectral image 115 and may also identify specific types of regions of the sample (e.g., leaves, seeds, pulp, etc.).

[0174] The segmentation module 160 receives a hyperspectral image 115 of object 101 as input. The segmentation module 160 processes the input and generates data that identifies the boundaries of different regions of the image as output. The segmentation module 160 processes the hyperspectral image 115 to generate image segments 625, 630, and 635, where image segment 625 represents the leaves and / or stem of the strawberry, image segment 630 represents the seeds on the outer surface of the strawberry, and image segment 635 represents the flesh outside the strawberry. Of these, only segment 635 indicates the sugar content of the strawberry, so only segment 635 is used to generate feature values ​​for input to the machine learning model 170.

[0175] Figure 6B is an exemplary block diagram illustrating the generation of a set of feature values ​​660 for input to a machine learning model 170. This may include generating extended bands from multiple spectral bands of image 115. The system uses a region identified as segment 635, rather than the entire image 115, ignoring pixels from the rest of image 115. After segmentation, segment 635 contains values ​​for the same spectral bands as the original hyperspectral image 115, but only pixels within a specific spatial region are considered.

[0176] The feature extraction module 150 may include an extended bandwidth generator 650 that processes image data from different spectral bands to generate extended bandwidths for the image data. These extended bandwidths can be based on a combination of image data from different bands, and one of various operations such as addition, subtraction, multiplication, or division of the per-pixel intensity values ​​of one band with the corresponding per-pixel intensity values ​​of another band may be applied. In some cases, the extended bandwidth may be calculated by calculating the intensity value of band 1, dividing it by the intensity value of band 2, and then squaring the quotient (e.g., (band 1 / band 2)). 2 This can include nonlinear combinations with different input bandwidths, such as outputting ).

[0177] The feature extraction module generates an input vector 660 for the machine learning model 170. For example, the extended band generator 650 takes a segmented hyperspectral image 640 as input and generates an input vector 660 containing the average intensity values ​​of the segmented images for each band within a given set of bands, e.g., bands 1, 3, 26 and extended bands 2 and 5. This set of bands may be a subset of the total number of available bands, and the subset has been determined through previous analysis to be the set that is most appropriate, or at least minimally relevant, to predicting the characteristics that the model 170 is to predict. In some implementations, in addition to or instead of the average intensity values, other characteristics may be generated and provided as input to the machine learning model 170. For example, values ​​indicating maximum intensity, minimum intensity, intermediate intensity, intensity distribution characteristics, etc., may be provided as features in addition to or instead of the arithmetic mean of the intensity values ​​in the segmented portions of the image data for the spectral bands. Other examples of features include whether or not peaks, troughs, zero crossings, or baseline crossings are detected in a given frequency band (e.g., spectrum, preprocessed spectrum, derivative of spectrum, derivative of preprocessed spectrum, second derivative of spectrum, second derivative of preprocessed spectrum, etc.), the integral of the convolution of the kernel and spectrum (or the derivative or second derivative of the kernel and spectrum), the number of peaks, etc.

[0178] Figure 7 is a flowchart showing an example of a process 700 for analyzing the chemical composition of a particular type of fruit. The operation of process 700 can be performed by one or more data processing devices or computing devices, such as the computer system 130 described above. The operation of process 700 can also be implemented as instructions stored in a computer-readable medium. The execution of the instructions can cause one or more data processing devices or computing devices to perform the operation of process 700. The operation of process 700 can also be implemented by a system including one or more data processing devices or computing devices and a memory device that stores instructions causing one or more data processing devices or computing devices to perform the operation of process 700.

[0179] Prior to process 700, a machine learning model 170 can be trained. For example, the model can be trained using a set of training images containing multiple images of the type of object to be analyzed, e.g., strawberries. Each image in the set of training images may have a corresponding ground truth value for each of the chemical properties to be predicted, e.g., the sugar content of the fruit depicted in the image. In some implementations, the images of the fruit used for training are first captured to show the invariant exterior of the object. The ground truth chemical property values ​​are then obtained by performing invasive and / or destructive chemical tests, such as cutting or crushing the fruit to determine the sugar content of the fruit.

[0180] The set of 540 training images may include parameters other than the ground truth sugar content of the strawberries. For example, each example in the set of training images may include the image, ground truth chemical properties (e.g., concentration, amount, or other level of a chemical or group of chemicals), an identifier of the object type, contextual data about the object's location or source, the angle at which the particular image was taken, the camera's position in the scene, and the specifications of the camera system used to capture the particular image. Any or all of these values ​​may be used to provide features as input to a machine learning model during training and inference processes.

[0181] Process 700 acquires image data of an image containing a representation of a specific fruit piece (710). For example, hyperspectral images may be captured for an object for which predictions about its chemical composition are desired.

[0182] Process 700 includes segmenting hyperspectral image data to identify specific types of regions on an object (715). For example, the outer surface of a strawberry generally includes both flesh and achenes. The system processes the hyperspectral image to perform image segmentation and separates the image regions representing the flesh portion of the fruit from the achenes.

[0183] Process 700 provides feature data derived from image data to a machine learning model trained to predict levels of chemical components based on hyperspectral image data (720). For example, the computer system 130 uses the feature extraction module 550 to determine feature values ​​that represent the characteristics of a particular type of identified region. For example, feature values ​​can describe image data for different wavelength bands within a region identified as corresponding to a particular type of region. For example, a particular type of region might be the fleshy outer part of a strawberry, excluding the achenes and calyx, as well as the background of the image. Feature values ​​can be aggregate values ​​of different bands, such as the average intensity of different bands. One exemplary type of feature value is the average intensity value taken across all pixels for one wavelength band located within the region identified as corresponding to the strawberry flesh. Aggregates for bands could be, for example, the mean, median, mode, geometric mean, or the percentage of pixels that fall within a certain range of values. The set of feature values ​​is then processed by the machine learning model to determine an output indicating the predicted levels of chemical components.

[0184] In addition to feature values ​​for a single wavelength band, the system can also obtain feature values ​​for a combined or composite band. For example, the system can combine certain predetermined bands using a given function to create a combination that the system identifies as predicting the chemical properties of the subject. The system can then derive one or more feature values ​​from the composite image, such as the average pixel intensity across a set of pixels located within a region identified as corresponding to the pulp of a strawberry.

[0185] Process 700 includes receiving an output (730) which includes a prediction of the levels of chemical composition of an object represented in the image. For example, an extracted feature representing the outer part of a strawberry is provided as input to a trained machine learning model 170. The machine learning model 170 processes the input feature based on its trained parameters and generates a value which is the predicted sugar content 170 of the strawberry shown in image 115.

[0186] Process 700 includes providing output data indicating predicted levels of the chemical composition of an object (740). For example, the prediction may be provided to a user device for presentation or to a sorting or packaging machine to control the handling of the object. Similarly, the predicted value or classification of an object may be stored in a database in relation to an object identifier or an identifier for a group of objects containing the object.

[0187] Examples The imaging system is configured to capture line-scan hyperspectral images of an object moving along a conveyor belt. Two reference hyperspectral images are collected, each representing a state where no object is present on the conveyor belt. One ("bright") reference hyperspectral image is collected with each light source of the imaging system turned on at maximum brightness to illuminate the object with spectrally flat and spatially uniform reflectivity, while one ("dark") reference hyperspectral image is collected with each light source of the imaging system cut off from illuminating the camera. Figures 8A–8C show exemplary hyperspectral data for the dark reference hyperspectral image. Figure 8A shows a color map where colors identify intensity, the x-axis (columns) spans location, and the y-axis (rows) spans frequency bands. Figure 8B shows a slice of the color map across frequency, and Figure 8C shows a slice of the color map across location. Figures 9A–9C show corresponding exemplary hyperspectral data for the bright reference hyperspectral image. The decrease in intensity shown in Figure 9B is due to ambient CO2. Reference hyperspectral images can be collected repeatedly (e.g., daily, weekly, monthly, every six months, etc.) because their intensity can change over time due to the light source and / or ambient light.

[0188] The spectra of one or more chemicals (e.g., pure plastics) are also collected. Figure 10 shows a normalized and pre-treated example of polyethylene plastic observed with a hyperspectral camera. The band between channels 75–125 corresponds to the effect of CH molecular absorbance on the observed signal. The band between 175–200 corresponds to the loss of light due to atmospheric CO2 present during the measurement. Although the spectrum is shown here for only one specific exemplary chemical, it will be understood that spectra for other chemicals (e.g., natural PVC, polyethylene black, natural polypropylene, PVC gray, polypropylene black, and / or natural polyethylene) can be collected.

[0189] (For example) each spectrum can contain 308 frequency bands. Detecting signals that can be used to reliably distinguish and / or detect materials with a very large number of values ​​can be difficult due to interband correlation, especially when a given object may contain multiple chemicals or when a single chemical may produce correlated signals across many bands. Therefore, a feature set with fewer values ​​than a given spectrum is defined. Specifically, 32 kernels are designed using a convolutional neural network machine learning model to reduce the dimensionality of the data input from 308 to 32. The convolutional action of the kernels using a hypercube projects the full set of bands at each pixel onto 32 feature maps that can be used to decode the spatial information in subsequent processing steps. The 32 kernels are selected to minimize the prediction loss across the training and evaluation sets.

[0190] Subsequently, a conveyor belt supporting the object begins to move, causing the object to move towards the imaging system's field of view. The imaging system acquires hyperspectral images by performing a line scan. The hyperspectral images are used by a computing system, which preprocesses them using two reference hyperspectral images to show that each measured intensity is close to its maximum value relative to its minimum value. The preprocessed hyperspectral images are then segmented so that pixels corresponding to the object are detected.

[0191] In the first analysis, for each pixel corresponding to an object, the pre-processed spectrum is convolved with each of 32 kernels. The 32 kernels result in 32 feature maps. The feature maps are then fed into a convolutional neural network to process the spatial correlations between pixel features. The output of this second convolutional neural network may consist of chemical composition, more detailed analytical chemistry data (such as a predicted absorbance spectrum), or the presence of contaminants.

[0192] Various implementations of the systems and technologies described herein may be realized in digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be dedicated or general-purpose and coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device.

[0193] These computer programs (also known as programs, software, software applications, or code) include machine instructions for programmable processors and may be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly language / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” mean any computer program product, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, and include machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signal” means any signal used to provide machine instructions and / or data to a programmable processor.

[0194] To provide user interaction, the systems and technologies described herein can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) on which the user can provide input to the computer. Other types of devices can similarly be used to provide user interaction; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user can be received in any form, including acoustic, voice, or tactile input.

[0195] The systems and technologies described herein may be implemented in computing systems that include backend components (e.g., as data servers), middleware components (e.g., application servers), or frontend components (e.g., client computers having a graphical user interface or web browser that allows users to interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., communication networks). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), peer-to-peer networks (with ad-hoc or static members), grid computing infrastructure, and the Internet.

[0196] A computing system can include clients and servers. Clients and servers are generally geographically separated from each other and typically interact via a communication network. The client-server relationship arises from computer programs running on each computer that have a client-server relationship with each other.

[0197] Several implementations have been described. Nevertheless, it should be understood that various modifications are possible. For example, the various forms of the flow shown above can be used with steps rearranged, added, or removed. Also, while several applications and methods for providing incentives for media sharing have been described, it should be recognized that numerous other applications are conceived. Therefore, other implementations are within the scope of the following claims.

Claims

1. A method performed by one or more computers, wherein the method is The acquisition of image data by one or more computers, which shows the degree to which one or more objects in the recycled line material reflect, scatter, or absorb light in each of a plurality of wavelength bands, wherein the image data is collected while the conveyor belt is moving the one or more objects, Preprocessing the aforementioned image data to generate preprocessed image data, wherein the preprocessing includes performing an analysis over frequency and / or an analysis over spatial dimension representation, The above-mentioned one or more computers generate a set of feature values ​​derived from the preprocessed image data, The process involves one or more computers, and based on the output generated by the machine learning model in accordance with the set of feature values ​​provided as input to the machine learning model, generating a prediction of the identity of a chemical substance in one or more objects or the level of one or more chemical substances in one or more objects, wherein the prediction includes the identity of a contaminant or a particular type of plastic or the level of a contaminant or a particular type of plastic. A method comprising providing data by one or more computers indicating the prediction of the identity of the chemical substance in one or more objects or the level of the one or more chemical substance in one or more objects.

2. Preprocessing the aforementioned image data is The method according to claim 1, comprising normalizing the hyperspectral data using one or more reference image datasets.

3. Preprocessing the aforementioned image data is Using the aforementioned image data, generate the derivative, Identifying the threshold, The method according to claim 1, comprising performing threshold cross-analysis using the derivative of the image data and the threshold.

4. The aforementioned image was collected by a camera, and the camera, A lens, wherein the optical axis of the lens is positioned at an angle of 40 to 50 degrees with respect to the surface of the conveyor belt, or An image sensor, wherein the optical axis of the image sensor is positioned at an angle of 40 to 50 degrees with respect to the surface of the conveyor belt. The method according to claim 1, comprising:

5. The aforementioned image was collected by a camera, and the camera, A lens, wherein the optical axis of the lens is positioned at an angle of 85 to 95 degrees with respect to the surface of the conveyor belt, or An image sensor, wherein the optical axis of the image sensor is positioned at an angle of 85 to 95 degrees with respect to the surface of the conveyor belt. The method according to claim 1, comprising:

6. The aforementioned image was collected by a camera, and the camera, A lens, wherein the optical axis of the lens is positioned such that it is at an angle of less than 15 degrees with respect to the surface of the conveyor belt, or An image sensor, wherein the optical axis of the image sensor is positioned such that it is at an angle of less than 15 degrees with respect to the surface of the conveyor belt. The method according to claim 1, comprising:

7. It is a method, The image data includes a value that identifies the reflectance, absorptiveness, or absorbance corresponding to each position in the set of positions along the dimensions of the conveyor belt, and each frequency in the set of frequencies. The method further includes using segmentation techniques to identify a subset of the set of locations as corresponding to a specific object, The method according to claim 1, wherein generating the set of feature values ​​includes generating the set of feature values ​​derived from a portion of the image data corresponding to the subset of the set of positions.

8. To generate the aforementioned set of feature values, Accessing the kernel set, The method according to claim 1, comprising convolving each of one or more portions of the preprocessed image data using each of the set of kernels.

9. The method according to claim 8, wherein each of the at least one of the set of kernels includes a frequency signature corresponding to a particular set of chemicals.

10. The method according to claim 1, wherein the image data includes hyperspectral image data.

11. The method according to claim 1, comprising sorting plastic objects from a waste stream based on the predicted identity of the chemical substance in at least one of the plastic objects, or based on the predicted level of one or more chemical substances in at least one of the plastic objects.

12. The method according to claim 1, wherein the machine learning model is a decision tree or a neural network.

13. The method according to claim 1, wherein the preprocessing of the image data is performed in the same computational workflow as the machine learning.

14. The method according to claim 1, wherein the prediction is a prediction of the identity of a major component within one of the one or more objects.

15. It is a system, One or more computers, One or more computer-readable media that, when executed by the one or more computers, stores instructions that can be operated to cause the system to perform an operation, wherein the operation is Image data indicating the degree to which one or more objects reflect, scatter, or absorb light in each of multiple wavelength bands, wherein the image data is collected while a conveyor belt is moving the one or more objects, and the acquisition of the image data. Preprocessing the aforementioned image data to generate preprocessed image data, wherein the preprocessing includes performing an analysis over frequency and / or an analysis over spatial dimension representation, To generate a set of feature values ​​derived from the aforementioned preprocessed image data, Based on the output generated by the machine learning model in response to the set of feature values ​​provided as input to the machine learning model, the following is to generate a prediction of the identity of a chemical substance in one or more objects or the level of one or more chemical substances in one or more objects: A system comprising providing data indicating the identity of the chemical substance in one or more objects or the predicted level of the chemical substance in one or more objects.

16. Preprocessing the aforementioned image data is The system according to claim 15, comprising normalizing the hyperspectral data using one or more reference image datasets.

17. Preprocessing the aforementioned image data is Using the aforementioned image data, generate the derivative, Identifying the threshold, The system according to claim 15, comprising performing threshold cross-analysis using the derivative of the image data and the threshold.

18. The aforementioned image was collected by a camera, and the camera, A lens, wherein the optical axis of the lens is positioned at an angle of 40 to 50 degrees with respect to the surface of the conveyor belt, or An image sensor, wherein the optical axis of the image sensor is positioned at an angle of 40 to 50 degrees with respect to the surface of the conveyor belt. The system according to claim 15, having the following features.

19. The aforementioned image was collected by a camera, and the camera, A lens, wherein the optical axis of the lens is positioned at an angle of 85 to 95 degrees with respect to the surface of the conveyor belt, or An image sensor, wherein the optical axis of the image sensor is positioned at an angle of 85 to 95 degrees with respect to the surface of the conveyor belt. The system according to claim 15, having the following features.

20. One or more non-temporary computer-readable media that store instructions operable to cause the system to perform an operation when executed by the one or more computers, wherein the operation is Image data indicating the degree to which one or more objects reflect, scatter, or absorb light in each of multiple wavelength bands, wherein the image data is collected while a conveyor belt is moving the one or more objects, and the acquisition of the image data. Preprocessing the aforementioned image data to generate preprocessed image data, wherein the preprocessing includes performing an analysis over frequency and / or an analysis over spatial dimension representation, To generate a set of feature values ​​derived from the aforementioned preprocessed image data, Based on the output generated by the machine learning model in response to the set of feature values ​​provided as input to the machine learning model, the following is to generate a prediction of the identity of a chemical substance in one or more objects or the level of one or more chemical substances in one or more objects: A non-temporary computer-readable medium that includes providing data indicating the identity of the chemical substance in one or more objects or the prediction of the level of the one or more chemical substance in one or more objects.