A system and a method for detection and recognition of materials

A multi-stage AI workflow with specialized neural networks enhances the speed and quality of material detection and recognition in recycling, improving sorting efficiency through precise material classification.

WO2025217742A1PCT designated stage Publication Date: 2025-10-23INDS MACHINEX INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2025/050567
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-18
Filing Date
2025-04-17
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing systems for automatic detection and recognition of materials in recycling lack speed and quality, necessitating improvements in classification and characterization.

Method used

A multi-stage artificial intelligence workflow utilizing distinct neural networks for object detection and classification, including convolutional neural networks (CNN), vision transformer-based models, and multilayer perceptron-based models, combined with sensor data from various systems to enhance precision in material identification.

Benefits of technology

Significantly improves the precision and accuracy of material detection and classification, enabling efficient sorting and mechanical manipulation of objects based on their composition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025050567_23102025_PF_FP_ABST
    Figure CA2025050567_23102025_PF_FP_ABST
Patent Text Reader

Abstract

A method and a system for improving precision in detection and recognition of material. The method comprises executing at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by a camera, the scene image illustrating the objects made of materials; and executing at least one object classification routine comprising at least one second neural network to generate classification results related to materials of the objects; and generating a classification identification for each one of the objects based on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.
Need to check novelty before this filing date? Find Prior Art

Description

A SYSTEM AND A METHOD FOR DETECTION AND RECOGNITION OF MATERIALSRELATED APPLICATION

[0001] The present application claims priority to or benefit of United States provisional patent application No. 63 / 635,779, filed April 18, 2024, and United States provisional patent application No. 63 / 635,778, filed April 18, 2024, which are incorporated herein by reference in their entirety.TECHNICAL FIELD

[0002] The present disclosure relates to systems and methods for processing objects made of various materials. More specifically, it relates to detection and recognition of materials.BACKGROUND

[0003] Various systems exist for characterization and sorting materials. Sorting systems may be important in various industries, including, but not limited to, the recycling industry. Proper recognition of the objects being recycled is the major task of sorting.

[0004] Currently known systems and methods for automatic detection of the materials and recognition of the objects still need to be improved, by increasing not only the speed but also the quality of such detection and recognition, which depends on classification of the materials.SUMMARY

[0005] According to one aspect of the disclosed technology, there is provided a method comprising: executing at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by a camera, the scene image illustrating the objects made of materials; and executing at least one object classification routine comprising at least one second neural network to generate classification results related to materials of the objects; and generating a classification identification for each one of the objects based on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.

[0006] According to another aspect of the disclosed technology, there is provided a system comprising: a camera configured to capture scene images of objects made of materials; a display; and a processor configured to: execute at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and togenerate segmented images of the objects using a scene image captured by the camera, the scene image illustrating the objects made of materials; execute at least one object classification routine comprising at least one second neural network to generate classification results related to materials of the objects; and generate a classification identification for each one of the objects based on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.

[0007] In at least one embodiment, each one of the at least one first neural network is: a first convolutional neural network (CNN) having at least two convolutional layers, a first vision transformer based (ViT-based) model, a first multilayer perceptron based (MLP-based) model, or a first hybrid model comprising at least two of: elements of the first CNN model, elements of the first ViT-based model, and elements of the first MLP-based model; and each one of the at least one second neural network is: a second CNN having at least two convolutional layers, a second ViT-based model, a second MLP-based model, or a second hybrid model comprising at least two of: elements of the second CNN, elements of the second ViT-based model, and elements of the second MLP-based model. In at least one embodiment, each one of the at least one first neural network is: a first convolutional neural network (CNN) having at least two convolutional layers, a first vision transformer-based (ViT-based) model, a first multilayer perceptron-based (MLP-based) model, a first autoencoder-based model, a first contrastive learning model, a first generative model which may be, for example, a generative adversarial network (GAN) or a diffusion-based model, or a first hybrid model comprising at least two components selected from: elements of the first CNN model, elements of the first ViT-based model, elements of the first MLP-based model, elements of the first autoencoder-based model, elements of the first contrastive learning model, and elements of the first generative model; and each one of the at least one second neural network is: a second convolutional neural network (CNN) having at least two convolutional layers, a second vision transformer-based (ViT-based) model, a second multilayer perceptron-based (MLP-based) model, a second autoencoder-based model, a second contrastive learning model, a second generative model, which may be, for example, a generative adversarial network (GAN) or a diffusion-based model; or a second hybrid model comprising at least two components selected from: elements of the second CNN model, elements of the second ViT-based model, elements of the second MLP-based model, elements of the second autoencoder-based model, elements of the second contrastive learning model, and elements of the second generative model. The first generative model may be a first generative adversarial network (GAN) or a first diffusion-based model, and the second generative model is a second GAN or a second diffusion-based model.

[0008] The at least one first neural network may be a first machine learning model configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a self-supervised learning mode, or a semi-supervised learning mode; and the at least one second neural network may be a second machine learning model configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a self-supervised learning mode, or a semi-supervised learning mode.

[0009] The at least one object detection routine may comprise a bounding boxes generation routine, and generating detected object locations may comprise generating bounding boxes coordinates of bounding boxes surrounding the objects by the bounding boxes generation routine. The at least one object detection routine may comprise a semantic segmentation routine configured to generate polygon coordinates of detected objects polygons representing the objects, and generating the detected object locations may comprise generating the polygon coordinates. The detected objects polygons representing the detected objects may be generated within the delimitations of the bounding boxes generated by the bounding boxes generation routine. Prior to generation of the classification identification, the method may generate a sorting decision using the classification results and complementary classification results received from at least one complementary classification unit.

[0010] Sensor data may be generated and captured at the time of acquisition of the scene image, and the sensor data may be used as input to the object detection routine and the object classification routine. The sensor data may be generated by at least one of a near-infrared (NIR) system, and a short-wave infrared (SWIR) system. In at least one embodiment, the sensor data may be generated by the near-infrared (NIR) system and the short-wave infrared (SWIR) system. The sensor data may be generated by a laser sensor, by a volumetric sensor, by a point measurement system for visible spectroscopy, by a middle wavelength infrared (MWIR) system, by a radiography X-ray system, by a fluoroscopy X-ray system, by a thermal camera, by a visible detector, by an invisible marker detector. The sensor data may be generated by at least one of a laser sensor, a volumetric sensor, a point measurement system for visible spectroscopy, a near-infrared (NIR) system, a short-wave infrared (SWIR) system, a middle wavelength infrared (MWIR) system, a radiography X-ray system, a fluoroscopy X-ray system, a thermal camera, a visible detector, or an invisible marker detector.

[0011] The classification results related to each object may be mapped with other complementary unit classification results generated by and received from at least one complementary classification unit. At least one of the detected object locations, the segmented images and the classification results may be further used to track items within a plurality of frames.

[0012] The classification results may be used to generate breakdown statistics of the classified objects. The method may comprise generating breakdown statistics of the classified objects using the classification results. The classification results are further used to estimate weight of the classifiedobjects. The method may comprise estimating weight of the classified objects using the classification results. The classification results may be further used in instructions to perform a mechanical sorting operation on each one or a subset of the classified objects. The method may comprise instructing performing a mechanical sorting operation on each one or a subset of the classified objects. The classification results may be transmitted to and used as an input and instructions in other systems to adjust operation of the other systems within a sorting facility. The method may comprise transmitting to and using as an input and instructions the classification results in other systems to adjust operation of the other systems within a sorting facility. The classification identification may be further used to generate instructions transmitted to a targeting system. The method may comprise generating instructions and transmitting the instructions to a targeting system. The camera may comprise an RGB camera or a grayscale camera.

[0013] According to a further aspect of the disclosed technology, there is provided a method comprising: executing at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by a camera, the scene image illustrating the objects made of materials, each one of the at least one first neural network being: a first convolutional neural network (CNN) having at least two convolutional layers, a first vision transformer based (ViT-based) model, a first multilayer perceptron based (MLP-based) model; or a first hybrid model comprising at least two of: elements of the first CNN model, elements of the first ViT-based model, and elements of the first MLP-based model; and executing at least one object classification routine comprising at least one second neural network to generate classification results related to materials of the objects, each one of the at least one second neural network being: a second CNN having at least two convolutional layers, a second ViT-based model, a second MLPbased model, or a second hybrid model comprising at least two of: elements of the second CNN, elements of the second ViT-based model, and elements of the second MLP-based model; and generating a classification identification for each one of the objects based on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.

[0014] A method and a system for improving precision in detection and recognition of material are provided. The method comprises executing at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by a camera, the scene image illustrating the objects made of materials; and executing at least one object classification routine comprising at least one second neural network to generate classification results related to materials of the objects; and generating a classification identification for each one of the objectsbased on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.

[0015] According to another aspect of the disclosed technology, there is provided a method comprising: executing at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by a camera, the scene image illustrating the objects made of materials, the at least one first machine learning model being: a first convolutional neural network (CNN) having at least two convolutional layers; a first vision transformer-based (ViT-based) model; a first multilayer perceptron-based (MLP-based) model; a first autoencoder-based model; a first contrastive learning model; a first generative model; or a first hybrid model comprising at least two components selected from: elements of the first CNN model, elements of the first ViT-based model, elements of the first MLP-based model, elements of the first autoencoder-based model, elements of the first contrastive learning model, and elements of the first generative model; executing at least one object classification routine comprising at least one second machine learning model to generate classification results related to materials of the objects, each one of the at least one second being: a second convolutional neural network (CNN) having at least two convolutional layers; a second vision transformer-based (ViT-based) model; a second multilayer perceptron-based (MLP-based) model; a second autoencoder-based model; a second contrastive learning model; a second generative model; and a second hybrid model comprising at least two components selected from: elements of the second CNN model, elements of the second ViT-based model, elements of the second MLP-based model, elements of the second autoencoder-based model, elements of the second contrastive learning model, and elements of the second generative model; generating a classification identification for each one of the objects based on the classification results, the at least one first neural network being trained differently from the at least one second neural network.

[0016] The first machine learning model may be configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a self-supervised learning mode, or a semi-supervised learning mode and the second machine learning model may be configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a selfsupervised learning mode, or a semi-supervised learning mode.

[0017] In at least one embodiment, a method and a system for improving precision in detection and classification of material are herein provided. The method may comprise executing at least one object detection routine, each one comprising at least one first convolutional neural network (CNN), and executing at least one object classification routine, each one comprising at least one secondCNN, to generate classification identification of each one of objects captured in a scene image by a camera, wherein the at least one first CNN is trained differently from the at least one second CNN. The method may generate classification identification and segmented and classified object representation which may be labeled according to the classification identification. The classification identification may be used, for example, to instruct a mechanical sorting system to manipulate the objects. According to one aspect of the disclosed technology, the method may comprise: executing at least one object detection routine comprising at least one first convolutional neural network (CNN), each one of the at least one first CNN having at least two convolutional layers, and generating detected object locations representing locations of objects based on a scene image received from a camera, the scene image illustrating the plurality of objects; and executing at least one object classification routine comprising at least one second CNN, each one of the at least one second CNN having at least two convolutional layers, and generating a classification identification of each one of the objects based on the detected object images, the at least one first CNN being trained differently from the at least one second CNN.

[0018] The at least one object detection routine may comprise a bounding boxes generation routine, and the generating detected object locations may comprise generating bounding boxes coordinates of bounding boxes surrounding the objects by the bounding boxes generation routine. The at least one object detection routine may comprise a semantic segmentation routine configured to generate polygon coordinates of detected objects polygons representing the objects, and generating the detected object locations may comprise generating the polygon coordinates. The detected objects polygons representing the detected objects may be generated within the delimitations of the bounding boxes generated by the bounding boxes generation routine. The routines may also receive input sensor data from a laser sensor. The routines may also receive input sensor data from a volumetric sensor. The routines may also receive input sensor data from a point measurement system for visible spectroscopy. The routines may also receive the input sensor data from a near-infrared (NIR) system and / or sensor data from a short-wave infrared (SWIR) system and / or sensor data from a middle wavelength infrared (MWIR) system. The routines may also receive the input sensor data from a radiography X-ray system. The routines may also receive the input sensor data from a fluoroscopy X-ray system. The routines may also receive the input sensor data from a thermal camera. The routines may also receive the input sensor data from a visible detector. The routines may also receive the input sensor data from an invisible marker detector.

[0019] The classification results related to each object may be matched with other classification results generated by and received from at least one other object classification system. The classification results may be used to produce breakdown statistics of the classified objects. The classification results may be further used to estimate weight of the classified objects. Theclassification results may be further used in instructions to perform a mechanical sorting operation on each one or a subset of the classified objects. The classification results may be transmitted to and used as an input and instructions in other systems to adjust operation of the other systems within a sorting facility. The classification identification may be further used to generate instructions transmitted to a targeting system. The camera may comprise an RGB camera or a grayscale camera.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Further features and advantages of the present disclosure will become apparent from the following detailed description, taken in combination with the appended drawings, in which:

[0021] Fig 1 is a schematic block diagram of a system for improving precision in detection and classification of objects that are being recycled, in accordance with at least one embodiment of the present disclosure;

[0022] Fig. 2 illustrates a flow diagram of a method for improving precision in detection and classification of objects, in accordance with at least one embodiment of the present disclosure;

[0023] Fig. 3 illustrates a non-limiting example of the final output image, in accordance with at least one embodiment of the present disclosure; and

[0024] Fig. 4 illustrates non-limiting examples of scene images, bounding boxes images, segmented images, and the final output image, in accordance with at least one embodiment of the present disclosure.

[0025] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION

[0026] Various aspects of the present disclosure generally address one or more of the problems of detecting the material, such as, for example, recycled or recyclable material or construction materials, such as, for example, plastics, paper, etc. The present description provides a system and a method for improving precision in detection and classification of materials that various objects are made of. The system and the method as described herein use a multi-stage artificial intelligence (Al) workflow for the improvement of the precision of material detection and recognition. The system and method as described herein helps in characterization and classification of materials that the objects are made of. The system and method as described herein may be used for sorting objects made ofvarious materials. For example, the system and method as described herein may be used for sorting objects made of recycled or recyclable material or construction materials, such as, for example, plastics, paper, etc.

[0027] Fig. 1 is a schematic block diagram of a system 100 for executing a method 200 for improving the precision in detection and classification of objects 105 that are being recycled, in accordance with at least one embodiment of the present disclosure. Fig. 2 illustrates the execution steps (in other words, a flow diagram) of the method 200 for improving the precision in detection and classification of objects 105, in accordance with at least one embodiment of the present disclosure. When discussing the system 100 and the method 200 herein below, reference will be made to Fig. 1 and Fig. 2.

[0028] The system 100 comprises at least one camera 110. In at least one embodiment, the camera 110 may be an RGB (red, green, blue) camera and / or a grayscale camera. For example, more than one camera 110 may comprise one or more RGB camera, one or more greyscale camera, or both RBG camera(s) and greyscale camera(s). The camera(s) 110 are configured to capture scene images 115, which may be RGB image(s) and / or grayscale image(s), respectively, obtained by the RGB camera(s) and / or the grayscale camera(s). The one or more camera 110 may operate on a line-scan or area-scan basis.

[0029] The system 100 may also have at least one additional sensor 112. In addition to the scene images 115, additional sensor data 117 may be generated by the additional sensors 112 (also referred to herein as “additional devices 112”). In the technology as described herein, various additional sensors 112 are configured to obtain and generate the additional sensor data 117 (which may be also referred to herein as the “complementary data 117” or the “input sensor data 117” or “sensor data 117”). The additional sensors 112 may be one or more of the following: a laser sensor, a volumetric sensor, an electromagnetic detector, a point measurement system for visible spectroscopy, a near-infrared (NIR) system, a short-wave infrared (SWIR) system, a middle wavelength Infrared (MWIR) system, a sensor for measuring X-rays, radiography X-ray system, a sensor for measuring X-ray fluorescence (fluoroscopy X-ray system), a thermal camera, a visible marker detector, an invisible marker detector. In at least one embodiment, the laser sensor is configured to determine the height and / or presence of items. Preferably, the additional sensors 112 are the NIR system and the SWIR system. In some embodiments, the system 100 does not have the additional sensors 112.

[0030] Still referring to Fig. 1 , in at least one embodiment, the system 100 comprises a processor 150 and a non-transitory computer readable medium with computer executable instructions stored thereon. In some embodiments, the processor 150, a memory, and the non-transitory computerreadable medium are located on a server. The processor 150 as described herein is configured to execute an application which executes or otherwise has access to routines as described herein. The camera 110 and the additional sensors 112 may communicate with the server 150 via a network (for example, a wireless or a wired network). In some other embodiments, the memory, the processor 150 and the non-transitory computer readable medium are located in an electronic device. For example, the electronic device may be a computer, an iPad, a phone, etc. In at least one embodiment, the electronic device may also comprise a screen (which may be the same or different from display 160 of Fig. 1) for displaying a final decision 250, generated by the processor 150, comprising classification identification.

[0031] The system 100 comprises at least two modules 120, each having and using at least one neural network for extraction of characteristics of the objects 105. In at least one embodiment, the neural network may be a convolutional neural network (also referred to as a “convolutional layer deep learning network” or “CNN”), a vision-transformer-based model (also referred to as a “ViT-based model”), or a multilayer perceptron-based model (also referred to as a “MLP-based model”), or a hybrid model described below.

[0032] In at least one embodiment, the neural network may be at least one CNN which is configured to extract characteristics of the objects 105. In at least one embodiment, each CNN has at least one convolutional layer. In at least one preferred embodiment, each CNN has at least two convolutional layers. As discussed below, each CNN is trained to execute specific tasks of the module that particular CNN belongs to, and is therefore narrowly focused and trained to perform specific tasks.

[0033] In at least one embodiment, the neural network may be the ViT-based model. The ViT- based model is a visual model based on the architecture of a transformer originally designed for textbased tasks. The ViT-based model represents an input scene image 115 as a plurality of scene image patches. The ViT-based model is configured to directly predict class labels for the input scene image 115.

[0034] The ViT-based model may utilize a transformer architecture adapted for visual tasks, in which input image data is divided into patches and embedded for sequential processing using selfattention mechanisms to generate segmentation maps.

[0035] In at least one embodiment, the neural network may be the MLP-based model. The MLPbased model is a feedforward neural network having fully connected neurons with nonlinear activation functions, organized in layers. The MLP-based model is configured to distinguish data thatis not linearly separable. The MLP-based model may employ a sequence of fully connected layers to learn hierarchical feature representations for pixel-wise classification.

[0036] The one or more neural networks may be also a hybrid model which has elements of one or more CNN and / or elements of one or more ViT-based model and / or elements of one or more MLPbased model. In other words, the hybrid model may combine elements (aspects) from CNN models and / or ViT-based models and / or MLP-based models. The hybrid model may incorporate convolutional layers for local feature extraction, followed by MLP and transformer-based components to enhance both spatial resolution and global context understanding. Each of these model types may be pre-trained on large-scale image datasets and fine-tuned for domain-specific segmentation tasks. For instance, the hybrid model may use a CNN for early visual processing (such as feature extraction), the MLP-based model for dense decision-making, and the ViT-based model for global context. The neural network may be a hybrid model comprising any combination of elements of (for example at least two of): at least one CNN, at least one ViT-based model, and / or at least one MLPbased model.

[0037] The modules 120 may be connected to or implement one or more transformers 130 which are routines also executed by the processor 150. In at least one embodiment, the transformers are implemented by the processor 150 using a combination of software and hardware, for example, for the routines 121 , 122, that are implemented using ViT-based or MLP-based models.

[0038] In at least one embodiment, the transformers (for the ViT-based or MLP-based models) or other elements of the neural network (such as, for example, the CNN or the hybrid model) are configured to process data to determine an object classification category of each object 105 in addition to other units and models described herein. The result of such classification - the object classification category - may be combined with one or more complementary unit classification results 141 (which may be also referred to herein as “additional unit classification results 141”) from one or more complementary classification unit 140 (also referred to as a “complementary sorting unit 140” or a “additional classification unit 140”) to determine a final classification decision 250.

[0039] Still referring to Figs. 1 and 2, in at least one embodiment, the system 100 may receive the complementary unit classification result 141 for any targeted object 105. The complementary unit classification results 141 may be received, for example, without receiving the data from the additional sensors 112 such as additional sensor data 117 used to determine the complementary unit classification results 141. The complementary unit classification result 141 may be received from the complementary classification unit 140 which may be any sorting software and / or hardware that returns a complementary unit classification result 141 for any targeted object 105. The complementary classification unit 140 may be, for example, a classification routine or a combinationof classification routines, implemented by the processor 250 or, alternatively, implemented by a third party, and configured to generate the complementary unit classification results 141 , as illustrated in Fig. 2. In some embodiments, the complementary classification unit 140 may be another system, similar to the system 100 as described herein, configured to generate the complementary unit classification result 141.

[0040] To generate the complementary unit classification results 141 , one or more complementary classification units 140 use data obtained using one or more of the following technologies: RGB camera that generates a RGB scene image, grayscale camera that generates a grayscale scene image, a laser sensor used to determine the height and / or presence of the items (objects 105), a volumetric sensor, an electromagnetic detector, a point measurement system for visible spectroscopy, the NIR system, the SWIR system and / or the MWIR system, a sensor for X- rays or for measuring X-ray fluorescence, the thermal camera, a detector for visible or invisible markers (a visible marker detector, an invisible marker detector). Preferably, the complementary classification unit 140 uses data obtained using the NIR system and the SWIR system.

[0041] In at least one embodiment, three steps are performed by three separate models 120. In at least one embodiment, the system 100 comprises 3 modules 120 which implement 4 steps. This is one non-limiting example of the implementation, and other architectures may be implemented (for example, having two or more than three modules implementing two, three, or more steps).

[0042] In at least one embodiment, to obtain the final decision 250, the following routines (which may be also referred to herein as “models”) are implemented (executed): a bounding boxes generation routine 121 (bounding boxes generation model 121), a semantic segmentation routine 122 (semantic segmentation model 122), an object classification routine 123 (object classification model 123), a sorting decision routine 124 (sorting decision model 124).

[0043] Referring to Fig. 2, in at least one embodiment, the method 200 comprises execution of an object detection routine 220 (object detection model 220) and execution of the object classification routine 123. The object detection routine 220 may comprise one or more routines: the object detection routine 220 may comprise the bounding boxes generation routine 221 ; and, in some preferred embodiments, the object detection routine 220 may comprise the semantic segmentation routine 122.

[0044] The “routines” as referred to herein are each configured to execute a series of specific tasks and functions by executing steps as described herein. In at least one embodiment, each corresponding module 120 implements its corresponding routine in software and hardware.

[0045] Each one of the routines 121 , 122, 123 uses its own neural network, such as, for example, CNN (one or more). The neural network of each routine may be trained differently from the neural networks of the other routines and has its own input and output. For example, CNN (one or more) of one routine may be trained differently from the CNNs of the other routines. Training each of the routines 121 , 122, 123 and therefore the neural networks, such as, for example, the CNN, the ViT- based model, the MLP-based model or the hybrid model, separately allows to specialize each neural network in a specific task, significantly improving the quality of the output of each routine and therefore significantly improving the quality of the final decision. Having two or more routines, each having at least one specifically trained neural network results (for example, CNN results) in a significant improvement to the final decision 250 compared to using one overall model and therefore one general neural network (for example, “general CNN”) executing all the tasks leading to the final decision regarding the object. In some embodiments, the neural networks of different routines may be trained similarly, but their architecture is different from each other as described below.

[0046] The object detection routine 220 comprises at least one first neural network and each one of the at least one first neural network may be: a first CNN, a first ViT-based model, a first MLPbased model; and a first hybrid model comprising at least two of: elements of the first CNN model, elements of the first ViT-based model, and elements of the first MLP-based model. In some embodiments, the at least one first neural network may be the first CNN having at least two convolutional layers, the first ViT-based model, the first MLP-based model, a first autoencoder-based model, a first contrastive learning model, a first generative model including but not limited to a generative adversarial network (GAN) or a diffusion-based model, or the first hybrid model comprising at least two of: elements of the first CNN model, elements of the first ViT-based model, elements of the first MLP-based model, elements of the first autoencoder-based model, elements of the first contrastive learning model, and elements of the first generative model. For example, the at least one first neural network may be selected from the group consisting of these models.

[0047] The one object classification routine 123 comprises at least one second neural network and each one of the at least one second neural network may be: a second CNN having at least two convolutional layers, a second ViT-based model, a second MLP-based model, or a second hybrid model comprising at least two of: elements of the second CNN, elements of the second ViT-based model, and elements of the second MLP-based model. In some embodiments, the at least one second neural network may be the second CNN having at least two convolutional layers, the second ViT-based model, the second MLP-based model, a second autoencoder-based model, a second contrastive learning model, a second generative model including but not limited to a generative adversarial network (GAN) or a diffusion-based model, or the second hybrid model comprising atleast two of: elements of the second CNN model, elements of the second ViT-based model, elements of the second MLP-based model, elements of the second autoencoder-based model, elements of the second contrastive learning model, and elements of the second generative model. For example, the at least one second neural network may be selected from the group consisting of these models.

[0048] The first generative model may be a first generative adversarial network (GAN) or a first diffusion-based model, and the second generative model may be a second GAN or a second diffusion-based model. The at least one first neural network may be a first machine learning model configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a self-supervised learning mode, or a semi-supervised learning mode; and the at least one second neural network may be a second machine learning model configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a self-supervised learning mode, or a semi-supervised learning mode.

[0049] In at least one embodiment, the at least one first neural network and the at least one second neural network are different models having different architecture from each other. The models may be of the same type (for example, CNN-type) but having different architecture from each other. For example, if the first neural network used in the object detection routine 220 is CNN, the second neural network used in the object classification routine 123 may be another CNN which has different architecture. In other words, the first and the second neural networks may be both CNNs but have different architecture. For example, and without limitation, the CNNs with different architectures may be YOLO and Faster R-CNN: YOLO treats object detection as a regression problem which directly predicts class probabilities and bounding boxes in a single pass through the network, while Faster R-CNN proposes regions likely to contain objects using a Region Proposal Network (RPN) then classifies those regions and refines bounding boxes. Similarly, the first neural network and the second neural network may be both MLP-based models but with different architectures.

[0050] The bounding boxes generation routine 121 (also referred to herein as “detection routine”) receives, as an input, a scene image 115 captured by the camera 110 (the RGB or grayscale camera). The scene image 115 may be merged with additional data from one or more additional sensors 112 (as discussed above). When processing the input scene image 115, the bounding boxes generation routine 121 locates objects 105 in the input scene image 115 by generating the coordinates of (which may be also referred to as “predicting”) bounding boxes 132 around the objects 105 presented in the scene image 115. The term “locating” as used herein comprises identifying that there is an object 105 present, estimating the coordinate of the center of the object 105, mapping the additional sensor data 117, obtained from the additional sensors 112, related to every detected object105. Each object 105 may be thus assigned a corresponding object-related additional sensor data cut out from and obtained from the additional sensor data 117.

[0051] For example, the additional sensor data 117 may be an X-ray image captured at the same time as (in other terms, at the time of acquisition of, or simultaneously with) the scene image 115. Such X-ray image provides X-ray data for the same objects 105 as in the scene image 115, and the X-ray data for each one (or at least several) of the objects of the scene image is mapped to the corresponding object 105 and its corresponding bounding box 132 by the bounding boxes generation routine 121.

[0052] In at least one embodiment, the additional sensor data 117 is used as input to the object detection routine 220. In at least one embodiment, the additional sensor data 117 is used as input to the object classification routine 123. In at least one embodiment, the additional sensor data 117 may be used as input by any routines described herein in order to improve the accuracy of the final classification decision 250.

[0053] In at least one embodiment, the bounding boxes generation routine 121 has at least one object detection CNN, which is specifically trained and configured to trace the bounding boxes 132 around the objects 105 on the scene images 115 to produce the cut images 131 (which may be also referred to as “belt images 131” or “bounding boxes images 131”). When executing the machine learning algorithm with the use of the object detection CNN of the bounding boxes generation routine 121 (which may be also referred to as a “detection module”), objects 105 in the input scene image 115 are located (in other words, the location of the objects 105 in the input scene image 115 is determined) by generating the coordinates of the bounding boxes 132 which can be traced around those objects 105.

[0054] Training of the bounding boxes generation routine 121 uses scene images 115 captured by the camera 110 with precise bounding boxes 132 surrounding the objects 105. In at least one embodiment, the bounding boxes 132 for training may be generated automatically based on one or more previous iterations of the bounding boxes generation routine 121. In some embodiments, the system 100 as described herein may request for a human input to annotate training images, in order to receive the annotations for the training images and to generate data (with the bounding boxes 132) for training of the bounding boxes generation routine 121.

[0055] As a result of the processing by the bounding boxes generation routine 121 , the cut images 131 are generated which have, instead of the scene images 115, the bounding boxes 132 surrounding the objects 105. These bounding boxes images 131 may be merged with additional sensor data 117 received from one or more additional sensors 112. The output of the bounding boxesgeneration routine 121 comprises the bounding boxes images 131 of the detected objects 105, in addition to the scene images 115, and the coordinates of the bounding boxes 132, also determined by the bounding boxes generation routine 121. In some embodiments, the output of the bounding boxes generation routine 121 may also comprise other information associated with the additional data 117. In some embodiments, bounding boxes images 131 may comprise, in addition to the bounding boxes 132, the scene images 115 illustrating the initial image.

[0056] In at least one embodiment, the method 200 as described herein comprises executing at least one object detection routine 220 which comprises the bounding boxes generation routine 121 and the semantic segmentation routine 122. In at least one embodiment, the at least one object detection routine 220 is configured to generate detected object locations representing locations of objects 105 based on a scene image 115 received from the camera 110. The bounding boxes generation routine 121 may generate, as its output, bounding boxes images 131 and / or bounding boxes coordinates for each object 105. In other terms, the bounding boxes generation routine 121 is the routine configured to generate detected object locations in the form of bounding boxes 132 surrounding the objects 105.

[0057] The semantic segmentation routine 122 is configured to generate detected objects polygons 133 and polygons’ coordinates (coordinates of the detected objects polygons) representing the objects 105. In other terms, the semantic segmentation routine 122 is a routine configured to generate the detected object locations in the form of masks 133 (which are also referred to herein as “polygons 133” or “detected object polygons 133”), delineating their shapes. In at least one embodiment, a mask 133 is made of line segments connected to form a closed polygonal chain, where the segments do not intersect each other. Each mask 133 represents the shape of the object 105. The outside edge of the mask 133 comprises a plurality of straight-line segments. Although the masks 133 are illustrated in Fig. 4 as being regular twelve-sided polygons for the simplicity of illustration, the number of the straight-line segments of the mask 133 may be more than 100, such that the mask 133 represents accurately the shape of the object 105. The straight-line segments may have same lengths or different lengths. The mask 133 may be partially concave and / or partially convex. The mask 133 (polygon 133) follows the edges of the objects 105 in order to draw the contours of the object 105 in the segmented image 125 and the final output image 260. Two or more masks 133 may overlap, due to the detection of corresponding overlapping objects 105 (as illustrated, for example, in Figs. 2, 3, for masks 133a, 133b).

[0058] In at least one embodiment, the detected objects polygons 133 representing the detected objects 105 are generated within the delimitations of the bounding boxes 132 generated by the bounding boxes generation routine 121. The generating of the detected object locations representinglocations of objects 105 may comprise generating bounding boxes coordinates of the bounding boxes 132 surrounding the objects 105 by the bounding boxes generation routine 121 and, in some embodiments, generating of the polygon coordinates (also referred to herein as “mask coordinates”) of detected objects polygons 133 (masks 133) representing the detected objects 105.

[0059] The semantic segmentation routine 122 receives, as an input, the data (for example, bounding boxes coordinates) and bounding boxes images 131 or directly scene images 115 received from the bounding boxes generation routine 121 executed previously or received from the camera 110 if the semantic segmentation routine 122 is executed by the object detection routine 220 directly. The semantic segmentation routine 122 generates, as an output, a segmented version of the input scene image 115, referred to herein as a segmented image 125. Each pixel of the segmented image 125 is classified according to the object 105 to which it belongs. The additional data trained, if present, is also predicted.

[0060] Fig 3 illustrates a non-limiting example of the final output image 260, in accordance with at least one embodiment of the present disclosure. Fig. 3 illustrates the masks 133 and the final decision indications 251 (also referred to as a “classification indication 251”) of the final output image 260, as well as the original objects 105 from the corresponding scene image 115 and some bounding boxes 132 (in dotted lines) for comparison. In at least one embodiment, original objects 105 from the corresponding scene image 115 and the bounding boxes 132 are not illustrated in the final output image 260.

[0061] Fig 4 illustrates non-limiting examples of the scene image 115 with objects 105, the bounding boxes images 131 with bounding boxes 132, the segmented images 125 with masks 133, and the final output image 260 illustrating the masks 133 and the final decision indications 251 inside the masks 133, in accordance with at least one embodiment of the present disclosure.

[0062] The scene images 115 and / or the bounding boxes images 131 are fed into the semantic segmentation model 122 in addition to the additional sensor data 117 measured by the additional sensors 112, if they were present at the time of acquisition of the scene images 115. To train the semantic segmentation model 122, images of the objects 105, each with a bounding box 132 around the object 105 (precise contour) contained in the scene image 115 are used. Additional data 117 from the additional sensors 112 present during the acquisition may also be associated with these traced objects 105. Referring to Figs. 2 and 3, the semantic segmentation routine 122 assigns a class pixel label to each pixel (also referred to as a “class label”) in the region inside the bounding box 132 (so-called “bounding box region”), effectively segmenting objects 105 from an empty bounding box or other nearby objects 105. For example, the empty bounding box may be the box that was cut, but does not have any object 105 therein. The objects within the bounding boxes are segmented(although the goal is to have every single object within bounding box following the detection model run).

[0063] The semantic segmentation routine 122 generates detected object locations which may be represented as detected object outlines (shapes) in the form of polygons (which may be also referred to as “detected objects polygons”). The detected object polygons 133 delineate the shape of each object 105. This improves considerably the representation of the real shape of the object 105 and allows to properly target the manipulations (such as, for example, grabbing of the object) later.

[0064] In at least one embodiment, if two objects 105 (a first object 105a, and a second object 105b) are located one over another, the semantic segmentation routine 122 is configured to recognize that the first object 105a overlaps with the second object 105b, and labeling each pixel with the class pixel label permits delineating the whole shapes of both objects 105a, 105b, such that one pixel on one scene image 115 may be related to both the first object 105a and the second object 105b (a series of points related to the outline of the polygon that represents the object 105a, 105b). The semantic segmentation routine 122 may be configured and trained to determine whether two objects 105a, 105b are overlapping with each other. The semantic segmentation routine 122 may be specifically trained to detect the overlap 135 of two or more objects 105 (for example, two objects 105a, 105b) and then to provide representations of the shapes of the objects 105 in the form of polygons 133a, 133b which are very close to real shapes of those objects 105.

[0065] The semantic segmentation routine 122 comprises its proper at least one segmentation- oriented CNN. The segmentation-oriented CNN used by the semantic segmentation routine 122 becomes significantly advanced when trained to determine the overlaps of the objects while being focused on the tasks of segmentation. Such segmentation-oriented CNN generates the segmented images that have better quality compared to the representation of the objects that could be generated using the general CNN discussed above.

[0066] In at least one embodiment, the semantic segmentation routine 122 may be implemented using the CNN, the ViT-based model, the MLP-based model, and / or the hybrid model as described for the object detection routine 220, while the neural networks used for the semantic segmentation routine 122 may be different and differently trained from the neural networks for the other routines.

[0067] The semantic segmentation routine 122 generates, as an output, a segmented version of the input scene image 115, referred to herein as a segmented image 125. Each pixel of the segmented image 125 is classified according to the object to which it belongs. The additional data trained, if present, is also predicted.

[0068] The object classification routine 123 receives, as an input, the segmented images 125 with segmented objects 133 from the semantic segmentation routine 122. These segmented images 125 may be merged with additional data from one or more additional sensors if present at the time of acquisition, as discussed above. Each segmented image 125 is analyzed by the classification routine 123 to determine the object classification category.

[0069] In at least one embodiment, the object classification category is related to the material of the detected object 105. For example, the object classification category may be polyethylene terephthalate (PET) plastics, high-density polyethylene (HDPE) plastics, polycarbonate (PC) plastics, various papers, which may include, for example, more detailed information about the material that may be determined by the system 100.

[0070] To train the object classification model / routine 123, training segmented objects 133 are sorted by the object classification category. The classification model / routine 123 may be retrained to readjust categories to adjust to customer needs or to improve performance of the model. As an output, the object classification model / routine 123 generates a classification identification which may comprise category labels 251 assigned to each segmented object 133 (also referred to herein as a “category object label”) on the segmented image 125. In other terms, the object classification model / routine 123 generates segmented and classified object representations.

[0071] The sorting decision routine 124 associates each of the objects 105 with the result of an complementary classification unit 140 using one or more of the technologies described above. For example, the complementary classification unit 140 may be implemented externally to the system 100, for example, by a third party, or on a different processor using routines different from those of system 100. The sorting decision routine 124 is optional. In some embodiments, the system 100 as described herein has no complementary classification unit 140 (or is not connected to the complementary classification unit 140 to receive complementary unit classification results 141 from it) or has (is connected to) more than one complementary classification units 140. In the embodiments using more than one complementary classification unit 140, each complementary classification unit 140 generates its proper complementary unit classification result 141. In at least one embodiment, the sorting decision routine 124 uses one or more complementary classification units 140. Each complementary classification unit 140 returns its proper at least one complementary classification result 141.

[0072] The complementary classification unit 140 may receive and use, as input, scene images 115 and / or additional sensor data 117. For example, with reference to Fig. 2, an additional sensor 112 implemented with the sensor for X-rays, may generate X-ray images. The X-ray images may be provided as the additional sensor data 117 to the object detection routine 220. In addition, the X-rayimages may be used by an additional classification unit 140 (for example, externally to the system 100) to generate an X-ray-related classification which may be used as complementary unit classification results 141 to generate a sorting decision 124, together with the classification results 127. The complementary classification unit 140 may also operate in an independent way.

[0073] A decision matrix may be then generated that cross-combines (merges) the result of the complementary classification unit(s) with the output of the object classification routine 123. The final classification result is determined based on (by analyzing) the decision matrix. The final decision 250 may be generated according to the decision matrix of the classification results.

[0074] In at least one embodiment, the classification results 127 related to each object are matched (and / or merged) with other classification results generated by and received from at least one other object classification system (which may be also referred to as an “external object classification system”). These other object classification systems may use pattern recognition algorithms or neural networks to analyze data from their sensors. The comparison is conducted within the decision matrix to generate a consolidated object classification outcome. In some non-limiting examples, weighting or a logic operation may be applied to merge the classification results of the method 200 and the external object classification systems.

[0075] The final decision 250 (classification identification) and / or a final output image 260 may be then displayed on the display 160.

[0076] In at least one embodiment, the classification results 127 and / or the final decision 250 (classification identification) are used to produce breakdown statistics of the classified objects 105 and / or the materials those objects 105 are made of. For example, the system 100 may determine and generate the indication of the quantity of particular materials from which the objects 105 are made, or, for example, the quantity of objects 105 made of specific materials. The classification results 127 and / or the final decision 250 may be further used to estimate weight and / or dimensions (and, for example, area used) of the classified objects 105.

[0077] Based on the classification results 127 and the final decision 250, the system 100 may generate and transmit instructions to perform a mechanical sorting operation on each one of the classified objects 105 or on a subset of the classified objects 105. The mechanical sorting operation may involve, for example, and without limitation, manipulation of a robotic arm, operation of a flip gate, etc. In a non-limiting example, the robotic arm may receive the instructions comprising actions to perform (for example, move), and the classification, as well as the coordinates of the object 105 with the object polygon 133 that was determined by the segmentation model 122.

[0078] In some embodiments, the classification results 127 and / or the final decision 250 are transmitted to and used as an input and instructions in other systems to adjust operation of the other systems within a sorting facility. In some embodiments, the classification results 127 may be further used in and / or to generate instructions transmitted to a targeting system to instruct the targeting system with regards to manipulation of the targeting system, which may help to target the classified object 105 before manipulation. For example, such targeting systems may use lasers to target the objects 105.

[0079] Based on the final decision 250 for each object 105, the final output image 260 is generated. In at least one embodiment, the final output image 260 comprises the detected objects polygons (a first polygon 133a and a second polygon 133b are illustrated for two objects in Fig. 2) and a final decision indication 251 which represents the final decision 250 for each object 105 (first final decision indication 251a for the first object 105a and second final decision indication 251 b for the second object 105b). For example, the final decision indication 251 may be positioned close to or in the middle of the detected objects polygon 133. For example, different colors of different objects 105 in these final output images 260 may represent different classification of the objects 105, while each object 105 is illustrated with its own shape of the detected object polygon 133. The final output images 260 may comprise the detected objects polygons 133 which are color-coded depending on the classification. Overlap 135 is also identified by the method 200 as described herein and therefore may be clearly represented in the final output image 260.

[0080] In at least one embodiment, the detection results (and therefore the bounding boxes 132), segmentation results (such as the segmented images 125) and / or classification results 127 or even final classification decisions 250 are further used to track objects 105 within a plurality of other scene images 115 (a plurality of frames). The system 100 is thus configured to track items (objects) using results from a plurality scene images 115 captured by the camera. The frame is an image representing what is under the camera at a given time.

[0081] The generated final output image 260 may be displayed on the display 160 for the operator to supervise the manipulation of the objects 105, and the colored representation may help to quickly recognize the type of the object 105, for example, and / or the size of the object 105 or other characteristics. In some preferred embodiments, the final output image 260 may be further used in instructions generated for and transmitted to the targeting systems and / or mechanical sorting systems which are configured to execute automatic targeting and / or mechanical sorting operation, respectively.

[0082] The computer-implemented method is described herein configured to execute the routines as described herein. In at least one embodiment, there is provided herein a computerprogram comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method 200 as described herein.

[0083] In at least one embodiment, the method 200 as described herein comprises: executing at least one object detection routine 220 comprising at least one first convolutional neural network (CNN), each one of the at least one first CNN having at least two convolutional layers, and generating detected object locations representing locations of objects based on a scene image received from a camera, the scene image illustrating the plurality of objects; and executing at least one object classification routine comprising at least one second CNN, each one of the at least one second CNN having at least two convolutional layers, and generating a classification identification of each one of the objects based on the detected object images. The at least one first CNN is trained differently from the at least one second CNN.

[0084] In some embodiments, the method 200 comprises: generating, by an object detection routine comprising at least one first CNN, bounding boxes images of objects located on a scene image, based on the scene image received from a camera and, in some embodiments, an additional data received from additional sensors, the bounding boxes images having a plurality of bounding boxes; assigning a class label to each pixel of each bounding box by a semantic segmentation routine comprising at least two second CNN, and generating a segmented image; and generating category labels assigned to each segmented object 133 of the segmented image 125 by an object classification routine comprising at least one third CNN. The method may further comprise associating each object with classification results obtained by complementary classification units, the association being performed by a sorting decision routine configured to generate a decision matrix which combines the classification results obtained by complementary classification units with the category labels generated by the object classification routine.

[0085] In at least one embodiment, the method 200 as described herein comprises: executing at least one object detection routine 220 comprising at least one first neural network to generate detected object locations representing locations of objects 105 and to generate segmented images 125 of the objects 105 using a scene image 115 captured by a camera 110, the scene image 115 illustrating the objects 105 made of materials; and executing at least one object classification routine 123 comprising at least one second neural network to generate classification results 127 related to materials of the objects 105; and generating a classification identification 251 for each one of the objects 105 based on the classification results 127, the at least one first neural network and the at least one second neural network being different models having different architecture from each other. In at least one embodiment, prior to generating the classification indication 251 , the classificationresults 127 related to each object 105 are mapped with other complementary unit classification results 141 generated by and received from at least one complementary classification unit 140.

[0086] In at least one embodiment, system 100 as described herein comprises: a camera 110 configured to capture scene images 115 of objects 105 made of materials; a display 160; and a processor 150 configured to: execute at least one object detection routine 220 comprising at least one first neural network to generate detected object locations representing locations of objects 105 and to generate segmented images 125 of the objects 105 using the scene image 115 captured by the camera 110, the scene image 115 illustrating the objects 105 made of materials; execute at least one object classification routine 123 comprising at least one second neural network to generate classification results 127 related to materials of the objects 105; and generate a classification identification 251 for each one of the objects 105 based on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.

[0087] While preferred embodiments have been described above and illustrated in the accompanying drawings, it will be evident to those skilled in the art that modifications may be made without departing from this disclosure. Such modifications are considered as possible variants comprised in the scope of the disclosure.

Claims

CLAIMS:

1. A method comprising: executing at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by a camera, the scene image illustrating the objects made of materials; and executing at least one object classification routine comprising at least one second neural network to generate classification results related to materials of the objects; and generating a classification identification for each one of the objects based on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.

2. The method of claim 1 , wherein each one of the at least one first neural network is: a first convolutional neural network (CNN) having at least two convolutional layers, a first vision transformer based (ViT-based) model, a first multilayer perceptron based (MLP-based) model, or a first hybrid model comprising at least two of: elements of the first CNN model, elements of the first ViT-based model, and elements of the first MLP-based model; and each one of the at least one second neural network is: a second CNN having at least two convolutional layers, a second ViT-based model, a second MLP-based model, or a second hybrid model comprising at least two of: elements of the second CNN, elements of the second ViT-based model, and elements of the second MLP-based model.

3. The method of claim 1 , wherein each one of the at least one first neural network is: a first convolutional neural network (CNN) having at least two convolutional layers, a first vision transformer-based (ViT-based) model, a first multilayer perceptron-based (MLP-based) model, a first autoencoder-based model, a first contrastive learning model, a first generative model, or a first hybrid model comprising at least two components selected from: elements of the first CNN model, elements of the first ViT-based model, elements of the first MLP-based model, elements of the first autoencoder-based model, elements of the first contrastive learning model, and elements of the first generative model; andeach one of the at least one second neural network is: a second convolutional neural network (CNN) having at least two convolutional layers, a second vision transformer- based (ViT-based) model, a second multilayer perceptron-based (MLP-based) model, a second autoencoder-based model, a second contrastive learning model, a second generative model, or a second hybrid model comprising at least two components selected from: elements of the second CNN model, elements of the second ViT-based model, elements of the second MLPbased model, elements of the second autoencoder-based model, elements of the second contrastive learning model, and elements of the second generative model.

4. The method of claim 3, wherein the first generative model is a first generative adversarial network (GAN) or a first diffusion-based model, and the second generative model is a second GAN or a second diffusion-based model.

5. The method of any one of claims 1 to 4, wherein the at least one first neural network is a first machine learning model configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a self-supervised learning mode, or a semi-supervised learning mode; and wherein the at least one second neural network is a second machine learning model configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a selfsupervised learning mode, or a semi-supervised learning mode.

6. The method according to any one of claims 1 to 5, wherein the at least one object detection routine comprises a bounding boxes generation routine, and generating detected object locations comprises generating bounding boxes coordinates of bounding boxes surrounding the objects by the bounding boxes generation routine.

7. The method according to any one of claims 1 to 6, wherein the at least one object detection routine comprises a semantic segmentation routine configured to generate polygon coordinates of detected objects polygons representing the objects, and generating the detected object locations comprises generating the polygon coordinates.

8. The method according to claim 7, wherein the detected objects polygons representing the detected objects are generated within the delimitations of the bounding boxes generated by the bounding boxes generation routine.

9. The method of any one of claims 1 to 8, further comprising, prior to generation of the classification identification, generating a sorting decision using the classification results and complementary classification results received from at least one complementary classification unit.

10. The method of any one of claims 1 to 9, wherein sensor data is generated and captured at the time of acquisition of the scene image, and the sensor data being used as input to the object detection routine and the object classification routine.

11. The method of claim 10, wherein the sensor data is generated by at least one of a near-infrared (NIR) system, a short-wave infrared (SWIR) system.

12. The method of any one of claims 10 or 11 , wherein the sensor data is generated by a laser sensor.

13. The method of any one of claims 10 to 12, wherein the sensor data is generated by a volumetric sensor.

14. The method of any one of claims 10 to 13, wherein the sensor data is generated by a point measurement system for visible spectroscopy.

15. The method of any one of claims 10 to 14, wherein the sensor data is generated by a middle wavelength infrared (MWIR) system.

16. The method of any one of claims 10 to 15, wherein the sensor data is generated by a radiography X-ray system.

17. The method of any one of claims 10 to 16, wherein the sensor data is generated by a fluoroscopy X-ray system.

18. The method of any one of claims 10 to 17, wherein the sensor data is generated by a thermal camera.

19. The method of any one of claims 10 to 18, wherein the sensor data is generated by a visible detector.

20. The method of any one of claims 10 to 19, wherein the sensor data is generated by an invisible marker detector.

21. The method of claim 10, wherein the sensor data is generated by at least one of a laser sensor, a volumetric sensor, a point measurement system for visible spectroscopy, a near-infrared (NIR) system, a short-wave infrared (SWIR) system, a middle wavelength infrared (MWIR) system, a radiography X-ray system, a fluoroscopy X-ray system, a thermal camera, a visible detector, or an invisible marker detector.

22. The method according to any one of claims 1 to 21, wherein the classification results related to each object are mapped with other complementary unit classification results generated by and received from at least one complementary classification unit.

23. The method according to any one of claims 1 to 22, wherein at least one of the detected object locations, the segmented images and the classification results are further used to track items within a plurality of frames.

24. The method according to any one of claims 1 to 23, wherein the classification results are used to generate breakdown statistics of the classified objects.

25. The method according to any one of claims 1 to 24, wherein the classification results are further used to estimate weight of the classified objects.

26. The method according to any one of claims 1 to 25, wherein the classification results are further used in instructions to perform a mechanical sorting operation on each one or a subset of the classified objects.

27. The method according to any one of claims 1 to 26, wherein the classification results are transmitted to and used as an input and instructions in other systems to adjust operation of the other systems within a sorting facility.

28. The method of any one of claims 1 to 27, wherein the classification identification is further used to generate instructions transmitted to a targeting system.

29. The method of any one of claims 1 to 28, wherein the camera comprises an RGB camera or a grayscale camera.

30. A system comprising: a camera configured to capture scene images of objects made of materials; a display; and a processor configured to: execute at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by the camera, the scene image illustrating the objects made of materials; execute at least one object classification routine comprising at least one second neural network to generate classification results related to materials of the objects; and generate a classification identification for each one of the objects based on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.

31. The system of claim 30, wherein each one of the at least one first neural network is: a first convolutional neural network (CNN) having at least two convolutional layers, a first vision transformer based (ViT-based) model, a first multilayer perceptron based (MLP-based) model; or a first hybrid model comprising at least two of: elements of the first CNN model, elements of the first ViT-based model, and elements of the first MLP-based model; and each one of the at least one second neural network is: a second CNN having at least two convolutional layers, a second ViT-based model, a second MLP-based model, or a second hybridmodel comprising at least two of: elements of the second CNN, elements of the second ViT-based model, and elements of the second MLP-based model.

32. The system of claim 30, wherein each one of the at least one first neural network is: a first convolutional neural network (CNN) having at least two convolutional layers, a first vision transformer-based (ViT-based) model, a first multilayer perceptron-based (MLP-based) model, a first autoencoder-based model, a first contrastive learning model, a first generative model, or a first hybrid model comprising at least two components selected from: elements of the first CNN model, elements of the first ViT-based model, elements of the first MLP-based model, elements of the first autoencoder-based model, elements of the first contrastive learning model, and elements of the first generative model; and each one of the at least one second neural network is: a second convolutional neural network (CNN) having at least two convolutional layers, a second vision transformer- based (ViT-based) model, a second multilayer perceptron-based (MLP-based) model, a second autoencoder-based model, a second contrastive learning model, a second generative model; or a second hybrid model comprising at least two components selected from: elements of the first CNN model, elements of the first ViT-based model, elements of the first MLP-based model, elements of the first autoencoder-based model, elements of the first contrastive learning model, and elements of the first generative model.

33. The method of claim 32, wherein the first generative model is a first generative adversarial network (GAN) or a first diffusion-based model, and the second generative model is a second GAN or a second diffusion-based model.

34. The method of any one of claims 30 to 33, wherein the at least one first neural network is a first machine learning model configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a self-supervised learning mode, or a semi-supervised learning mode; and wherein the at least one second neural network is a second machine learning model configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a selfsupervised learning mode, or a semi-supervised learning mode.

35. The system of any one of claims 30 to 34, wherein the at least one object detection routine comprises a bounding boxes generation routine, and generating detected object locations comprisesgenerating bounding boxes coordinates of bounding boxes surrounding the objects by the bounding boxes generation routine.

36. The system of any one of claims 30 to 35, wherein the at least one object detection routine comprises a semantic segmentation routine configured to generate polygon coordinates of detected objects polygons representing the objects, and generating the detected object locations comprises generating the polygon coordinates.

37. The system of any one of claims 30 to 36, wherein the detected objects polygons representing the detected objects are generated within the delimitations of the bounding boxes generated by the bounding boxes generation routine.

38. The system of any one of claims 30 to 37, further comprising, prior to generation of the classification identification, generating a sorting decision using the classification results and complementary classification results received from at least one complementary classification unit.

39. The system of any one of claims 30 to 38, wherein sensor data is generated and captured at the time of acquisition of the scene image, and the sensor data being used as input to the object detection routine and the object classification routine.

40. The system of claim 39, wherein the sensor data is generated by at least one of a near-infrared (NIR) system, a short-wave infrared (SWIR) system.

41. The system of any one of claims 39 or 40, wherein the sensor data is generated by a laser sensor.

42. The system of any one of claims 39 to 41 , wherein the sensor data is generated by a volumetric sensor.

43. The system of any one of claims 39 or 42, wherein the sensor data is generated by a point measurement system for visible spectroscopy.

44. The system of any one of claims 39 or 43, wherein the sensor data is generated by a middle wavelength infrared (MWIR) system.

45. The system of any one of claims 39 or 44, wherein the sensor data is generated by a radiography X-ray system.

46. The system of any one of claims 39 or 45, wherein the sensor data is generated by a fluoroscopy X-ray system.

47. The system of any one of claims 39 or 46, wherein the sensor data is generated by a thermal camera.

48. The system of any one of claims 39 or 47, wherein the sensor data is generated by a visible detector.

49. The system of any one of claims 39 or 48, wherein the sensor data is generated by an invisible marker detector.

50. The method of claim 39, wherein the sensor data is generated by at least one of a laser sensor, a volumetric sensor, a point measurement system for visible spectroscopy, a near-infrared (NIR) system, a short-wave infrared (SWIR) system, a middle wavelength infrared (MWIR) system, a radiography X-ray system, a fluoroscopy X-ray system, a thermal camera, a visible detector, or an invisible marker detector.

51. The system of any one of claims 30 to 50, wherein the classification results related to each object are mapped with other complementary unit classification results generated by and received from at least one complementary classification unit.

52. The system according to any one of claims 30 to 51 , wherein at least one of the detected object locations, the segmented images and the classification results is further used to track items within a plurality of frames.

53. The system according to any one of claims 30 to 52, wherein the classification results are used to generate breakdown statistics of the classified objects.

54. The system according to any one of claims 30 to 53, wherein the classification results are further used to estimate weight of the classified objects.

55. The system according to any one of claims 30 to 54, wherein the classification results are further used in instructions to perform a mechanical sorting operation on each one or a subset of the classified objects.

56. The system according to any one of claims 30 to 55, wherein the classification results are transmitted to and used as an input and instructions in other systems to adjust operation of the other systems within a sorting facility.

57. The system according to any one of claims 30 to 56, wherein the classification identification is further used to generate instructions transmitted to a targeting system.

58. The system according to any one of claims 30 to 57, wherein the camera comprises an RGB camera or a grayscale camera.

59. A method comprising: executing at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by a camera, the scene image illustrating the objects made of materials, each one of the at least one first neural network being: a first convolutional neural network (CNN) having at least two convolutional layers, a first vision transformer based (ViT-based) model, a first multilayer perceptron based (MLP-based) model; ora first hybrid model comprising at least two of: elements of the first CNN model, elements of the first ViT-based model, and elements of the first MLP-based model; and executing at least one object classification routine comprising at least one second neural network to generate classification results related to materials of the objects, each one of the at least one second neural network being: a second CNN having at least two convolutional layers, a second ViT-based model, a second MLP-based model, or a second hybrid model comprising at least two of: elements of the second CNN, elements of the second ViT-based model, and elements of the second MLP-based model; and generating a classification identification for each one of the objects based on the classification results, the at least one first neural network and the at least one second neural network being different models having different architecture from each other.

60. A method comprising: executing at least one object detection routine comprising at least one first neural network to generate detected object locations representing locations of objects and to generate segmented images of the objects using a scene image captured by a camera, the scene image illustrating the objects made of materials, the at least one first machine learning model being: a first convolutional neural network (CNN) having at least two convolutional layers; a first vision transformer-based (ViT-based) model; a first multilayer perceptron-based (MLP-based) model; a first autoencoder-based model; a first contrastive learning model; a first generative model; or a first hybrid model comprising at least two components selected from: elements of the first CNN model, elements of the first ViT-based model, elements of the first MLPbased model, elements of the first autoencoder-based model, elements of the first contrastive learning model, and elements of the first generative model; executing at least one object classification routine comprising at least one second machine learning model to generate classification results related to materials of the objects, each one of the at least one second being: a second convolutional neural network (CNN) having at least two convolutional layers; a second vision transformer-based (ViT-based) model; a second multilayer perceptron-based (MLP-based) model;a second autoencoder-based model; a second contrastive learning model; a second generative model; and a second hybrid model comprising at least two components selected from: elements of the second CNN model, elements of the second ViT-based model, elements of the second MLP-based model, elements of the second autoencoder-based model, elements of the second contrastive learning model, and elements of the second generative model; generating a classification identification for each one of the objects based on the classification results, the at least one first neural network being trained differently from the at least one second neural network.

61. The method of claim 60, wherein the first machine learning model is configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a self-supervised learning mode, or a semi-supervised learning mode and the second machine learning model is configured to operate in at least one of: a supervised learning mode, an unsupervised learning mode, a selfsupervised learning mode, or a semi-supervised learning mode.

Citation Information

Patent Citations

  • Systems and methods for optical material characterization of waste materials using machine learning

    US20190130560A1

  • Neural network for bulk sorting

    US20230011383A1

  • Defect detection using one or more neural networks

    US20230125477A1

  • Waste management system

    WO2023229538A1

  • System labeling objects on a conveyor using machine vision monitoring and data combination

    WO2024026562A1