Visible and near infrared image analysis system
The system enhances object classification accuracy and reduces computational resources by using synchronized hyperspectral and color image capture devices with separate neural networks, addressing the inefficiencies of existing systems.
Patent Information
- Authority / Receiving Office
- RU · RU
- Patent Type
- Patents
- Current Assignee / Owner
- OBSHCHESTVO S OGRANICHENNOJ OTVETSTVENNOSTYU VINGZBI
- Filing Date
- 2025-11-05
- Publication Date
- 2026-07-02
Smart Images

Figure 00000001_ABST
Abstract
Description
[0001] FIELD OF TECHNOLOGY
[0002] The invention relates to the field of computer technology, specifically to a system for analyzing images in the visible and near-infrared ranges using neural networks. The present invention can be used for sorting objects in various industrial sectors, such as waste processing, mining, food processing, metallurgy, etc.
[0003] STATE OF THE ART
[0004] Object analysis technology using hyperspectral and RGB images combined with neural networks overcomes the critical limitations of traditional systems that rely solely on RGB visualization by capturing a broader range of information about object properties. RGB images focus on the color and shape of an object, while hyperspectral images focus on the spectral characteristics of the material. However, algorithms and models for jointly processing hyperspectral and RGB images are often excessively complex to achieve high object classification accuracy, resulting in significant computational costs.
[0005] From patent application US 2023214982 A1, published 06.07.2023, class G06T 7 / 00; G06V 20 / 68, a system is known for determining food quality indicators using image data, such as time-lapse RGB, hyperspectral, thermal and / or multispectral images. The method may include receiving image data of food products from imaging devices, performing object detection on these images to identify a bounding box around each food product, and determining a quality indicator of each food product by applying trained models to the bounding boxes. The models were trained using training data from images of other food products that were annotated based on previous determinations of a first portion of the other food products as having low quality indicators and a second portion as having good quality indicators.Other food products and the food products themselves are classified as the same type. The overall quality of each food product is assessed using a quality assessment module. Any combination of rules and / or thresholds can be used to determine the overall quality score of the analyzed food product based on quality metrics determined by trained modules. However, the existing system requires extensive computational resources to determine food quality metrics due to the large number of computational modules.
[0006] In patent application US 2024125712 A1, published 18.04.2024, class G01N 21 / 17, G01N 21 / 27, G01N 21 / 88, G01N 21 / 94, G01N 33 / 44, detection of macroplastics and microplastics is carried out based on RGB and hyperspectral images. The detection is carried out as follows: using a high-resolution color image scanner and a hyperspectral camera, RGB and hyperspectral images are obtained, the RGB and hyperspectral images are combined using the Gram-Schmidt method, and based on the combined images, automatic classification and identification of macroplastics and microplastics is carried out using a supervised classification model trained on the combined images. However, the Gram-Schmidt method requires large computational resources to fuse RGB and hyperspectral images, and the classification accuracy of the supervised classification model is only 90%.
[0007] The closest analogue is a system for identifying the type of material of an object using several types of sensors (patent application US 2023196132 A1, published 06 / 22 / 2023, class G06N 5 / 02, G06V 10 / 58, G06V 10 / 764, G06V 10 / 774).A known system comprises a processor configured to: obtain a machine learning model, in which the machine learning model was trained using training data containing image frames showing a set of objects obtained using a visual sensor (e.g., a camera), and in which the image frames showing the set of objects are associated with material characteristic labels that are determined based on non-visual sensors (e.g., a hyperspectral camera) corresponding to the set of objects; receive a signal from a visual vision sensor; receive a signal from a non-visual sensor; and use the machine learning model, the visual sensor signal, and the non-visual sensor signal to determine the type of material associated with the object.However, training models that classify images with labels can be challenging due to the need to simultaneously predict multiple labels and the need for a large training dataset to ensure high classification accuracy, which also requires significant computational resources. Furthermore, when classifying material types, visual sensor data is prioritized, focusing on identifying object characteristics based on their color and shape rather than spectral properties, which can also reduce classification accuracy.
[0008] The technical problem is to eliminate the above deficiencies.
[0009] DISCLOSURE OF THE INVENTION
[0010] The technical problem solved by the present invention consists in developing a system for analyzing images in the visible and near infrared ranges using neural network models, which classifies objects with high accuracy and saves computing resources.
[0011] The technical result achieved by the present invention consists in increasing the accuracy of classifying objects and saving computing resources.
[0012] The above technical result is achieved by an image analysis system containing at least the following:
[0013] a hyperspectral image capturing device and a color image capturing device configured to capture at least one hyperspectral image and at least one color image of at least one object;
[0014] at least one processor, associated with the hyperspectral image capturing device and the color image capturing device and configured to perform operations in accordance with instructions stored in at least one memory, wherein the operations include at least the following:
[0015] analyzing the obtained at least one hyperspectral image using at least one trained neural network for analyzing hyperspectral images; and if hyperspectral data is determined in the analysis process, then creating a hyperspectral classification mask of an object based on the determined hyperspectral data;
[0016] under the condition of creating a hyperspectral object classification mask, analyzing the obtained at least one color image using at least one trained neural network for analyzing color images; and if color data is determined in the analysis process, then creating a color object classification mask based on the determined color data;
[0017] under the condition of creating a color mask for classifying an object, superimposing the created hyperspectral and color masks on each other; and if the hyperspectral and color masks for classifying an object intersect, then classifying at least one object.
[0018] Improved object classification accuracy and reduced computational resources are achieved by processing hyperspectral and color images using different neural networks. This allows for the specific characteristics of hyperspectral and color images to be taken into account during processing, thereby increasing accuracy. Computing resources are also saved because the neural networks process images that do not contain labels, etc. Hyperspectral data is prioritized during classification, allowing for more accurate object classification, which also improves classification accuracy. Furthermore, the binary-based image processing algorithm also improves classification accuracy and reduces computational resources.
[0019] The hyperspectral imaging device may be a hyperspectral camera with an operating range of 900-1700nm.
[0020] The color image capture device can be an RGB camera with an operating range of 400-800nm.
[0021] The neural network for hyperspectral image analysis can be a neural network based on a fully connected multilayer architecture.
[0022] The neural network for color image analysis can be a neural network based on the cascaded encoder-decoder architecture.
[0023] The hyperspectral image capturing device and the color image capturing device may be configured to synchronizedly capture at least one hyperspectral and at least one color image of at least one object.
[0024] Preferably, the training of neural networks is carried out separately on the basis of pairs of images of objects, wherein each pair of images of at least one object contains a hyperspectral image and a color image of at least one object obtained using a hyperspectral image capture device and a color image capture device, wherein the training of the neural network for analyzing hyperspectral images is carried out on the basis of hyperspectral images from each pair, and the training of the neural network for analyzing color images is carried out on the basis of color images from each pair.
[0025] The hyperspectral mask may be a material classification mask of the object, and the color mask may be a color classification mask of the material of the object.
[0026] The hyperspectral image capturing device and the color image capturing device can be configured to be installed in a sorting zone on a production line along which objects to be sorted are transported, wherein the system is configured to generate a control signal that is transmitted to actuators for sorting at least one object, wherein the system is configured to skip the object without sorting if the hyperspectral data is not determined or if the color data is not determined or if the hyperspectral and color classification masks of the object do not intersect.
[0027] The control signal may contain at least a material, a color of the material, and coordinates of the object.
[0028] BRIEF DESCRIPTION OF DRAWINGS
[0029] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the foregoing general description of the invention and the following detailed description of the embodiments, serve to explain the principles of the present invention.
[0030] The attached drawings are presented to explain the essence of the invention and in no way limit other, particular embodiments of its implementation that do not go beyond the scope of the requested scope of legal protection and are obvious to a specialist in this field of technology.
[0031] The present invention is illustrated by figures 1-8, which show:
[0032] Fig. 1 illustrates a block diagram of joint processing of hyperspectral and color images according to the present invention;
[0033] Fig. 2 illustrates a block diagram of training a neural network for analyzing hyperspectral images and a neural network for analyzing color images according to the present invention;
[0034] Fig. 3 illustrates an example of application of the visible and near infrared image analysis system according to the present invention for sorting objects;
[0035] Fig. 4 illustrates an example of the arrangement of a hyperspectral camera and an RGB camera inside a cabinet;
[0036] Fig. 5 illustrates an example of a hyperspectral image (left) and a neural network prediction (right);
[0037] Fig. 6 illustrates an example of RGB image segmentation;
[0038] Fig. 7 illustrates an example of a neural network architecture for processing hyperspectral camera images;
[0039] Fig. 8 illustrates an example of the architecture of a neural network for processing RGB camera images.
[0040] The figures in the drawings indicate: 1 - a cabinet in which the hyperspectral camera and the RGB camera are located; 2 - a cabinet in which the computing device is located; 3 - a conveyor belt; 4 - an RGB camera; 5 - a hyperspectral camera; 6 - a hyperspectral image; 7 - material prediction by a neural network for processing hyperspectral data; 8 - an RGB image; 9 - a segmented image (material prediction by a neural network for processing RGB images); 10 - input data of the neural network for processing hyperspectral data (for example, a hyperspectral image); 11 - hidden layers of the neural network for processing hyperspectral images; 12 - output data of the neural network for processing hyperspectral images (for example, material prediction); 13 - the structure of the encoder of the neural network for processing color images; 14 - the structure of the decoder of the neural network for processing color images.
[0041] IMPLEMENTATION OF THE INVENTION
[0042] The detailed description of the embodiment of the invention provides numerous implementation details to ensure a clear understanding of the present invention. However, it is obvious to those skilled in the art how the present invention can be used with or without these implementation details. Furthermore, the present invention is not limited to the described implementation. Numerous possible modifications, changes, variations, and substitutions, while maintaining the spirit and form of the present invention, are obvious to those skilled in the art.
[0043] The image analysis system according to the present invention comprises a hyperspectral image capture device and a color image capture device, configured to capture hyperspectral and color images of various objects, such as plastic products, glass fragments, metal parts, rocks, etc. The hyperspectral image capture device is, for example, a hyperspectral camera. The color image capture device is, for example, an RGB camera. The hyperspectral image capture device and the color image capture device are connected to at least one processor, which is configured to perform operations in accordance with instructions stored in at least one memory.The operations include at least the following: analyzing the obtained at least one hyperspectral image using at least one trained neural network for analyzing hyperspectral images; and if hyperspectral data is determined during the analysis, then creating a hyperspectral mask based on the determined hyperspectral data; provided that the hyperspectral mask is created, analyzing the obtained at least one color image using at least one trained neural network for analyzing color images; and if color data is determined during the analysis, then creating a color mask based on the determined color data; provided that the color mask is created, superimposing the created hyperspectral and color masks on each other; and if the hyperspectral mask and the color mask intersect, then classifying at least one object.A hyperspectral image capture device and a color image capture device may be configured to synchronize the capture of hyperspectral and color images. The hyperspectral mask may be a mask for classifying the object's material, and the color mask may be a mask for classifying the color of the object's material. Fig. 1 illustrates an algorithm for jointly processing hyperspectral and color images according to the present invention, which is used for sorting objects.
[0044] The neural network for hyperspectral image analysis can be a neural network of any architecture that is specialized in processing hyperspectral images, such as a neural network based on a fully connected multilayer architecture.
[0045] The neural network for color image analysis can be a neural network of any architecture that specializes in processing color images, such as a neural network based on a cascaded encoder-decoder architecture.
[0046] Training of neural networks is carried out separately on the basis of pairs of images of objects, wherein each pair of images of at least one object contains a hyperspectral image and a color image of at least one object obtained using a hyperspectral image capture device and a color image capture device, wherein training of the neural network for analyzing hyperspectral images is carried out on the basis of hyperspectral images from each pair, and training of the neural network for analyzing color images is carried out on the basis of color images from each pair.
[0047] Fig. 2 shows a block diagram of training a neural network for analyzing hyperspectral images and a neural network for analyzing color images according to the present invention.
[0048] The process is an iterative training cycle, beginning with initial data collection. The "Data" stage generates a primary dataset, including paired images obtained from a hyperspectral image capture device and a color image capture device, collected in real-world production conditions.
[0049] Data processing and preparation includes a set of procedures to ensure the quality of the training sample. Hyperspectral data undergoes noise correction using background subtraction, while color images undergo augmentation. This stage also includes labeling and database creation.
[0050] "Creating Train / Val / Test Databases" involves dividing the data into three independent sets. The training set (Train) is used for training the models, the validation set (Val) is used to control overfitting, and the test set (Test) is used for final evaluation of the model's quality on previously unseen data.
[0051] "Neural network training" is performed separately for the hyperspectral image processing model and the color image processing model using specialized architectures. The process includes iterative optimization of weight coefficients and regular monitoring of metrics on the validation set.
[0052] "Testing" is a comprehensive assessment of the quality of trained models on an independent test set.
[0053] "Importing into a usable format" is the final step, where trained models are converted into a standardized format for subsequent integration into the production system. This ensures compatibility with various computing platforms and optimized performance.
[0054] "Error detection" is a key element of the process. If insufficient model quality or specific errors are detected during the testing phase, the process returns to the data processing and preparation phase. This allows for supplementing the training set with problematic cases, adjusting the augmentation, or revising the labeling approach, after which the entire cycle is repeated until the required accuracy is achieved.
[0055] Below is presented an embodiment of the present invention for sorting objects on a production line, which should not be used as limiting other, particular embodiments of the implementation of the present invention that do not go beyond the scope of the requested scope of legal protection and are obvious to a person skilled in the art.
[0056] The sorting system is a single hardware and software complex with a single control interface (Fig. 3). The system contains two synchronized cameras (with synchronized capture): an RGB camera (4) and a hyperspectral camera (5), covering the visible (400-800 nm) and near infrared (900-1700 nm) spectra (Fig. 4). Cameras (4) and (5) are installed in the sorting zone near the conveyor belt (3), along which various objects to be sorted are transported: plastic products, glass fragments, metal parts, rocks, etc. The cameras are placed in a cabinet (1), which is installed in such a way that the cameras can scan the objects that are transported along the conveyor belt (3). The chambers (4) and (5) are connected to a computing device, which is located in the cabinet (2) and is intended for processing data according to the present invention.The hyperspectral camera (5), operating in the 900-1700 nm range, scans each object and, using a trained neural network for processing hyperspectral images (Fig. 7), determines the type of material based on spectral analysis; for example, the color of the image (on the right) indicates the predicted material (in this case, ABS plastic) (Fig. 5). At the same time, the RGB camera (4) records the color characteristics of the object (Fig. 6) in the visible range of 400-800 nm using a trained neural network for processing RGB images (Fig. 8).
[0057] The data processing method is characterized by a cascade neural network architecture for RGB data and a fully connected 4-layer neural network for hyperspectral data, a joint analysis algorithm, a decision-making method, and the ability to work in real time.
[0058] The neural network training method is characterized by the use of pixel-level labeling of hyperspectral data, background subtraction, and similarity search techniques for the hyperspectral model, as well as pixel-level labeling for RGB images. The training process also utilizes augmentation—artificially expanding the training dataset by creating modified copies of existing images.
[0059] During system operation, objects are processed as follows. The hyperspectral camera begins recording data in 256 spectral channels with a step size of approximately 4 nm (6). Simultaneously, the RGB camera captures a color image (8) with a resolution of 3840×2160 pixels. The cameras are synchronized via a hardware system. The acquired data is sent to the computing device (2), where it undergoes preliminary processing.
[0060] The system's data processing algorithm combines hyperspectral and RGB image analysis using neural network models, ensuring highly accurate material classification. The basic operating principle is based on the prioritization of hyperspectral data, followed by refinement of color characteristics through RGB analysis (Fig. 1).
[0061] The hyperspectral processing model is a neural network with four hidden layers (11) (Fig. 7), containing approximately 280,000 trainable parameters. The model receives spectral data (10) in the 900-1700 nm range, represented by 256 spectral channels, as input. Before processing, the data undergoes a pre-processing step, including normalization and noise removal.
[0062] The RGB model (Fig. 8) is built on the basis of a deep encoder (13) and a decoder (14) with five layers and contains about 14 million parameters. The deep encoder is a component that converts the input data sequence (e.g., an image) into a numerical representation). The decoder is a component that converts the encoded representation into an output data sequence (e.g., back to an image). It processes standard three-channel color images (8) in the visible range of 400-800 nm. The result is a segmented image (9).
[0063] The algorithm for joint processing of hyperspectral images and RGB images (Fig. 1) in real time combines data from both cameras (4) and (5), creating a comprehensive characteristic of the object - both in material and color.
[0064] "HSC and RGB Data Capture." The hyperspectral camera (HSC) captures images in the 900-1700 nm range with a resolution of 256 spectral channels. The RGB camera simultaneously captures color images in the visible range of 400-800 nm with a spatial resolution of 3840×2160 pixels.
[0065] "GSC Mask Generation." Spectral data undergoes preprocessing, including normalization and noise correction using background subtraction. A neural network model with four hidden layers analyzes the spectral characteristics of each pixel and generates a material classification mask. The mask highlights areas corresponding to different material types (PET, PVC, PP, etc.).
[0066] "Creating an RGB mask." A color image is processed by an encoder-decoder neural network model. The encoder extracts spatial features of the image, and the decoder generates semantic segmentation. The result is a color feature mask that identifies regions by their color parameters in the CIELAB color space.
[0067] "Mask Overlay." Spatially superimposes the GSK and RGB masks. The overlay algorithm combines material information from the GSK mask and color characteristics from the RGB mask.
[0068] "Control signal generation." Based on the resulting mask, control signals for sorting mechanisms are generated.
[0069] "Sorting mechanisms." Actuators receive control signals and perform physical separation of objects.
[0070] The algorithm operates on a binary principle: if the data analysis is successful, the system proceeds to the next stage; if not, the process is terminated. The system begins with analysis of hyperspectral camera data. If spectral analysis fails to classify the material, the object is marked as unidentified, and further processing is terminated without generating a control signal. If the material is successfully identified using hyperspectral data, a corresponding classification mask is created, after which RGB data analysis is activated. If RGB analysis fails to determine the color characteristics of the object, processing is terminated, and no control signal is generated. If color characteristics are successfully determined, an RGB data mask is created. The HSC and RGB masks are then superimposed. If the masks do not overlap, the object is considered unidentified, and no control signal is generated.Only if all previous steps—material identification using GSK data, color characterization using RGB data, and mask intersection—are successfully completed does the system generate a control signal for the actuators (sorting mechanisms). Depending on the settings, the sorting mechanisms can include pneumatic nozzles, mechanical pushers, robotic manipulators, and so on, which direct objects into the correct container. The control signal includes the precise coordinates of the object, the material's classification parameters, and its color characteristics, ensuring precise positioning and sorting of the object into the appropriate category. The sorting system can distinguish, for example, transparent PET plastic from PVC, even if they appear identical, and accurately determine the color of painted parts.
[0071] The sorting system operates in real time at speeds up to 3 m / s, achieving a sorting accuracy of more than 98%, while the number of objects on the conveyor does not affect the speed and accuracy of image processing.
[0072] A computing device capable of processing the data necessary for implementing the present invention generally comprises components such as one or more processors, at least one memory, a data storage medium, interfaces, input / output means, and networking means. When executing machine-readable instructions contained in the RAM, one or more processors are configured to perform the basic computing operations necessary for the operation of the system or the functionality of one or more of its components. The memory is typically implemented in the form of RAM, into which the necessary program logic is loaded to provide the desired functionality. When implementing the present invention, the memory volume necessary to perform the operations according to the claimed solution is allocated. The data storage medium may be implemented in the form of an HDD, SSD disks, RAID array, network storage, flash memory, etc.The means enables long-term storage of various types of information. The interfaces are standard means for connecting and operating peripheral and other devices, such as USB, RS232, RJ45, COM, HDMI, PS / 2, Lightning, etc. A keyboard, joystick, display (touch screen), projector, touchpad, mouse, trackball, light pen, speakers, microphone, etc. can be used as data input / output means in any embodiment of the system according to the present invention. The network interaction means are selected from devices that provide network reception and transmission of data, such as an Ethernet card, WLAN / Wi-Fi module, Bluetooth module, BLE module, NFC module, IrDa, RFID module, GSM modem, etc. The network interaction means enable data exchange via a wired or wireless data transmission channel, such as a WAN, PAN, LAN, Intranet, Internet, WLAN, WMAN or GSM, etc.The device components are connected via a common data bus.
[0073] These application materials present a preferred disclosure of the embodiment of the present invention, which should not be used as limiting other, particular embodiments of its implementation that do not go beyond the scope of the requested scope of legal protection and are obvious to a person skilled in the art.
[0074] It should be clear to a person skilled in the art that various variations of the disclosed technical solution do not change the essence of the invention, but only determine its specific embodiments and applications.
Claims
1. An image analysis system comprising at least the following: a hyperspectral image capture device and a color image capture device configured to capture at least one hyperspectral image and at least one color image of at least one object; at least one processor, coupled to the hyperspectral image capture device and the color image capture device and configured to perform operations in accordance with instructions stored in at least one memory, wherein the operations include at least the following: analyzing the obtained at least one hyperspectral image using at least one trained neural network for analyzing hyperspectral images; and if hyperspectral data is determined during the analysis, then creating a hyperspectral mask for classifying an object based on the determined hyperspectral data; subject to the creation of a hyperspectral mask for classifying an object, analyzing the obtained at least one color image using at least one trained neural network for analyzing color images; and if color data is determined during the analysis, then creating a color mask for classifying an object based on the determined color data; provided that a color mask for classifying an object is created, superimposing the created hyperspectral and color masks on each other; and if the hyperspectral and color masks for classifying an object intersect, then classifying at least one object.
2. The system according to claim 1, characterized in that the device for capturing hyperspectral images is a hyperspectral camera with an operating range of 900-1700 nm.
3. The system according to claim 1, characterized in that the color image capture device is an RGB camera with an operating range of 400-800 nm.
4. The system according to claim 1, characterized in that the neural network for analyzing hyperspectral images is a neural network based on a fully connected multilayer architecture.
5. The system according to claim 1, characterized in that the neural network for analyzing color images is a neural network based on a cascade encoder-decoder architecture.
6. The system according to claim 1, characterized in that the hyperspectral image capture device and the color image capture device are configured to synchronize the capture of at least one hyperspectral and at least one color image of at least one object.
7. The system according to claim 1, characterized in that the training of the neural networks is carried out separately on the basis of pairs of images of objects, wherein each pair of images of at least one object contains a hyperspectral image and a color image of at least one object obtained using a hyperspectral image capture device and a color image capture device, wherein the training of the neural network for analyzing the hyperspectral images is carried out on the basis of the hyperspectral images from each pair, and the training of the neural network for analyzing the color images is carried out on the basis of the color images from each pair.
8. The system according to claim 1, characterized in that the hyperspectral classification mask of the object is a classification mask of the material of the object, and the color classification mask of the object is a classification mask of the color of the material of the object.
9. The system according to claim 1, characterized in that the hyperspectral image capturing device and the color image capturing device are configured to be installed in a sorting zone on a production line along which objects to be sorted are transported, and the system is configured to generate a control signal that is transmitted to actuators for sorting at least one object, and the system is configured to skip the object without sorting if the hyperspectral data is not determined, or if the color data is not determined, or if the hyperspectral and color classification masks of the object do not intersect.
10. The system according to claim 9, characterized in that the control signal contains at least the material, the color of the material, and the coordinates of the object.