Target object classification method and device based on hyperspectral information, equipment and medium

By directly analyzing the spectral characteristics of objects from sparse sampling of hyperspectral signals, and using affine transformation alignment and instance segmentation models for feature extraction, the problems of large amount of hyperspectral data and complex reconstruction algorithms in the prior art are solved, real-time high-efficiency spectral analysis and fast object classification are achieved.

CN120182731AInactive Publication Date: 2025-06-20NANJING UNIV

Patent Information

Application Number
CN202510661863.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the high-spectral data volume is huge and the reconstruction algorithm is complex, resulting in long processing time of algorithms and low task execution efficiency, which cannot meet the needs of real-time imaging and analysis in visual classification tasks such as industrial sorting.

Method used

By directly analyzing the spectral characteristics of the object from sparse sampling of the signal, the affine transformation alignment and instance segmentation model are used to extract the contour area of ​​the target object, thereby obtaining the spectral characteristics of the target object and using the classification model for feature classification.

Benefits of technology

It realizes efficient spectral analysis of real-time acquisition and accelerated processing, can quickly classify target objects, save algorithm processing time, and improve the execution efficiency of industrial sorting and other visual classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182731A_ABST
    Figure CN120182731A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral information-based target object classification method, apparatus and device, and a medium. The method comprises the steps of obtaining hyperspectral information of sparse sampling points of a scene where a to-be-classified target object is located; obtaining visible light information of a scene where the to-be-classified target object is located; performing affine transformation alignment on the hyperspectral information and the visible light information; extracting a contour region of a to-be-classified target object from the visible light information by using an instance segmentation model; acquiring a spectral curve of each sparse sampling point in the contour area from the hyperspectral information according to the contour area of the target object and an affine transformation alignment result; and obtaining spectral features of the target object according to the spectral curves of all the sparse sampling points in the contour region, and performing feature classification on the target object by using a classification model based on the spectral features. The target objects can be quickly classified, the processing time is saved, and the execution efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method, device, equipment, and medium for classifying target objects based on hyperspectral information. Background Art

[0002] In today's hyperspectral imaging applications, in order to obtain a complete hyperspectral image, a scanning method is often used to acquire the three-dimensional hyperspectral data of the entire scene, and then the spectral characteristics of the region of interest are analyzed. The scanning process requires sacrificing time, and the hyperspectral data volume is huge. The subsequent transmission, storage, and analysis processes also consume high costs.

[0003] For the spectral video acquisition method of dynamic scenes, a snapshot spectral imaging technology based on computational photography theory has been developed. However, this method requires dimensionality reduction acquisition first and then computational reconstruction to recover the complete high-dimensional spectral data. Although this method can shoot at video rate, the subsequent reconstruction algorithm has high complexity and cannot meet the application requirements of real-time imaging and analysis in visual classification tasks such as industrial sorting. How to achieve rapid classification of target objects, real-time acquisition, and rapid and efficient spectral analysis is an urgent problem to be solved currently. Summary of the Invention

[0004] In view of the above problems in the prior art, the present invention provides a method, device, equipment, and medium for classifying target objects based on hyperspectral information, which directly analyzes the spectral characteristics of an object from the sparse sampling of signals, realizes real-time acquisition and accelerated processing of efficient spectral analysis, can also rapidly classify target objects, and does not require complex spectral acquisition and calculation processes, saving algorithm processing time and improving the execution efficiency of industrial sorting and other visual classification tasks.

[0005] To achieve the above object, the first aspect of the present application provides a method for classifying target objects based on hyperspectral information, including:

[0006] Obtaining the hyperspectral information of sparse sampling points in the scene where the target object to be classified is located; and obtaining the visible light information of the scene where the target object to be classified is located;

[0007] Performing an affine transformation alignment on the hyperspectral information and the visible light information;

[0008] Using an instance segmentation model to extract the contour region of the target object to be classified from the visible light information;

[0009] According to the contour region of the target object and the result of the affine transformation alignment, obtaining the spectral curve of each sparse sampling point within the contour region from the hyperspectral information;

[0010] The spectral characteristics of the target object are obtained from the spectral curves of all sparse sampling points within the contour region, and the target object is classified based on the spectral characteristics using a classification model.

[0011] As a possible implementation manner of the first aspect, obtaining the hyperspectral information of the sparse sampling points in the scene where the target object to be classified is located includes:

[0012] Sampling the optical signal of the scene where the target object is located using a mask to obtain the optical signal of the sparse sampling points;

[0013] Dispersing the optical signal of the sparse sampling points using a prism and restricting the dispersion length through a broadband filter;

[0014] Obtaining the hyperspectral information of the sparse sampling points from the dispersed optical signal.

[0015] As a possible implementation manner of the first aspect, aligning the hyperspectral information and the visible light information through affine transformation includes:

[0016] Performing affine transformation alignment according to pre-calibrated optical mapping parameters; wherein, the optical mapping parameters include: the sampling point coordinates of the mask, the dispersion drift distance corresponding to each spectral segment, and the affine transformation matrix.

[0017] As a possible implementation manner of the first aspect, the instance segmentation model includes a U-net model or a MaskR-CNN model.

[0018] As a possible implementation manner of the first aspect, according to the contour region of the target object and the result of the affine transformation alignment, obtaining the spectral curve of each sparse sampling point within the contour region from the hyperspectral information includes:

[0019] Extracting the coordinates of the sparse sampling points within the contour region of the target object from the hyperspectral information according to the contour region of the target object and the result of the affine transformation alignment;

[0020] Obtaining the spectral curve of each sparse sampling point within the contour region from the hyperspectral information according to the coordinates of the sparse sampling points.

[0021] As a possible implementation manner of the first aspect, obtaining the spectral characteristics of the target object according to the spectral curves of all sparse sampling points within the contour region includes:

[0022] Calculating the average value of the spectral curves of all sparse sampling points within the contour region;

[0023] Use the average value as the spectral feature of the target object.

[0024] As a possible implementation of the first aspect, the classification model includes a support vector machine or a decision tree model.

[0025] The second aspect of the present application provides a target object classification device based on hyperspectral information, including:

[0026] A first acquisition unit, configured to acquire hyperspectral information of sparse sampling points in the scene where the target object to be classified is located; and acquire visible light information of the scene where the target object to be classified is located;

[0027] An alignment unit, configured to perform affine transformation alignment on the hyperspectral information and the visible light information;

[0028] A segmentation unit, configured to use an instance segmentation model to extract the contour region of the target object to be classified from the visible light information;

[0029] A second acquisition unit, configured to acquire the spectral curve of each sparse sampling point within the contour region from the hyperspectral information according to the contour region of the target object and the result of the affine transformation alignment;

[0030] A classification unit, configured to obtain the spectral feature of the target object based on the spectral curves of all sparse sampling points within the contour region, and perform feature classification on the target object using the classification model based on the spectral feature.

[0031] As a possible implementation of the second aspect, the first acquisition unit is configured to:

[0032] Sample the optical signal of the scene where the target object is located using a mask to obtain the optical signal of the sparse sampling points;

[0033] Disperse the optical signal of the sparse sampling points using a prism, and limit the dispersion length through a broadband filter;

[0034] Acquire the hyperspectral information of the sparse sampling points from the dispersed optical signal.

[0035] As a possible implementation of the second aspect, the alignment unit is configured to:

[0036] Perform affine transformation alignment according to pre-calibrated optical mapping parameters; wherein, the optical mapping parameters include: the sampling point coordinates of the mask, the dispersion drift distance corresponding to each spectral segment, and the affine transformation matrix.

[0037] As a possible implementation of the second aspect, the instance segmentation model includes a U-net model or a MaskR-CNN model.

[0038] As a possible implementation of the second aspect, the second acquisition unit is configured to:

[0039] Extract the coordinates of the sparse sampling points within the contour region of the target object from the hyperspectral information according to the contour region of the target object and the result of the affine transformation alignment;

[0040] Obtain the spectral curve of each sparse sampling point within the contour region from the hyperspectral information according to the coordinates of the sparse sampling points.

[0041] As a possible implementation of the second aspect, the classification unit is configured to:

[0042] Calculate the average value of the spectral curves of all the sparse sampling points within the contour region;

[0043] Take the average value as the spectral feature of the target object.

[0044] As a possible implementation of the second aspect, the classification model includes a support vector machine or a decision tree model.

[0045] The third aspect of the present application provides a computing device, including:

[0046] A communication interface;

[0047] At least one processor, which is connected to the communication interface; and

[0048] At least one memory, which is connected to the processor and stores program instructions, and when the program instructions are executed by the at least one processor, the at least one processor executes the method according to any one of the first aspects above.

[0049] The fourth aspect of the present application provides a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a computer, the computer executes the method according to any one of the first aspects above.

[0050] Compared with the prior art, the significant advantages of the present invention are as follows: The present invention directly analyzes the spectral features of an object only from the sparse sampling of signals, realizes efficient spectral analysis with real-time acquisition and accelerated processing, can also quickly classify the target object, and does not require a complex spectral acquisition and calculation process, saving the algorithm processing time and improving the execution efficiency of industrial sorting and other visual classification tasks, thereby solving the technical problems of huge hyperspectral data volume, high complexity of reconstruction algorithms, long algorithm processing time, and low execution efficiency of tasks in the prior art. Description of the Drawings

[0051] The various features of the present invention and the relationships between the various features will be further described below with reference to the accompanying drawings. The accompanying drawings are all exemplary. Some features are not shown in actual proportions, and in some of the accompanying drawings, conventional features in the field related to the present application that are not necessary for the present application may be omitted, or features that are not necessary for the present application may be shown additionally. The combinations of the various features shown in the accompanying drawings are not used to limit the present application. Additionally, throughout this specification, the content referred to by the same reference numerals is also the same. The specific description of the accompanying drawings is as follows:

[0052] Figure 1 Schematic diagram of an embodiment of the method for classifying target objects based on hyperspectral information provided by an embodiment of the present application;

[0053] Figure 2 Schematic diagram of the optical path of the hyperspectral sparse information and visible light information acquisition module for an embodiment of the method for classifying target objects based on hyperspectral information provided by an embodiment of the present application;

[0054] Figure 3 Schematic diagram of instance segmentation for an embodiment of the method for classifying target objects based on hyperspectral information provided by an embodiment of the present application;

[0055] Figure 4 Schematic diagram of instance segmentation for an embodiment of the method for classifying target objects based on hyperspectral information provided by an embodiment of the present application;

[0056] Figure 5 Schematic diagram of an embodiment of the method for classifying target objects based on hyperspectral information provided by an embodiment of the present application;

[0057] Figure 6 Schematic diagram of spectral features for an embodiment of the method for classifying target objects based on hyperspectral information provided by an embodiment of the present application;

[0058] Figure 7 Schematic diagram of quickly classifying target objects for an embodiment of the method for classifying target objects based on hyperspectral information provided by an embodiment of the present application;

[0059] Figure 8 Schematic diagram of an embodiment of the device for classifying target objects based on hyperspectral information provided by an embodiment of the present application;

[0060] Figure 9 Schematic diagram of the computing device provided by an embodiment of the present application. Detailed implementation manners

[0061] The terms "first", "second", "third", etc. or similar terms such as Module A, Module B, Module C, etc. in the specification and claims are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that, where permitted, the specific order or sequence can be interchanged so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0062] In the following description, the reference numerals representing steps, such as S110, S120, etc., do not necessarily mean that the steps will be executed in this order. Where permitted, the order of the front and back steps can be interchanged, or the steps can be executed simultaneously.

[0063] The term "comprising" used in the specification and claims should not be construed as being limited to the content listed thereafter; it does not exclude other elements or steps. Therefore, it should be construed as specifying the presence of the stated features, wholes, steps or components, but does not exclude the presence or addition of one or more other features, wholes, steps or components and their groups. Thus, the expression "a device comprising device A and B" should not be limited to a device consisting only of components A and B.

[0064] As used herein, the term "one embodiment" or "an embodiment" means that the specific features, structures, or characteristics described in connection with that embodiment are included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment" or "in an embodiment" that appear throughout this specification do not necessarily all refer to the same embodiment, but may refer to the same embodiment. In addition, in one or more embodiments, the various specific features, structures, or characteristics can be combined in any suitable manner, as will be apparent to those of ordinary skill in the art from this disclosure.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. In case of inconsistency, the meaning set forth in this specification or the meaning derived from the content recorded in this specification shall prevail. In addition, the terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application. For the purpose of accurately describing the technical content in this application and for accurately understanding the present invention, the following explanations or definitions of the terms used in this specification are given before describing the specific embodiments:

[0066] 1) ROI (Region of Interest): In machine vision and image processing, a region that needs to be processed is outlined in the processed image in the form of a rectangle, circle, ellipse, irregular polygon, etc., and is called the region of interest.

[0067] 2) U-net: The U-net network is an image segmentation network based on CNN (Convolutional Neural Network). The first half (left side) of the U-Net network is used for feature extraction, and the second half (right side) is used for upsampling. In some literature, such a structure is also called the Encoder-Decoder (encoding-decoding) structure. Since the overall structure of this network resembles the capital letter U in English, it is named U-Net.

[0068] 3) Mask R-CNN (Mask Region Convolutional Neural Network): Mask R-CNN is an instance segmentation model that can determine the positions and categories of various objects in an image and give pixel-level predictions. The so-called "instance segmentation" refers to segmenting each object of interest in a scene, regardless of whether they belong to the same category. For example, the model can identify individual objects such as vehicles and people from a street view video. Instance segmentation is a task of identifying the object contours at the pixel level. Mask R-CNN mainly performs segmentation on the basis of object detection. Mask R-CNN is a two-stage framework. In the first stage, the image is scanned and proposals (i.e., regions that may contain an object) are generated. In the second stage, the proposals are classified and bounding boxes and masks are generated.

[0069] 4) SVM (Support Vector Machine): Support Vector Machine is a class of generalized linear classifiers that perform binary classification on data in a supervised learning manner. Its decision boundary is the maximum margin hyperplane solved for the learning samples. SVM uses the hinge loss function to calculate the empirical risk and adds a regularization term to the solution system to optimize the structural risk. It is a classifier with sparsity and robustness.

[0070] 5) Decision tree model: The decision tree model is a tree diagram composed of decision points, policy points (event points), and results. The decision tree model is generally applied to sequential decision-making. Usually, the maximum expected return value or the lowest expected cost is used as the decision criterion. By solving the benefit values of various solutions under different conditions in a graphical way and then comparing them, a decision is made. The decision tree model is a simple and easy-to-use non-parametric classifier. It does not require any prior assumptions about the data, has a fast calculation speed, the results are easy to interpret, and it is robust.

[0071] 6) RGB: The RGB color model is a color standard in the industrial field. By varying the three color channels of red (R), green (G), and blue (B) and their mutual superposition, a variety of colors can be obtained. RGB represents the colors of the three channels of red, green, and blue. This standard covers almost all the colors that the human vision can perceive and is one of the most widely used color systems.

[0072] First, the existing methods will be introduced below, and then the technical solutions of this application will be introduced in detail.

[0073] A spectral image with a spectral resolution in the order of 10λ is called a hyperspectral image (Hyperspectral Image). The hyperspectral imaging system acquires a three-dimensional data cube of the two-dimensional space and one-dimensional spectrum of the scene. The spectral characteristics of each pixel can reveal the inherent properties of the interaction between the light source and the substance, and it has a wide range of application scenarios in agricultural remote sensing, life sciences, and industrial sorting.

[0074] For a long time, simultaneously collecting high-resolution data in the spectral, spatial, and temporal dimensions has been a major challenge in the design of spectral cameras. In today's hyperspectral imaging applications, in order to obtain a complete hyperspectral image, a scanning method is often used to acquire the complete three-dimensional hyperspectral data of the scene, and then the spectral characteristics of the region of interest are analyzed. Its scanning process requires sacrificing time, and the hyperspectral data volume is huge. The subsequent transmission, storage, and analysis processes also require high costs.

[0075] For the method of spectral video acquisition of dynamic scenes, a snapshot spectral imaging technology based on computational photography theory has been developed. However, this method requires first reducing the dimension for acquisition and then computationally reconstructing to restore the complete high-dimensional spectral data. Although this method can capture at video rate, the subsequent reconstruction algorithm has a high complexity, a long algorithm processing time, and a low task execution efficiency. It cannot meet the application requirements of real-time imaging and analysis in visual classification tasks such as industrial sorting. How to quickly classify the target object, collect in real time, and quickly perform efficient spectral analysis is an urgent problem to be solved currently.

[0076] The existing technologies have the following defects: the hyperspectral data volume is huge, the reconstruction algorithm has a high complexity, the algorithm processing time is long, and the task execution efficiency is low.

[0077] Based on the technical problems existing in the above-mentioned prior art, the present application provides a method, device, equipment and medium for classifying target objects based on hyperspectral information, which directly analyzes the spectral characteristics of objects from the sparse sampling of signals, realizes efficient spectral analysis of real-time acquisition and accelerated processing, can also quickly classify target objects, and does not require complex spectral acquisition and calculation processes, saving the algorithm processing time and improving the execution efficiency of industrial sorting and other visual classification tasks, thereby solving the technical problems of huge hyperspectral data volume, high complexity of reconstruction algorithms, long algorithm processing time and low execution efficiency of tasks in the prior art.

[0078] Figure 1 FIG. is a schematic diagram of an embodiment of a method for classifying target objects based on hyperspectral information provided by an embodiment of the present application. As Figure 1 shown, the method may specifically include:

[0079] Step S110, obtaining hyperspectral information of sparse sampling points in the scene where the target object to be classified is located; and obtaining visible light information of the scene where the target object to be classified is located;

[0080] Step S120, performing affine transformation alignment on the hyperspectral information and the visible light information;

[0081] Step S130, using an instance segmentation model to extract the contour area of the target object to be classified from the visible light information;

[0082] Step S140, according to the contour area of the target object and the result of the affine transformation alignment, obtaining the spectral curve of each sparse sampling point in the contour area from the hyperspectral information;

[0083] Step S150, obtaining the spectral characteristics of the target object according to the spectral curves of all sparse sampling points in the contour area, and performing feature classification on the target object based on the spectral characteristics by using a classification model.

[0084] In order to solve the problem of low task execution efficiency caused by huge hyperspectral data volume, an embodiment of the present application provides a high-speed spectral video system for fast classification. Using this system, target spectral classification is performed only from hyperspectral sparse sampling, and reconstruction algorithms do not need to be used, which can meet the requirements of real-time spectral analysis in visual classification tasks.

[0085] In the embodiments of the present application, the high-speed spectral video system for rapid classification may specifically include: a hyperspectral sparse information acquisition module, a visible light information acquisition module, a visible light instance segmentation module, an image registration module, a spectral information extraction module, and a spectral feature classification module. When using this system to perform a visual classification task, the above-mentioned modules can be initialized first. The visual classification of the target object may specifically include the following steps:

[0086] In step S110, a camera device can be used to photograph the target object to obtain the optical signals of the target object to be classified and its surrounding scene. Then, a beam splitter is used to divide the above optical signals into two identical parts, one of which is transmitted to the visible light information acquisition module, and the other is transmitted to the hyperspectral sparse information acquisition module. Among them, the hyperspectral sparse information acquisition module uses a prism mask spectral video imaging system (PMVIS, Prism Mask Multispectral Video Imaging System) to obtain the sparse sampling information of the target object scene spectrum. Specifically, in the PMVIS, a mask can be used to limit the amount of information acquisition data, and spectral information is only collected for evenly distributed acquisition points. Through the above limitations, the amount of hyperspectral information data is greatly reduced, thereby saving processing time and improving execution efficiency. These evenly distributed acquisition points can be called dimensionality reduction sampling points or sparse sampling points.

[0087] On the one hand, the hyperspectral sparse information acquisition module is used to obtain the hyperspectral information of the sparse sampling points in the scene where the target object to be classified is located. On the other hand, the visible light information acquisition module is used to obtain the visible light information of the scene where the target object to be classified is located. That is to say, the visible light information acquisition module can perform synchronous acquisition with the hyperspectral sparse information acquisition module to obtain the visible light information of the target object. In one example, the visible light information may include an RGB image.

[0088] In step S120, the image registration module performs synchronous spatial affine transformation alignment on the two video frames of the hyperspectral information and the visible light information according to the pre-calibrated optical mapping parameters. In one example, the hyperspectral information and the visible light information can be aligned at the pixel level according to the optical calibration information. Through the affine transformation alignment, a one-to-one correspondence can be established between the pixels in the hyperspectral information and the visible light information. Among them, the visible light information includes the spatial position coordinates of each pixel in the image. After the affine transformation alignment, the visible light information can be used to determine the spatial position coordinates of the sparse sampling points in the hyperspectral information.

[0089] In step S130, the visible light instance segmentation module processes the visible light information using a pre-trained instance segmentation model, and extracts the contour regions of one or more target objects to be classified from the visible light information.

[0090] In step S140, the spectral information extraction module extracts the coordinates of all hyperspectral dimensionality reduction sampling points within the contour region of each target object according to the contour region of the target object and the result of the affine transformation alignment. Then, according to the coordinates of the sparse sampling points in all the hyperspectral information within each contour region, the spectral dispersion bands of each sparse sampling point are obtained, and further, the spectral curves of each sparse sampling point are generated based on the spectral dispersion bands.

[0091] In step S150, the spectral feature classification module performs data analysis on the spectral curves of all sparse sampling points within the contour region of each target object to obtain the spectral features of the target object. For example, the average value of the spectral curves of all sparse sampling points can be calculated as the spectral feature of the instance, and the feature classification is performed on it through a pre-trained classification model to obtain the category to which the target object belongs.

[0092] The embodiment of the present application improves the high-speed spectral video system according to actual application requirements, directly analyzes the spectral features of objects only from the sparse sampling of signals, realizes efficient spectral analysis of real-time acquisition and accelerated processing, can also quickly classify target objects, and does not require complex spectral acquisition and calculation processes, saving algorithm processing time and improving the execution efficiency of industrial sorting and other visual classification tasks.

[0093] In one implementation, Figure 1 in step S110, the obtaining of the hyperspectral information of the sparse sampling points in the scene where the target object to be classified is located specifically may include:

[0094] Sampling the optical signal in the scene where the target object is located by using a mask to obtain the optical signal of the sparse sampling points;

[0095] Dispersing the optical signal of the sparse sampling points by using a prism and restricting the dispersion length through a broadband filter;

[0096] Obtaining the hyperspectral information of the sparse sampling points from the dispersed optical signal.

[0097] Figure 2 This is a schematic optical path diagram of the hyperspectral sparse information and visible light information acquisition module for the target object classification method provided by the embodiment of the present application. Refer to Figure 2 , the hyperspectral sparse information acquisition module includes a beam splitter, an objective lens, a mask, a broadband filter, a prism or a grating, an eyepiece, and a grayscale camera. Figure 2Among them, Scene represents the scene where the target object to be classified is located; Beam Splitter represents a beam splitter; Objective-Lens represents an objective lens; Mask represents a mask; Relay-Lens represents a relay lens; Bands-Selection represents a broadband filter; Prism represents a prism; 900-1700nm Sensor represents a grayscale camera sensor; Mirror represents a reflector; RGB Sensor represents an RGB guiding camera. The eyepiece in the hyperspectral sparse information acquisition module is behind the prism and in front of the grayscale camera sensor. Figure 2 It is not drawn in the figure.

[0098] See Figure 2 , the objective lens forms an image of the scene on the mask plane; the eyepiece converges the image on the sensor target surface of the grayscale camera; the relay lens turns the output optical signal into parallel light; the main function of the reflector is to keep the two cameras parallel in the optical path. In another example, the reflector can also be omitted, and the eyepiece and the RGB guiding camera can be directly set behind the beam splitter. The grating, like the prism, plays the role of dispersion.

[0099] See Figure 2 , the visible light information acquisition module includes an eyepiece and an RGB guiding camera. The beam splitter divides the optical signal into two identical parts, and one of the beams is transmitted to the hyperspectral sparse information acquisition module. The visible light information acquisition module is used to receive the other beam of the above beam splitter to obtain an RGB video with high spatial resolution. The image collected by the RGB guiding camera is mainly used to determine the contour area of the target object and the spatial position coordinates of the sampling spectra of the sparse sampling points therein. The eyepiece in the visible light information acquisition module is in front of the RGB guiding camera and is used to converge the image on the sensor target surface of the RGB camera. Figure 2 It is not drawn in the figure.

[0100] See Figure 2 , the beam splitter divides the optical signal into two identical parts, one beam is transmitted to the visible light information acquisition module, and the other beam is spatially sampled using a mask with evenly distributed acquisition points. The role of the mask here is: to sample the optical signal of the scene where the target object is located to obtain the optical signal of the sparse sampling points. That is to say, information is not collected for all pixel points, and spectral information is collected only for the evenly distributed acquisition points through the mask. By using the mask to limit the amount of information acquisition data, the amount of hyperspectral information data is greatly reduced through the above limitation, thereby saving processing time and improving execution efficiency.

[0101] See Figure 2, after spatial sampling using a mask, a prism is used to disperse the optical signals of the sparse sampling points, and a broadband filter is used to limit the dispersion length. Among them, the broadband filter can be composed of a high-pass filter and a low-pass filter. For example, if the customized band is 450nm - 650nm, a 450nm high-pass filter and a 650nm low-pass filter are used to filter out the light outside this band range to eliminate interference; at the same time, it prevents the dispersion stripes of the sampling points from being too long and causing information aliasing between adjacent sampling points.

[0102] After limiting the dispersion length by the broadband filter, since the pixels of the mask acquisition points are far enough apart from each other, the sampled spectral dispersion does not overlap after reaching the sensor. After tiling the information of the sampling points onto the grayscale camera sensor in a non-aliasing manner, the spectral intensity of the dispersion can be directly read from the image, thereby sacrificing the scene spatial resolution to obtain the hyperspectral information of the sparse sampling points and transmitting the hyperspectral information to the spectral information extraction module.

[0103] In one implementation Figure 1 in step S120 of , the affine transformation alignment of the hyperspectral information and the visible light information includes:

[0104] Performing affine transformation alignment according to pre-calibrated optical mapping parameters; among them, the optical mapping parameters include: the sampling point coordinates of the mask, the dispersion drift distance corresponding to each spectral segment, and the affine transformation matrix.

[0105] After obtaining the hyperspectral information and the visible light information, further, the image registration module performs synchronous spatial affine transformation alignment on the two video frames of the hyperspectral information and the visible light information according to the pre-calibrated optical mapping parameters. Through the transformation alignment, the sparse sampling points captured by the hyperspectral information acquisition module can all be indexed in the RGB coordinate system, so that the spectral information extraction module can find the number of spectral samplings and the initial sampling points within the contour area of the target object output by the visible light instance segmentation module.

[0106] Among them, the optical mapping parameters mainly include the following parts:

[0107] (1) The sampling point coordinates of the mask in the hyperspectral sparse information acquisition module. By obtaining this parameter through pre-calibration, it can be known which pixel point in the scene the dispersion stripe is the spectral information of.

[0108] (2) The dispersion drift distance of each spectral segment relative to the initial sampling point. By pre-calibrating this parameter, it can be known the starting position and ending position of the dispersion stripe, and which spectral segment's intensity the pixel value at each position of the dispersion stripe corresponds to.

[0109] (3) An affine transformation matrix for aligning two-channel system images. After performing the affine transformation, a pair of matching data can be obtained from the hyperspectral information and the visible light information.

[0110] In one embodiment, the instance segmentation model includes a U-net model or a Mask R-CNN model.

[0111] The visible light instance segmentation module uses an image instance segmentation algorithm to obtain the region labels of the target objects to be classified. The U-net or Mask R-CNN instance segmentation model is used to extract the contour regions of the target objects to be classified. The execution result of the instance segmentation algorithm is as Figure 3 shown, where "bottle" represents a bottle, "cup" represents a cup, and "cube" represents a cube.

[0112] In another example, compared with the bounding box of object detection, instance segmentation can be accurate to the edge of the object, obtain a finer target ROI, and transmit it to the spectral information extraction module. The U-net or Mask R-CNN instance segmentation model is used to obtain the ROI of the target objects to be classified. The execution result of the instance segmentation algorithm is as Figure 4 shown.

[0113] As Figure 5 shown, in one embodiment, Figure 1 in step S140 of , obtaining the spectral curve of each sparse sampling point within the contour region from the hyperspectral information according to the contour region of the target object and the result of the affine transformation alignment includes:

[0114] Step S310, extracting the coordinates of the sparse sampling points within the contour region of the target object from the hyperspectral information according to the contour region of the target object and the result of the affine transformation alignment;

[0115] Step S320, obtaining the spectral curve of each sparse sampling point within the contour region from the hyperspectral information according to the coordinates of the sparse sampling points.

[0116] After performing the affine transformation alignment and extracting the contour regions of the target objects to be classified, further, the spectral information extraction module obtains the initial coordinates of all the hyperspectral sparse sampling points within each object contour region according to the contour region output by the visible light instance segmentation module and the spectral sampling point coordinates after registration and alignment, and then obtains the spectral curve of each sparse sampling point through the dispersion stripe in its hyperspectral information.

[0117] The spectral features of each object obtained by the spectral information extraction module are as Figure 6 shown. Figure 6On the left side in the middle is the ROI output by the visible light instance segmentation module, and on the right side are the spectral curves of the sparse sampling points in the ROIs of the target objects object1 and object2 respectively. The vertical coordinate Value of the spectral curve represents the spectral intensity, and the horizontal coordinate Wavelength represents the wavelength (unit: nm).

[0118] In one implementation, Figure 1 in step S150 of, obtaining the spectral feature of the target object according to the spectral curves of all sparse sampling points in the contour region includes:

[0119] Calculating the average value of the spectral curves of all sparse sampling points in the contour region;

[0120] Taking the average value as the spectral feature of the target object.

[0121] For a target object, the ROI usually includes the spectral curves of multiple sampling points. For most classification tasks, it is not necessary to perform redundant analysis on each sampling point. Therefore, calculating the average value of the spectral curves in this set of sampling points can be used as the spectral feature of the target object and output to the spectral feature classification module. The spectral feature of the target object can be calculated using the following formula:

[0122]

[0123] Wherein, represents the spectral feature information of the th object, is the set of sampling points in its ROI, is the total number of sampling points in the set, is the spectral curve of the th sampling point in the set.

[0124] Refer to Figure 6 again. The polygonal area in the upper left subfigure is a schematic diagram of the ROI of object 1. Several sampling points are marked inside the ROI. The lower left subfigure is the image collected by the hyperspectral sparse information acquisition module. The spectral curve of each sampling point can be obtained through the dispersion strip in the image. The spectral curve includes the spectral intensities of different wavelengths of the light at the sampling point. Calculating the average value of the spectral curves of all sparse sampling points in the polygonal area, finally obtaining the spectral feature information in the above formula .

[0125] In one implementation, the classification model includes a support vector machine or a decision tree model.

[0126] The spectral feature classification module models and classifies spectral features through a machine learning model or a deep learning model to obtain the category to which the target object belongs. In one example, the target object is a picked fruit. By using the target object classification method provided in the embodiments of the present application to classify the picked fruit, the categories to which the target object belongs can be obtained, including apples, pears, oranges, etc.

[0127] Figure 7 FIG. is a schematic diagram for quickly classifying a target object according to an embodiment of the target object classification method based on hyperspectral information provided in the embodiments of the present application. As Figure 7 shown, the image of the scene where the target object to be classified is located is input into the high-speed spectral video system. The system classifies the target object to obtain the spectral curves of sparse sampling points and the category to which the target object belongs. In Figure 7 the example, the scene image includes target object A and target object B. The classification result obtained by the system is: the category to which target object A belongs is non-pear, and the category to which target object B belongs is pear.

[0128] As Figure 8 shown, the present application also provides a corresponding embodiment of a target object classification device based on hyperspectral information. Regarding the beneficial effects or technical problems solved by the device, reference can be made to the descriptions in the methods corresponding to each device, or to the descriptions in the summary of the invention, which will not be elaborated here one by one.

[0129] In this embodiment of the target object classification device based on hyperspectral information, the device includes:

[0130] A first acquisition unit 100, configured to acquire the hyperspectral information of sparse sampling points of the scene where the target object to be classified is located; and acquire the visible light information of the scene where the target object to be classified is located;

[0131] An alignment unit 200, configured to perform affine transformation alignment on the hyperspectral information and the visible light information;

[0132] A segmentation unit 300, configured to extract the contour region of the target object to be classified from the visible light information by using an instance segmentation model;

[0133] A second acquisition unit 400, configured to acquire the spectral curve of each sparse sampling point within the contour region from the hyperspectral information according to the contour region of the target object and the result of the affine transformation alignment;

[0134] A classification unit 500, configured to obtain the spectral features of the target object according to the spectral curves of all sparse sampling points within the contour region, and perform feature classification on the target object based on the spectral features by using a classification model.

[0135] In one embodiment, the first acquisition unit 100 is configured to:

[0136] Sample the optical signal of the scene where the target object is located by using a mask to obtain the optical signal of the sparse sampling points;

[0137] Disperse the optical signal of the sparse sampling points by using a prism, and limit the dispersion length through a broadband filter;

[0138] Obtain the hyperspectral information of the sparse sampling points from the dispersed optical signal.

[0139] In one embodiment, the alignment unit 200 is configured to:

[0140] Perform affine transformation alignment according to pre-calibrated optical mapping parameters; wherein, the optical mapping parameters include: the sampling point coordinates of the mask, the dispersion drift distance corresponding to each spectral segment, and the affine transformation matrix.

[0141] In one embodiment, the instance segmentation model includes a U-net model or a Mask R-CNN model.

[0142] In one embodiment, the second acquisition unit 400 is configured to:

[0143] Extract the coordinates of the sparse sampling points within the contour region of the target object from the hyperspectral information according to the contour region of the target object and the result of the affine transformation alignment;

[0144] Obtain the spectral curve of each sparse sampling point within the contour region from the hyperspectral information according to the coordinates of the sparse sampling points.

[0145] In one embodiment, the classification unit 500 is configured to:

[0146] Average the spectral curves of all the sparse sampling points within the contour region;

[0147] Use the average value as the spectral feature of the target object.

[0148] In one embodiment, the classification model includes a support vector machine or a decision tree model.

[0149] Figure 9 It is a structural schematic diagram of a computing device 900 provided by an embodiment of the present application. The computing device 900 includes: a processor 910, a memory 920, and a communication interface 930.

[0150] It should be understood that Figure 9 the communication interface 930 in the computing device 900 shown in

[0151] Among them, the processor 910 can be connected to the memory 920. The memory 920 can be used to store the program code and data. Therefore, the memory 920 can be a storage unit inside the processor 910, an external storage unit independent of the processor 910, or a component including a storage unit inside the processor 910 and an external storage unit independent of the processor 910.

[0152] Optionally, the computing device 900 may further include a bus. Among them, the memory 920 and the communication interface 930 can be connected to the processor 910 through the bus. The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0153] It should be understood that in the embodiments of the present application, the processor 910 can adopt a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Or the processor 910 adopts one or more integrated circuits to execute relevant programs to implement the technical solutions provided by the embodiments of the present application.

[0154] The memory 920 can include a read-only memory and a random access memory, and provide instructions and data to the processor 910. A part of the processor 910 can also include a non-volatile random access memory. For example, the processor 910 can also store information about the device type.

[0155] When the computing device 900 is running, the processor 910 executes the computer-executable instructions in the memory 920 to perform the operation steps of the above method.

[0156] It should be understood that the computing device 900 according to the embodiments of the present application may correspond to the corresponding subject executing the methods according to the embodiments of the present application, and the above and other operations and / or functions of each module in the computing device 900 respectively implement the corresponding processes of the methods in each of the present embodiments. For the sake of brevity, they will not be elaborated herein.

[0157] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0158] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0159] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0160] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the present embodiment.

[0161] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0162] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0163] Embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it is used to execute a method for classifying target objects based on hyperspectral information, and this method includes at least one of the solutions described in the above-mentioned various embodiments.

[0164] The computer storage medium of the embodiments of this application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component.

[0165] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or component.

[0166] The program code contained on a computer-readable medium can be transmitted with any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.

[0167] The computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0168] Note that the above is only the preferred embodiment of the present application and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present application has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, all of which fall within the protection scope of the present invention.

Claims

1. A method for classifying target objects based on hyperspectral information, characterized in that, Including: Obtaining hyperspectral information of sparse sampling points in the scene where the target object to be classified is located; And obtaining visible light information of the scene where the target object to be classified is located; Performing affine transformation alignment on the hyperspectral information and the visible light information; Using an instance segmentation model to extract the contour region of the target object to be classified from the visible light information; According to the contour region of the target object and the result of the affine transformation alignment, obtaining the spectral curve of each sparse sampling point within the contour region from the hyperspectral information; Obtaining the spectral feature of the target object based on the spectral curves of all sparse sampling points within the contour region, and classifying the features of the target object using a classification model based on the spectral feature.

2. The method according to claim 1, characterized in that, The obtaining of the hyperspectral information of the sparse sampling points in the scene where the target object to be classified is located includes: Sampling the optical signal in the scene where the target object is located using a mask to obtain the optical signal of the sparse sampling points; Dispersing the optical signal of the sparse sampling points using a prism and restricting the dispersion length through a broadband filter; Obtaining the hyperspectral information of the sparse sampling points from the dispersed optical signal.

3. The method according to claim 2, characterized in that, The performing of the affine transformation alignment on the hyperspectral information and the visible light information includes: Performing affine transformation alignment according to pre-calibrated optical mapping parameters; wherein, the optical mapping parameters include: the sampling point coordinates of the mask, the dispersion drift distance corresponding to each spectral segment, and the affine transformation matrix.

4. The method according to claim 1, characterized in that, The instance segmentation model includes a U-net model or a MaskR-CNN model.

5. The method according to claim 1, characterized in that, According to the contour region of the target object and the result of the affine transformation alignment, obtaining the spectral curve of each sparse sampling point within the contour region from the hyperspectral information includes: According to the contour region of the target object and the result of the affine transformation alignment, extracting the coordinates of the sparse sampling points within the contour region of the target object from the hyperspectral information; According to the coordinates of the sparse sampling points, obtaining the spectral curve of each sparse sampling point within the contour region from the hyperspectral information.

6. The method according to claim 1, characterized in that, Obtaining the spectral feature of the target object based on the spectral curves of all sparse sampling points within the contour region includes: Calculating the average value of the spectral curves of all sparse sampling points within the contour region; Taking the average value as the spectral feature of the target object.

7. The method according to claim 1, characterized in that, The classification model includes a support vector machine or a decision tree model.

8. A device for classifying target objects based on hyperspectral information, characterized in that, For implementing the method according to any one of claims 1-7, the apparatus includes: A first acquisition unit, configured to acquire hyperspectral information of sparse sampling points in the scene where the target object to be classified is located; and acquire visible light information of the scene where the target object to be classified is located; An alignment unit, configured to perform affine transformation alignment on the hyperspectral information and the visible light information; A segmentation unit, configured to use an instance segmentation model to extract the contour region of the target object to be classified from the visible light information; A second acquisition unit, configured to obtain the spectral curve of each sparse sampling point within the contour region from the hyperspectral information according to the contour region of the target object and the result of the affine transformation alignment; A classification unit is used to obtain the spectral characteristics of the target object based on the spectral curves of all sparse sampling points within the contour region, and classify the features of the target object using a classification model based on the spectral characteristics.

9. A computing device, characterized in that, It includes: A communication interface; At least one processor, which is connected to the communication interface; And At least one memory, which is connected to the processor and stores program instructions. When the program instructions are executed by the at least one processor, the at least one processor executes the method according to any one of claims 1-7.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by a computer, the computer executes the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-feature multi-level visible light and high-spectrum image high-precision registering method

    CN102800099A

Cited By

  • Cultural relic hyperspectral image reconstruction method and device, hyperspectral area-array camera and medium

    CN120932112A