Multispectral intelligent sensing method and device

By building an intelligent perception network, the high-dimensional semantic information is directly obtained from the spectral coupled data of the spectral imaging system, and the problems of high computing and time cost in the prior art are solved, and efficient multi-spectral intelligent perception is achieved.

CN114998623BActive Publication Date: 2025-08-08BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210648807.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-08-08
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

The existing RGB and grayscale imaging methods have low perception dimensions and are greatly affected by light. The spectral perception system based on deep learning requires complex imaging processes, resulting in high computing and time costs and hindering the promotion of multi-spectral intelligent perception systems.

Method used

By building an intelligent perception network, high-dimensional semantic information is directly obtained from the spectral coupled data of the computational spectral imaging system, and end-to-end perceptual calculations are used to use deep learning networks, including high-dimensional semantic feature extraction and semantic analysis modules, avoiding the spectral imaging process.

Benefits of technology

Reduces the computational complexity, improves perception accuracy and response speed, and achieves efficient multi-spectral intelligent perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998623B_ABST
    Figure CN114998623B_ABST
Patent Text Reader

Abstract

This invention proposes a multispectral intelligent perception method and device. The method comprises obtaining a training perception dataset comprising multiple sets of spectral coupling data and intelligent perception data pairs; constructing an intelligent perception network and training the intelligent perception network using the training perception dataset to obtain a trained intelligent perception network; using a computational spectral imaging system to collect target spectral coupling data; and inputting the target spectral coupling data into the trained intelligent perception network for end-to-end perception computation to obtain intelligent perception results. This method eliminates the need to reconstruct spectral information for all channels and regions. Instead, it uses a deep learning network to efficiently and intelligently perceive the semantic information of the target scene directly from the data collected by the computational spectral imaging system, effectively reducing computational complexity and achieving fast response speed and high perception accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computational photography. Background Art

[0002] Traditional RGB and grayscale imaging have low perceptual dimensionality, are significantly affected by lighting, and exhibit the phenomenon of "metachromatic heterochromaticity." Intelligent perception methods based on these imaging methods are limited by the limited perceptual information and perform poorly in perception areas such as classification, segmentation, and object detection. Spectra, with their unique representational properties, can exceed the spectral perception range and resolution of the human eye, and are widely used in aerospace remote sensing, biomedicine, agricultural inspections, and industrial quality control.

[0003] However, spectral imaging data is large and contains redundant information. Image quality depends on complex imaging hardware (such as high-sensitivity sensors and achromatic lenses) and highly complex reconstruction algorithms (such as hyperspectral reconstruction algorithms). This results in high cost and computational complexity for existing perception systems, as well as low perception rates and heavy communication overhead. These shortcomings have hindered the practical application and widespread adoption of multispectral intelligent perception systems.

[0004] Due to the sparse priors of natural scenes, redundant information and high-fidelity reconstruction are wasteful hardware and software costs in most intelligent perception tasks. Existing deep learning-based spectral perception systems still require multispectral image reconstruction, which introduces a complex imaging process and incurs high computational and time costs. Therefore, bypassing the complex multispectral imaging process and directly acquiring target features rather than acquiring features through images or videos is a viable solution to improve information acquisition and perception efficiency. Summary of the Invention

[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0006] To this end, the first object of the present invention is to propose a spectral intelligent perception method for directly acquiring high-dimensional semantic information of a scene from spectral coupling data of a computational spectral imaging system.

[0007] The second object of the present invention is to provide a spectral intelligent sensing device.

[0008] A third object of the present invention is to provide a computer device.

[0009] A fourth object of the present invention is to provide a computer-readable storage medium.

[0010] To achieve the above objectives, a first embodiment of the present invention provides a spectrum intelligent perception method, comprising:

[0011] Acquire a training perception data set, wherein the perception data set includes multiple sets of spectral coupling data-intelligent perception data pairs;

[0012] Constructing an intelligent perception network, and training the intelligent perception network using the training perception data set to obtain a trained intelligent perception network;

[0013] Use a computational spectral imaging system to collect target spectral coupling data;

[0014] The target spectrum coupling data is input into the trained intelligent perception network to perform end-to-end perception calculation to obtain an intelligent perception result.

[0015] In addition, the spectrum intelligent perception method according to the above embodiment of the present invention may also have the following additional technical features:

[0016] Furthermore, in one embodiment of the present invention, the intelligent perception network includes a high-dimensional semantic feature extraction module and a semantic parsing module.

[0017] Furthermore, in one embodiment of the present invention, the design method of the high-dimensional semantic feature extraction module includes but is not limited to any one or more of the following:

[0018] Fully connected networks, convolutional networks, recurrent neural networks, and transformer networks based on attention mechanisms.

[0019] Furthermore, in one embodiment of the present invention, the semantic parsing module takes high-dimensional semantic features as input and outputs intelligent perception results of the scene. The design method includes but is not limited to any one of the following:

[0020] ResNet-based classification network, DenseNet-based classification network, Inception-based classification network, MobileNet-based classification network, EfficientNet-based classification network, Unet-based segmentation network, R-CNN-based target detection network, SPP-Net-based target detection network, Yolo-based target detection network and MaskR-CNN-based target detection network.

[0021] Furthermore, in one embodiment of the present invention, the using the training perception data set to train the intelligent perception network includes:

[0022] In the first stage, the high-dimensional semantic feature extraction capability of the high-dimensional semantic feature extraction module is trained on a large-scale self-supervised dataset; in the second stage, the semantic parsing module is replaced with a classification network, a segmentation network or a target detection network, and trained on the training perception dataset. At the same time, the parameters of the high-dimensional semantic feature extraction module and the semantic parsing module are updated to perform end-to-end classification, segmentation or target detection training.

[0023] Furthermore, in one embodiment of the present invention, the first stage includes:

[0024] The input multispectral information is weightedly fused, and a method for obtaining the weight of the weighted fusion includes a learning-based method and a camera parameter-based method.

[0025] Furthermore, in one embodiment of the present invention, it further comprises:

[0026] By using classification, segmentation or target detection perception data sets to perform two-stage training on the intelligent perception network, the optimal high-dimensional semantic feature extraction module and semantic parsing module are finally obtained.

[0027] To achieve the above objectives, a second embodiment of the present invention provides a multispectral intelligent sensing device, comprising the following modules:

[0028] An acquisition module is used to acquire a training perception data set, wherein the perception data set includes multiple sets of spectral coupling data-intelligent perception data pairs;

[0029] A training module is used to build an intelligent perception network and train the intelligent perception network using the training perception data set to obtain a trained intelligent perception network;

[0030] An acquisition module, configured to acquire target spectral coupling data using a computational spectral imaging system;

[0031] The perception module is used to input the target spectrum coupling data into the trained intelligent perception network to perform end-to-end perception calculation to obtain an intelligent perception result.

[0032] To achieve the above-mentioned objectives, a third aspect of the present invention provides a computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the multispectral intelligent perception method described above is implemented.

[0033] To achieve the above-mentioned objectives, a fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the multispectral intelligent sensing method described above when executed by a processor.

[0034] The multispectral intelligent perception method and device proposed in the present invention eliminate the need for spectral imaging. Instead, they use a deep learning network to directly extract high-dimensional semantic information about a scene from the spectral coupling data of a computational spectral imaging system. This allows for intelligent scene perception, including but not limited to classification, segmentation, and object detection. This method reduces computational complexity, achieves rapid response, and achieves high perception accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0036] Figure 1 A schematic diagram of a multispectral intelligent sensing method provided by an embodiment of the present invention.

[0037] Figure 2 A schematic diagram of a multispectral intelligent sensing device provided by an embodiment of the present invention.

[0038] Figure 3 This is a schematic diagram of the first stage training provided by an embodiment of the present invention.

[0039] Figure 4 This is a schematic diagram of the second stage training provided by an embodiment of the present invention.

[0040] Figure 5 This is a schematic diagram of the test phase provided by an embodiment of the present invention.

[0041] Figure 6 This is a structural diagram of a multispectral intelligent sensing device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0043] The multispectral intelligent sensing method and device according to embodiments of the present invention will be described below with reference to the accompanying drawings.

[0044] Example 1

[0045] Figure 1 A schematic diagram of a multispectral intelligent sensing method provided by an embodiment of the present invention.

[0046] like Figure 1 As shown, the multispectral intelligent perception method includes the following steps:

[0047] S101: Acquire a training perception data set, where the perception data set includes multiple sets of spectral coupling data-intelligent perception data pairs;

[0048] S102: Building an intelligent perception network, and training the intelligent perception network using a training perception data set to obtain a trained intelligent perception network;

[0049] S103: using a computational spectral imaging system to collect target spectral coupling data;

[0050] S104: Input the target spectral coupling data into the trained intelligent perception network to perform end-to-end perception calculation to obtain an intelligent perception result.

[0051] Furthermore, in one embodiment of the present invention, the intelligent perception network includes a high-dimensional semantic feature extraction module and a semantic parsing module.

[0052] Furthermore, in one embodiment of the present invention, the design method of the high-dimensional semantic feature extraction module includes but is not limited to any one or more of the following:

[0053] Fully connected networks, convolutional networks, recurrent neural networks, and transformer networks based on attention mechanisms.

[0054] Furthermore, in one embodiment of the present invention, the semantic parsing module takes high-dimensional semantic features as input and outputs intelligent perception results of the scene. The design method includes but is not limited to any of the following:

[0055] ResNet-based classification network, DenseNet-based classification network, Inception-based classification network, MobileNet-based classification network, EfficientNet-based classification network, Unet-based segmentation network, R-CNN-based target detection network, SPP-Net-based target detection network, Yolo-based target detection network and MaskR-CNN-based target detection network.

[0056] Furthermore, in one embodiment of the present invention, the intelligent perception network is trained using a training perception data set, including:

[0057] In the first stage, the high-dimensional semantic feature extraction capability of the high-dimensional semantic feature extraction module is trained on a large-scale self-supervised dataset; in the second stage, the semantic parsing module is replaced with a classification network, a segmentation network or a target detection network, and trained on the training perception dataset. At the same time, the parameters of the high-dimensional semantic feature extraction module and the semantic parsing module are updated to perform end-to-end classification, segmentation or target detection training.

[0058] Furthermore, in one embodiment of the present invention, the first stage includes:

[0059] The input multispectral information is weightedly fused, and the methods for obtaining the weights of the weighted fusion include a learning-based method and a camera parameter-based method.

[0060] Furthermore, in one embodiment of the present invention, it further comprises:

[0061] By using classification, segmentation or target detection perception data sets to perform two-stage training on the intelligent perception network, the optimal high-dimensional semantic feature extraction module and semantic parsing module are finally obtained.

[0062] Example 2

[0063] like Figure 2 Shown is the basic block diagram of the system of the present invention.

[0064] like Figure 3 FIG. 1 shows the first training phase of the present invention.

[0065] The feature extraction stage consists of two parts: a multispectral image weighted fusion module and a high-dimensional semantic feature extraction module. In this example, we train the perception network by using spectral coupled data synthesized by weighted fusion of an existing multispectral dataset. That is, the weighted fusion module is used to fuse multispectral images into coupled spectral data. Its input is a multispectral target, and its output is coupled multispectral data. The methods for determining the weights of each channel during weighted fusion in this module include but are not limited to camera parameter-based determination and learning-based determination. The high-dimensional semantic feature extraction module is responsible for extracting multi-channel high-dimensional semantic features that are effective for scene intelligent perception from the coupled data. The high-dimensional semantic feature extraction network includes but is not limited to convolutional layers and Transformer layers based on the attention mechanism.

[0066] The reconstruction component consists of a spectral image reconstruction module. Its input is multi-channel high-dimensional semantic features, and its output is a reconstructed multispectral image. In the first training phase, we enhance the feature extraction capabilities of the high-dimensional semantic feature extraction module by training the perception network's spectral image reconstruction capabilities.

[0067] In the second training phase, we replace the reconstruction module with a semantic parsing module to parse multi-channel, high-dimensional semantic features and perceive scene information from them. The semantic parsing module takes multi-channel, high-dimensional semantic features as input and outputs intelligent perception results such as object category information, location information, and pixel-level segmentation results. Its design methods include, but are not limited to, the following architectures: ResNet-based classification networks, DenseNet-based classification networks, Inception-based classification networks, MobileNet-based classification networks, EfficientNet-based classification networks, Unet-based segmentation networks, R-CNN-based object detection networks, SPP-Net-based object detection networks, Yolo-based object detection networks, and Mask R-CNN-based object detection networks.

[0068] The training of intelligent perception network is divided into two stages. Figure 3 As shown in the figure, the first training phase focuses on training the high-dimensional semantic feature extraction module. This phase is performed on a large-scale self-supervised training dataset. The decoder's function at this stage is to reconstruct the original multispectral image. During this phase, the parameters of the high-dimensional semantic feature extraction module and the spectral image reconstruction module are updated simultaneously.

[0069] The second phase of training is Figure 4 As shown, the parameters of the high-dimensional semantic feature extraction module are first fixed, and only the semantic parsing module is trained independently. Once the semantic parsing module can achieve preliminary intelligent perception functions such as classification, segmentation, or object detection, the training strategy is changed to simultaneously updating the parameters of the high-dimensional semantic feature extraction module and the semantic parsing module to find the optimal overall network structure. The second stage of training is performed on multispectral classification, segmentation, or object detection datasets.

[0070] Testing phase Figure 5 In actual use, the computational spectral imaging system is first used to collect spectral coupling data, which is then input into the high-dimensional semantic feature extraction module. The output multi-channel high-dimensional semantic features are then sent to the semantic parsing module to perform intelligent perception of the scene.

[0071] The multispectral intelligent perception method proposed in this embodiment does not require spectral imaging. Instead, it uses a deep learning network to directly extract high-dimensional semantic information about a scene from the spectral coupling data of a computational spectral imaging system. This method then performs intelligent scene perception, including but not limited to classification, segmentation, and object detection. This method reduces computational complexity, achieves rapid response, and achieves high perception accuracy.

[0072] In order to implement the above embodiment, the present invention also proposes a multi-spectral intelligent sensing device.

[0073] Figure 6 A schematic diagram of the structure of a multispectral intelligent sensing device provided by an embodiment of the present invention.

[0074] like Figure 6 As shown, the multispectral intelligent sensing device includes: an acquisition module 100, a training module 200, a collection module 300, and a sensing module 400, wherein:

[0075] An acquisition module is used to acquire a training perception data set, where the perception data set includes multiple sets of spectral coupling data-intelligent perception data pairs;

[0076] The training module is used to build an intelligent perception network and train the intelligent perception network using a training perception data set to obtain a trained intelligent perception network.

[0077] An acquisition module, configured to acquire target spectral coupling data using a computational spectral imaging system;

[0078] The perception module is used to input the target spectral coupling data into the trained intelligent perception network for end-to-end perception calculation to obtain intelligent perception results.

[0079] To achieve the above-mentioned objectives, the third aspect of the present invention proposes a computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the multispectral intelligent perception method as described above is implemented.

[0080] To achieve the above-mentioned objectives, a fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the multispectral intelligent sensing method described above when executed by a processor.

[0081] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0082] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0083] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limiting the present invention. A person skilled in the art may change, modify, replace, and modify the above embodiments within the scope of the present invention.

Claims

1. A multispectral intelligent perception method, characterized in that: The following steps are involved: Constructing an intelligent perception network and training the intelligent perception network to obtain a trained intelligent perception network, wherein the intelligent perception network includes a high-dimensional semantic feature extraction module and a semantic parsing module, the input of the intelligent perception network is spectral coupling data, and the output is an intelligent perception result; Use a computational spectral imaging system to collect target spectral coupling data; Inputting the target spectrum coupling data into the trained intelligent perception network to perform end-to-end perception calculation to obtain an intelligent perception result; The step of training the intelligent perception network includes: In the first stage, the high-dimensional semantic feature extraction capability of the high-dimensional semantic feature extraction module is trained on a large-scale self-supervised dataset. In the first stage, the multispectral image is fused into coupled spectral data through a weighted fusion module. Multi-channel high-dimensional semantic features effective for scene intelligent perception are extracted from the coupled spectral data through the high-dimensional semantic feature extraction module. The multi-channel high-dimensional semantic features are input into a spectral image reconstruction module, and a multispectral image is output. In the first stage, the parameters of the high-dimensional semantic feature extraction module and the spectral image reconstruction module are simultaneously updated; In the second stage, the semantic parsing module is used instead of the spectral image reconstruction module to perform end-to-end classification, segmentation or target detection training on the multispectral classification, segmentation or target detection dataset. In the second stage, the parameters of the high-dimensional semantic feature extraction module are fixed, and the semantic parsing module is trained separately. After the semantic parsing module realizes the preliminary intelligent perception function of classification, segmentation or target detection, the parameters of the high-dimensional semantic feature extraction module and the semantic parsing module are updated at the same time.

2. The method according to claim 1, characterized in that The high-dimensional semantic feature extraction module may be designed by methods including but not limited to any one or more of the following: Fully connected networks, convolutional networks, recurrent neural networks, and transformer networks based on attention mechanisms.

3. The method according to claim 1, characterized in that The semantic parsing module takes high-dimensional semantic features as input and outputs intelligent perception results of the scene. The design methods include but are not limited to any of the following: ResNet-based classification network, DenseNet-based classification network, Inception-based classification network, MobileNet-based classification network, EfficientNet-based classification network, Unet-based segmentation network, R-CNN-based target detection network, SPP-Net-based target detection network, Yolo-based target detection network and MaskR-CNN-based target detection network.

4. A multispectral intelligent sensing device, characterized in that: Includes the following modules: A training module is used to construct an intelligent perception network and train the intelligent perception network to obtain a trained intelligent perception network, wherein the intelligent perception network includes a high-dimensional semantic feature extraction module and a semantic parsing module, the input of the intelligent perception network is spectral coupling data, and the output is an intelligent perception result; An acquisition module, configured to acquire target spectral coupling data using a computational spectral imaging system; A perception module, configured to input the target spectral coupling data into the trained intelligent perception network to perform end-to-end perception calculations to obtain intelligent perception results; The step of training the intelligent perception network includes: In the first stage, the high-dimensional semantic feature extraction capability of the high-dimensional semantic feature extraction module is trained on a large-scale self-supervised dataset. In the first stage, the multispectral image is fused into coupled spectral data through a weighted fusion module. Multi-channel high-dimensional semantic features effective for scene intelligent perception are extracted from the coupled spectral data through the high-dimensional semantic feature extraction module. The multi-channel high-dimensional semantic features are input into a spectral image reconstruction module, and a multispectral image is output. In the first stage, the parameters of the high-dimensional semantic feature extraction module and the spectral image reconstruction module are simultaneously updated; In the second stage, the semantic parsing module is used instead of the spectral image reconstruction module to perform end-to-end classification, segmentation or target detection training on the multispectral classification, segmentation or target detection dataset. In the second stage, the parameters of the high-dimensional semantic feature extraction module are fixed, and the semantic parsing module is trained separately. After the semantic parsing module realizes the preliminary intelligent perception function of classification, segmentation or target detection, the parameters of the high-dimensional semantic feature extraction module and the semantic parsing module are updated at the same time.

5. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method implements the multispectral intelligent perception method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multispectral intelligent sensing method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Hyperspectral image camouflage target detection method based on deep learning

    CN111368712A

  • Multispectral remote sensing image urban village multi-class building semantic segmentation method and system

    CN112633140A