Multi-Spectral Image Scene Recognition Method, Apparatus, Electronic Device, and Storage Medium

By grouping and feature extraction of multispectral images, and fusing feature information in combination with cross-band attention fusion networks, the problem of under-exploring band relationships in the existing technology is solved, and the accuracy of scene recognition of multispectral images is improved.

CN118097347BActive Publication Date: 2025-06-10TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311800660.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-10
Estimated Expiration
2043-12-25

AI Technical Summary

Technical Problem

The interrelationships between different bands have not been fully explored in the prior art, resulting in the accuracy of multispectral image scene classification needs to be improved.

Method used

By grouping multi-spectral images, the band feature information is extracted using a preset network structure, and the cross-band attention fusion network is used to fuse these feature information, and finally the fusion feature information is input to the scene recognition model for classification.

Benefits of technology

By fully extracting and fusing feature information from different bands, complementary information between bands can be more effectively mined and the accuracy of multispectral image scene recognition can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118097347B_ABST
    Figure CN118097347B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-spectral image scene recognition method, apparatus, electronic device and storage medium, belonging to the technical field of image processing. The method includes: acquiring a multi-spectral image to be recognized, grouping the multi-spectral image according to the bands of the multi-spectral image to obtain at least one multi-spectral image group; extracting features from the multi-spectral image group based on a preset network structure to obtain band feature information corresponding to each multi-spectral image group; performing feature fusion on the band feature information based on a preset cross-band attention fusion network to obtain fused band feature information; inputting the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result. By fully extracting the feature information of different spectral bands and then fusing the feature information of different bands, the present invention fully excavates the complementary information between different bands, realizes the mutual enhancement of different features, and thus achieves better recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a multi-spectral image scene recognition method, apparatus, electronic device, and storage medium. Background Art

[0002] Remote sensing scene classification is an important task in the field of remote sensing. It mainly aims at a given remote sensing image and predicts the semantic categories of ground object targets contained in the image through certain algorithms.

[0003] Most of the existing technologies for remote sensing image scene classification are designed based on RGB images, and there is less research on scene classification of multi-spectral remote sensing images. Currently, the existing multi-spectral scene recognition methods can be roughly summarized into two technical frameworks. The first technical framework is to upsample the images of different bands to the same spatial resolution, then splice them in the channel dimension to form a multi-channel image data, and then directly input it into a classic convolutional neural network. The second technical framework is to group all bands according to different resolutions, where the bands with the same resolution are in one group, then input all groups into a multi-branch convolutional neural network, and then fuse the features of different groups through a simple fusion method, and finally input the fused features into the classification head to complete the final category prediction.

[0004] Deep learning relies on large-scale data sets. Due to the lack of multi-spectral image data sets, there are few scene classification methods for multi-spectral images currently. For the above two existing technical frameworks, the mutual relationships between different bands are not fully explored, so the accuracy needs to be improved. Summary of the Invention

[0005] The present invention provides a multi-spectral image scene recognition method, apparatus, electronic device, and storage medium to solve the defect in the prior art that the mutual relationships between different bands are not fully explored and the accuracy needs to be improved.

[0006] In a first aspect, the present invention provides a multi-spectral image scene recognition method, including:

[0007] Obtain a multi-spectral image to be recognized, and group the multi-spectral image according to the bands of the multi-spectral image to obtain at least one multi-spectral image group;

[0008] Extract features from the multi-spectral image group based on a preset network structure corresponding to the bands of the multi-spectral image group to obtain band feature information corresponding to each multi-spectral image group;

[0009] Fuse the band feature information based on a preset cross-band attention fusion network to obtain fused band feature information;

[0010] Input the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result.

[0011] According to a multi-spectral image scene recognition method provided by the present invention, the feature extraction of the multi-spectral image group based on a preset network structure corresponding to the bands of the multi-spectral image group to obtain band feature information corresponding to each multi-spectral image group includes:

[0012] Based on a preset backbone network, perform feature extraction on the multi-spectral image group in the RGB band to obtain RGB band feature information;

[0013] Based on a preset multi-stage grouped spectral feature extraction network, perform feature extraction on the multi-spectral image group in non-RGB bands to obtain band feature information corresponding to each multi-spectral image group.

[0014] According to a multi-spectral image scene recognition method provided by the present invention, the multi-stage grouped spectral feature extraction network includes an image patch embedding layer, a grouped image patch embedding layer corresponding to the bands of the multi-spectral image, and a multi-stage grouped spectral feature extraction network unit. Both the image patch embedding layer and the grouped image patch embedding layer are connected to one multi-stage grouped spectral feature extraction network unit; the multi-stage grouped spectral feature extraction network unit includes a grouped convolution module, a layer normalization module, an activation function module, and a residual connection module connected in sequence. The residual connection module is used to connect the input end and the output end of the multi-stage grouped spectral feature extraction network unit.

[0015] According to a multi-spectral image scene recognition method provided by the present invention, the backbone network includes at least one local feature extraction unit and at least one global feature extraction unit. The local feature extraction unit includes a first position encoding module, a local multi-head relationship aggregation module, and a first feed-forward neural network module arranged in sequence. A batch normalization module is provided between the first position encoding module and the local multi-head relationship aggregation module, and a batch normalization module is provided between the local multi-head relationship aggregation module and the first feed-forward neural network module; the global feature extraction unit includes a second position encoding module, a global multi-head relationship aggregation module, and a second feed-forward neural network module arranged in sequence. A layer normalization module is provided between the second position encoding module and the global multi-head relationship aggregation module, and a layer normalization module is provided between the global multi-head relationship aggregation module and the second feed-forward neural network module.

[0016] A multi - spectral image scene recognition method provided by the present invention, the cross - band attention fusion network includes a first linear layer, a second linear layer, a third linear layer and an attention module. The first linear layer is used to perform feature transformation on the RGB - band feature information to obtain feature query information. The second linear layer is used to perform feature transformation on the band feature information to obtain feature key information. The third linear layer is used to perform feature transformation on the band feature information to obtain feature value information. The attention module is used to determine the attention score corresponding to each band feature information according to the feature query information and the feature key information, and obtain the fused band feature information according to the attention score and the feature value information.

[0017] A multi - spectral image scene recognition method provided by the present invention, the scene recognition model includes a classifier. The classifier is used to classify the fused band feature information according to a preset single - label or multi - label classification task to determine a classification score. If the classifier has not been trained yet, then based on the classification score and a preset loss function, determine a loss value, and update the parameters of the preset network structure, the cross - band attention fusion network and the classifier according to the loss value until the classifier is trained. If the classifier is trained, then use the classification score as the scene recognition result.

[0018] In a second aspect, the present invention further provides a multi - spectral image scene recognition device, including:

[0019] A grouping module, configured to obtain a multi - spectral image to be recognized, and group the multi - spectral image according to the bands of the multi - spectral image to obtain at least one multi - spectral image group;

[0020] An extraction module, configured to perform feature extraction on the multi - spectral image group based on a preset network structure corresponding to the bands of the multi - spectral image group to obtain the band feature information corresponding to each multi - spectral image group;

[0021] A fusion module, configured to perform feature fusion on the band feature information based on a preset cross - band attention fusion network to obtain fused band feature information;

[0022] An identification module, configured to input the fused band feature information into a pre - trained scene recognition model to obtain a scene recognition result.

[0023] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of any one of the above - mentioned multi - spectral image scene recognition methods.

[0024] Fourthly, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-spectral image scene recognition method as described in any one of the above are implemented.

[0025] The multi-spectral image scene recognition method, device, electronic device and storage medium provided by the present invention fully extract the feature information of different spectral bands, then fuse the feature information of different bands, fully excavate the complementary information between different bands, and realize the mutual enhancement of different features, thereby achieving better recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 is one of the flow diagrams of the multi-spectral image scene recognition method provided by the present invention;

[0028] Figure 2 is another flow diagram of the multi-spectral image scene recognition method provided by the present invention;

[0029] Figure 3 is the structural diagram of the backbone network provided by the present invention;

[0030] Figure 4 is the structural diagram of the multi-stage grouped spectral feature extraction network provided by the present invention;

[0031] Figure 5 is the structural diagram of the multi-stage grouped spectral feature extraction network unit provided by the present invention;

[0032] Figure 6 is the structural diagram of the cross-band attention fusion network provided by the present invention;

[0033] Figure 7 is the structural diagram of the multi-spectral image scene recognition device provided by the present invention;

[0034] Figure 8 is the structural diagram of the electronic device provided by the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0036] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. The orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation to the present invention. Unless otherwise clearly defined and limited, the terms "mount", "connect" and "couple" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0037] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type and do not limit the number of objects. For example, the first object may be one or more. In addition, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0038] The following Figure 1 - Figure 8 describes the multi-spectral image scene recognition method, device, electronic device and storage medium provided by the present invention.

[0039] Figure 1 is one of the flow diagrams of the multi-spectral image scene recognition method provided by the present invention.Figure 2 This is the second schematic flow diagram of the multi-spectral image scene recognition method provided by the present invention. As Figure 1 and Figure 2 shown, it includes but is not limited to the following steps:

[0040] Step S110: Obtain the multi-spectral image to be recognized, and group the multi-spectral image according to the bands of the multi-spectral image to obtain at least one multi-spectral image group;

[0041] The bands include ultraviolet, visible - blue, visible - green, visible - red, near-infrared, short-wave infrared, panchromatic, thermal infrared, etc. An ordinary visible light camera only records the information of the visible - blue, visible - green, visible - red, i.e., the RGB bands, and the information of other bands will be lost. The multi-spectral image includes several or more than a dozen bands.

[0042] The multi-spectral image provided by the present invention includes information of multiple bands and is obtained based on satellite photography. Taking the Sentinel satellite as an example, the Sentinel satellite is a satellite used for surveillance and reconnaissance, usually for military and security purposes, and can provide real-time images and intelligence for monitoring enemy activities, border security, natural disaster monitoring, etc. The Sentinel satellite is usually equipped with high-resolution cameras and other sensors and can perform precise surveillance and tracking on ground targets.

[0043] Optionally, before grouping the multi-spectral image according to the bands of the multi-spectral image, it further includes:

[0044] Performing upsampling on the multi-spectral image to complete the preprocessing of the multi-spectral image.

[0045] The multi-spectral images taken by satellites may have different resolutions in different bands. For the convenience of processing, the present invention upsamples each band of the input multi-spectral image to the same spatial resolution. Specifically, the bicubic interpolation upsampling algorithm can be used to upsample each band to 128×128. Bicubic interpolation is a commonly used interpolation method and can be used for upsampling. In bicubic interpolation, the value of each new sampling point is calculated by a cubic interpolation function from its surrounding 16 neighboring sampling points. This interpolation method can, to a certain extent, maintain the smoothness and details of the original signal, so that a more real and clear result can be obtained during upsampling.

[0046] After completing the image preprocessing, the present invention groups the multi-spectral images according to the bands, combines the RGB bands as a group, and each of the other bands is used as a separate group. The RGB bands are closer to the visual perception system of the human eye, contain rich visual information, and the technology for feature extraction of RGB images is already very mature. Using each of the other bands as a separate group can also prevent the information of different bands from influencing each other and avoid the feature information of each band being submerged by other bands.

[0047] Step S120: Based on a preset network structure corresponding to the bands of the multi-spectral image group, perform feature extraction on the multi-spectral image group to obtain band feature information corresponding to each multi-spectral image group.

[0048] Specifically, set a corresponding preset network for the multi-spectral images of each group to extract the band feature information of the multi-spectral images. A feature extraction network can be set for the RGB group, and a feature extraction network can be set for the images of the bands other than the RGB bands.

[0049] Step S130: Based on a preset cross-band attention fusion network, perform feature fusion on the band feature information to obtain fused band feature information; the cross-band attention fusion network is implemented based on the attention mechanism and can perform feature fusion on the band feature information of different bands to obtain fused band feature information.

[0050] Step S140: Input the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result. The scene recognition model can be a classifier, which is used to recognize the scene category in the multi-spectral image according to the fused band feature information and output a preset number of scene categories and the probabilities corresponding to the scene categories.

[0051] It can be understood that the present invention fully extracts the feature information of different spectral bands, then fuses the feature information of different bands, fully excavates the complementary information between different bands, realizes the mutual enhancement of different features, and thus achieves better recognition accuracy.

[0052] Based on the above embodiments, as an optional embodiment, the performing feature extraction on the multi-spectral image group based on a preset network structure corresponding to the bands of the multi-spectral image group to obtain band feature information corresponding to each multi-spectral image group includes:

[0053] Based on a preset backbone network, perform feature extraction on the multi-spectral image group of the RGB bands to obtain RGB band feature information;

[0054] Based on a preset multi-stage grouped spectral feature extraction network, feature extraction is performed on the multi-spectral image group in non-RGB bands to obtain the band feature information corresponding to each multi-spectral image group.

[0055] Specifically, the present invention performs feature extraction on the multi-spectral image group corresponding to the RGB band based on a backbone network, and uniformly performs feature extraction on the multi-spectral image group in bands other than the RGB band based on the multi-stage grouped spectral feature extraction network.

[0056] The backbone network is the backbone network in the RGB feature extraction network in the prior art, such as UniFormer. UniFormer is a unified image feature extraction network that organically unifies the ideas of convolutional networks and self-attention.

[0057] It can be understood that the present invention obtains the RGB band feature information through the backbone network and obtains the band feature information of other bands through the multi-stage grouped spectral feature extraction network, and can learn the unique spectral information of each band. Instead of improving the accuracy of multi-spectral image recognition by increasing the number of training data, the present invention improves the multi-spectral image understanding ability by deeply excavating image features.

[0058] Based on the above embodiment, as an optional embodiment, the backbone network includes at least one local feature extraction unit and at least one global feature extraction unit. The local feature extraction unit includes a first position encoding module, a local multi-head relationship aggregation module, and a first feed-forward neural network module arranged in sequence. A batch normalization module is provided between the first position encoding module and the local multi-head relationship aggregation module, and a batch normalization module is provided between the local multi-head relationship aggregation module and the first feed-forward neural network module; the global feature extraction unit includes a second position encoding module, a global multi-head relationship aggregation module, and a second feed-forward neural network module arranged in sequence. A layer normalization module is provided between the second position encoding module and the global multi-head relationship aggregation module, and a layer normalization module is provided between the global multi-head relationship aggregation module and the second feed-forward neural network module.

[0059] Figure 3 is a schematic structural diagram of the backbone network provided by the present invention. As Figure 3 shown, the UniFormer backbone network of the present invention can include two local feature extraction units and two global feature extraction units to achieve multi-stage feature extraction. Figure 3 In

[0060] In Stage 1 and Stage 2, the idea of convolution is adopted to extract features from the image and perform local multi-head relationship aggregation; in Stage 3 and Stage 4, the idea of self-attention is adopted to model the long-range dependence relationship of the image and perform global multi-head relationship aggregation. Figure 3 In the local / global multi-head relationship aggregation module in Figure 3 , if it is a local multi-head relationship aggregation module, the network structure within the dotted line is a local feature extraction unit; if it is a global multi-head relationship aggregation module, it is a global feature extraction unit.

[0061] The first position encoding module and the second position encoding module can be dynamic position encoding modules, and the dynamic position encoding module is implemented through a convolutional layer with zero padding. Both the batch normalization module and the layer normalization module can make the training of the backbone network more stable. The first feed-forward neural network module and the second feed-forward neural network module both adopt fully connected networks to transform the feature dimensions.

[0062] There is a residual connection module between Stage 1, Stage 2, Stage 3 and Stage 4 to further improve the performance of the backbone network.

[0063] It can be understood that the present invention constitutes a multi-stage backbone network through at least one local feature extraction unit and at least one global feature extraction unit, and better RGB band feature information can be obtained.

[0064] Figure 4 is a schematic structural diagram of the multi-stage grouped spectral feature extraction network provided by the present invention. As Figure 4 shown, on the basis of the above embodiment, as an optional embodiment, the multi-stage grouped spectral feature extraction network includes an image patch embedding layer, a grouped image patch embedding layer corresponding to the bands of the hyperspectral image, and a multi-stage grouped spectral feature extraction network unit. Both the image patch embedding layer and the grouped image patch embedding layer are connected to one of the multi-stage grouped spectral feature extraction network units.

[0065] The present invention uses images in other bands except the RGB band as the input of a Multi-stage Grouped Spectral Feature Extractor (MGSFE), and extracts features from the images through multiple stages. The first stage consists of a common image patch embedding layer and MGSFE units. In the subsequent stages, in order to prevent interference between the feature information of different bands and maintain the independence of the feature information of each band, a grouped image patch embedding layer and MGSFE units are used. Specifically, the number of stages of the multi-stage grouped spectral feature extraction network can be determined according to the number of bands, and is at least 2. For example, setting the number of groups to the number of bands of the input image can enable the network to learn a unique data pattern for each band during training. The present invention is illustrated by taking four stages as an example.

[0066] Figure 5 It is a schematic structural diagram of the multi-stage grouped spectral feature extraction network unit provided by the present invention, as Figure 5 shown. The multi-stage grouped spectral feature extraction network unit includes a grouped convolution module, a layer normalization module, an activation function module, and a residual connection module connected in sequence. The residual connection module is used to connect the input end and the output end of the multi-stage grouped spectral feature extraction network unit.

[0067] In the MGSFE unit, the present invention mainly uses the grouped convolution module to extract features by grouping the bands, and uses pointwise convolution for spectral information fusion. The layer normalization module is to maintain the stability of the network training process, and the GELU activation function enhances the nonlinear characteristics of the network. In addition, a residual connection structure is introduced at the input and output positions of each unit. Through the MGSFE unit, unique feature information of each band can be extracted.

[0068] It can be understood that the present invention proposes a multi-stage grouped spectral feature extraction network, which can fully extract the information contained in different spectral bands.

[0069] Figure 6 It is a schematic structural diagram of the cross-band attention fusion network provided by the present invention, as Figure 6As shown, based on the above embodiments, as an alternative embodiment, the cross-band attention fusion network includes a first linear layer, a second linear layer, a third linear layer, and an attention module. The first linear layer is configured to perform feature transformation on the RGB band feature information to obtain feature query information. The second linear layer is configured to perform feature transformation on the band feature information to obtain feature key information. The third linear layer is configured to perform feature transformation on the band feature information to obtain feature value information. The attention module is configured to determine the attention score corresponding to each band feature information according to the feature query information and the feature key information, and obtain the fused band feature information according to the attention score and the feature value information.

[0070] The UniFormer and MGSFE networks can obtain the feature information of the RGB band and each other band. The present invention uses a cross-band attention fusion network (Cross-Band Attention Fusion, CBAF) to dynamically fuse the feature information of different bands / groups, realizing an adaptive band combination.

[0071] The CBAF network is mainly implemented based on the cross-attention mechanism. Since the cross-attention mechanism is good at processing the input sequences of two modalities, it can achieve a better feature aggregation effect. And because the fusion weights of the cross-attention mechanism are different according to different input data, it also realizes an adaptive feature fusion. Specifically, in the present invention, the RGB band feature information extracted by the UniFormer is subjected to feature transformation through a linear layer to obtain feature query information Query, and the features of other bands are respectively input into two linear layers to obtain feature key information Key and feature value information Value. Then, matrix multiplication is performed on Key and Query, and then through scaling and the Softmax function, the attention score of the band feature information is obtained. According to the obtained attention score, matrix multiplication is performed with Value again to obtain the fused feature. The formula expression of the above process is as follows:

[0072]

[0073] where Q, K, and V represent Queries, Keys, and Values respectively, and O c represents the output of the CBAF network. Using for scaling is to prevent gradient disappearance, ensure the stability of the network training process, and also realize the dynamic change of the attention score.

[0074] It can be understood that the present invention dynamically fuses the feature information of different bands / groups through a cross-band attention fusion network, fully excavates the complementary information between different bands, thereby achieving better classification performance and having a more powerful multi-spectral image understanding ability.

[0075] Based on the above embodiments, as an optional embodiment, the scene recognition model includes a classifier, and the classifier is used to classify the fused band feature information according to a preset single-label or multi-label classification task to determine a classification score; if the classifier has not been trained yet, then determine a loss value based on the classification score and a preset loss function, and update the parameters of the preset network structure, the cross-band attention fusion network, and the classifier according to the loss value until the classifier is trained; if the classifier is trained, then use the classification score as the scene recognition result.

[0076] Optionally, if the loss value reaches a set threshold or the number of training times is completed, it can be considered that the training is completed, and the way to update the parameters is backpropagation weight update.

[0077] The training of the preset network structure, the cross-band attention fusion network, and the classifier is obtained based on a multi-spectral image training set with labels. After training, multi-spectral image scene recognition can be achieved based on the preset network structure, the cross-band attention fusion network, and the classifier.

[0078] Specifically, the scene recognition model can be a classification head, and input the fused band feature information O c obtained by the CBAF network into the classification head, and perform final processing using the softmax or sigmoid function according to the single-label or multi-label classification task to obtain the classification score of each category, and then update the parameters or output the inference result according to the current training or inference stage.

[0079] It can be understood that the present invention inputs the fused feature information into the classification head to complete the final multi-spectral image scene recognition task, with higher classification accuracy and stronger multi-spectral image understanding ability.

[0080] The technical effects of the present invention will be described in detail below with reference to experimental data.

[0081] Existing network structures are basically designed for RGB images. When directly taking multispectral images as input, although existing networks can process multispectral images through techniques such as transfer learning, this simple transfer method still cannot fully utilize the spectral information of multispectral images. Based on the characteristics between different spectra of multispectral images, the present invention fully extracts the information contained in different spectral bands and dynamically fuses the feature information of different bands. It can fully exploit the complementary information between different bands, realize the mutual enhancement of different features, have a higher classification accuracy than existing methods, and enable the network to have a stronger multispectral image understanding ability.

[0082] The comparison of the classification performance quantification indicators of different methods on the BigEarthNet multispectral dataset is shown in Table 1 as follows.

[0083] Method mAP OP OR OF1 CP CR CF1 WRN - B4 - ECA - 82.4 75.5 78.8 - - - Resnet - 15(RGB) 81.7 78.1 65.3 71.1 64.5 52.8 56.6 Resnet - 152(ALWL) 88.0 83.0 72.9 77.6 71.7 60.2 64.1 ViT - L(RGB) 85.5 79.5 71.9 75.4 67.8 59.8 62.2 ViT - L(ALL) 89.0 80.9 77.8 79.3 71.7 65.8 67.4 Swin - B(RGB) 87.6 81.8 73.6 77.5 70.7 60.7 64.0 Swin - B(ALL) 89.6 83.3 76.4 79.7 73.0 64.8 67.5 The method of the present invention 89.7 85.7 81.3 83.4 77.6 70.8 73.1

[0084] Table 1

[0085] Among them, WRN-B4-ECA, Resnet-152(RGB), Resnet-152(ALL), ViT-L(RGB), ViT-L(ALL), Swin-B(RGB), Swin-B(ALL) are existing multispectral image scene recognition methods. The evaluation indicators mAP, OP, OR, OF1, CP, CR, CF1 respectively represent the mean average precision, overall precision, overall recall rate, overall F1-score, class precision, class recall rate, and class F1-score. RGB and ALL in the brackets respectively represent classifying using only the RGB band and using all bands.

[0086] The comparison of the classification performance quantification indicators of different methods on the EuroSAT multispectral dataset is shown in Table 2 as follows.

[0087]

[0088] Table 2

[0089] Among them, the evaluation indicators Acc, P, R, F1 respectively represent the classification accuracy rate, precision, recall rate, and F1-score. Resnet-152(RGB), Resnet-152(ALL), ViT-L(RGB), ViT-L(ALL), Swin-B(RGB), Swin-B(ALL) are used as comparison methods. RGB and ALL in the brackets respectively represent classifying using only the RGB band and using all bands.

[0090] Based on the experimental data in Table 1 and Table 2, it can be seen that the present invention has better classification performance.

[0091] It should be noted that the execution subject of the task construction method provided by the present invention can be a server or a computer device, such as a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc.

[0092] Figure 7 is a schematic structural diagram of the multi-spectral image scene recognition device provided by the present invention. As Figure 7 shown, the present invention also provides a multi-spectral image scene recognition device, including:

[0093] A grouping module 710, configured to obtain a multi-spectral image to be recognized, group the multi-spectral image according to the bands of the multi-spectral image, and obtain at least one multi-spectral image group;

[0094] An extraction module 720, configured to perform feature extraction on the multi-spectral image group based on a preset network structure corresponding to the bands of the multi-spectral image group, and obtain band feature information corresponding to each multi-spectral image group;

[0095] A fusion module 730, configured to perform feature fusion on the band feature information based on a preset cross-band attention fusion network to obtain fused band feature information;

[0096] An identification module 740, configured to input the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result.

[0097] It should be noted that the multi-spectral image scene recognition device provided by the present invention, during specific operation, can execute the multi-spectral image scene recognition method described in any of the above embodiments, and details thereof are not elaborated in this embodiment.

[0098] Figure 8 is a schematic structural diagram of the electronic device provided by the present invention. As Figure 8As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 may call logic instructions in the memory 830 to execute a multi-spectral image scene recognition method, which includes: obtaining a multi-spectral image to be recognized, grouping the multi-spectral image according to the bands of the multi-spectral image to obtain at least one multi-spectral image group; performing feature extraction on the multi-spectral image group based on a preset network structure corresponding to the bands of the multi-spectral image group to obtain band feature information corresponding to each multi-spectral image group; performing feature fusion on the band feature information based on a preset cross-band attention fusion network to obtain fused band feature information; inputting the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result.

[0099] In addition, when the logic instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0100] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the multi-spectral image scene recognition method provided in each of the above embodiments. The method includes: acquiring a multi-spectral image to be recognized, grouping the multi-spectral image according to the bands of the multi-spectral image to obtain at least one multi-spectral image group; extracting features from the multi-spectral image group based on a preset network structure corresponding to the bands of the multi-spectral image group to obtain band feature information corresponding to each multi-spectral image group; performing feature fusion on the band feature information based on a preset cross-band attention fusion network to obtain fused band feature information; inputting the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result.

[0101] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the execution of the multi-spectral image scene recognition method provided in each of the above embodiments. The method includes: acquiring a multi-spectral image to be recognized, grouping the multi-spectral image according to the bands of the multi-spectral image to obtain at least one multi-spectral image group; extracting features from the multi-spectral image group based on a preset network structure corresponding to the bands of the multi-spectral image group to obtain band feature information corresponding to each multi-spectral image group; performing feature fusion on the band feature information based on a preset cross-band attention fusion network to obtain fused band feature information; inputting the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result.

[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multispectral image scene recognition method, characterized in that, it includes: Obtain a multispectral image to be recognized, group the multispectral image according to the bands of the multispectral image to obtain at least one multispectral image group; Extract features from the multispectral image group based on a preset network structure corresponding to the bands of the multispectral image group to obtain band feature information corresponding to each multispectral image group; Perform feature fusion on the band feature information based on a preset cross-band attention fusion network to obtain fused band feature information; Input the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result; The extracting features from the multispectral image group based on a preset network structure corresponding to the bands of the multispectral image group to obtain band feature information corresponding to each multispectral image group includes: Extract features from the multispectral image group in the RGB band based on a preset backbone network to obtain RGB band feature information; Extract features from the multispectral image group in non-RGB bands based on a preset multi-stage grouped spectral feature extraction network to obtain band feature information corresponding to each multispectral image group; Use the images in bands other than the RGB band as the input of the multi-stage grouped spectral feature extraction network MGSFE, extract features from the images through multiple stages. The first stage consists of an image patch embedding layer and an MGSFE unit, and the other stages consist of grouped image patch embedding layers and MGSFE units. The number of stages of the multi-stage grouped spectral feature extraction network can be determined according to the number of bands; In the MGSFE unit, use a grouped convolution module to extract grouped features of the bands, and use pointwise convolution to fuse spectral information; The cross-band attention fusion network is implemented based on a cross-attention mechanism. The RGB band feature information extracted by the backbone network is transformed through a linear layer to obtain feature query information. The features of other bands are respectively input into two linear layers to obtain feature key information and feature value information. Perform matrix multiplication on the feature key information and the feature query information, and through scaling and the Softmax function, obtain the attention score of the band feature information. According to the obtained attention score, perform matrix multiplication with the feature value information again to obtain the fused feature.

2. The multispectral image scene recognition method according to claim 1, characterized in that, The multi-stage grouped spectral feature extraction network includes an image patch embedding layer, a grouped image patch embedding layer corresponding to the bands of the multispectral image, and a multi-stage grouped spectral feature extraction network unit. Both the image patch embedding layer and the grouped image patch embedding layer are connected to a multi-stage grouped spectral feature extraction network unit; the multi-stage grouped spectral feature extraction network unit includes a grouped convolution module, a layer normalization module, an activation function module, and a residual connection module connected in sequence. The residual connection module is used to connect the input end and the output end of the multi-stage grouped spectral feature extraction network unit.

3. The multispectral image scene recognition method according to claim 1, characterized in that, The backbone network includes at least one local feature extraction unit and at least one global feature extraction unit. The local feature extraction unit includes a first position encoding module, a local multi-head relationship aggregation module, and a first feed-forward neural network module arranged in sequence. A batch normalization module is provided between the first position encoding module and the local multi-head relationship aggregation module, and a batch normalization module is provided between the local multi-head relationship aggregation module and the first feed-forward neural network module. The global feature extraction unit includes a second position encoding module, a global multi-head relationship aggregation module, and a second feed-forward neural network module arranged in sequence. A layer normalization module is provided between the second position encoding module and the global multi-head relationship aggregation module, and a layer normalization module is provided between the global multi-head relationship aggregation module and the second feed-forward neural network module.

4. The multi-spectral image scene recognition method according to claim 1, wherein, the cross-band attention fusion network includes a first linear layer, a second linear layer, a third linear layer, and an attention module. The first linear layer is used to perform feature transformation on the RGB band feature information to obtain feature query information. The second linear layer is used to perform feature transformation on the band feature information to obtain feature key information. The third linear layer is used to perform feature transformation on the band feature information to obtain feature value information. The attention module is used to determine the attention score corresponding to each band feature information according to the feature query information and the feature key information, and obtain the fused band feature information according to the attention score and the feature value information.

5. The multi-spectral image scene recognition method according to claim 1, wherein, the scene recognition model includes a classifier. The classifier is used to classify the fused band feature information according to a preset single-label or multi-label classification task to determine a classification score. If the classifier has not been trained yet, then based on the classification score and a preset loss function, a loss value is determined, and the parameters of the preset network structure, the cross-band attention fusion network, and the classifier are updated according to the loss value until the classifier is trained. If the classifier is trained, then the classification score is used as the scene recognition result.

6. The multi-spectral image scene recognition method according to claim 1, wherein, before grouping the multi-spectral image according to the bands of the multi-spectral image, it further includes: upsampling the multi-spectral image to complete the preprocessing of the multi-spectral image.

7. A multi-spectral image scene recognition device, wherein, it includes: a grouping module, configured to obtain a multi-spectral image to be recognized, and group the multi-spectral image according to the bands of the multi-spectral image to obtain at least one multi-spectral image group; an extraction module, configured to perform feature extraction on the multi-spectral image group based on a preset network structure corresponding to the bands of the multi-spectral image group to obtain the band feature information corresponding to each multi-spectral image group; A fusion module, configured to perform feature fusion on the band feature information based on a preset cross-band attention fusion network to obtain fused band feature information; An identification module, configured to input the fused band feature information into a pre-trained scene recognition model to obtain a scene recognition result; The method for extracting, based on a preset network structure corresponding to the bands of the multi-spectral image group, the band feature information corresponding to each multi-spectral image group includes: Performing feature extraction on the multi-spectral image group in the RGB band based on a preset backbone network to obtain RGB band feature information; Performing feature extraction on the multi-spectral image group in non-RGB bands based on a preset multi-stage grouped spectral feature extraction network to obtain the band feature information corresponding to each multi-spectral image group; Taking the images in bands other than the RGB band as the input of the multi-stage grouped spectral feature extraction network MGSFE, and performing feature extraction on the images through multiple stages. The first stage is composed of an image patch embedding layer and an MGSFE unit, and the other stages are composed of a grouped image patch embedding layer and an MGSFE unit. The number of stages of the multi-stage grouped spectral feature extraction network can be determined according to the number of bands; In the MGSFE unit, a grouped convolution module is used to perform grouped feature extraction on the bands, and pointwise convolution is used for spectral information fusion; The cross-band attention fusion network is implemented based on a cross-attention mechanism. The RGB band feature information extracted by the backbone network is subjected to feature transformation through a linear layer to obtain feature query information. The features of other bands are respectively input into two linear layers to obtain feature key information and feature value information. Matrix multiplication is performed on the feature key information and the feature query information, and after scaling and the Softmax function, the attention score of the band feature information is obtained. According to the obtained attention score, matrix multiplication is performed with the feature value information again to obtain the fused feature.

8. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the multi-spectral image scene recognition method according to any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the multi-spectral image scene recognition method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Gender identification method and system based on multispectral fusion, storage medium and terminal

    CN111695407A